Instagram tells every AI model not to read your business
Instagram's robots file names the crawlers of OpenAI, Anthropic, Google, Perplexity, Apple and Amazon and blocks every one of them from the entire site. The live fetch that is not blocked returns about 939 kilobytes and roughly fifty characters of readable business fact. No address, no prices, no hours.
QUICK ANSWER
Instagram's robots file names GPTBot, ClaudeBot, Google-Extended, PerplexityBot, Applebot-Extended and Amazonbot, and blocks all of them from the entire site. If your business exists only on Instagram, Meta has instructed every major AI model not to read it, and you were not asked. Fetching a profile the way a model would returns about 939 kilobytes and roughly 50 characters of readable business fact. No address, no prices, no hours.
Open instagram.com/robots.txt in any browser. It is a public file and it takes ten seconds.
Near the top, by name, are the crawlers belonging to OpenAI, Anthropic, Google, Perplexity, Apple and Amazon. Each one is followed by Disallow: /, which means the whole site, every profile on it, including yours.
That is not a setting in your account. It is a decision Meta made on behalf of every business on the platform, and there is no version of your profile where it does not apply.
Instagram's own robots file blocks every major AI crawler by name
Here is the file, read on 6 September 2026. The relevant lines are these, verbatim:
User-agent: Amazonbot Disallow: /
User-agent: Applebot-Extended Disallow: /
User-agent: ClaudeBot Disallow: /
User-agent: Google-Extended Disallow: /
User-agent: GPTBot Disallow: /
User-agent: PerplexityBot Disallow: /
The header above them states the position plainly: collection of data on Instagram through automated means is prohibited without express written permission from Instagram.
This is a defensible thing for Meta to do. Its data is its asset. The point is only that the decision is theirs and the consequence is yours.
[ 01 / instagram.com/robots.txt ]
Named, and blocked from the entire site
Read on 6 September 2026. It is a public file and it takes ten seconds to check.
Blocked with Disallow: /
- GPTBot · OpenAI, training
- ClaudeBot · Anthropic, training
- Google-Extended · Google, AI training
- PerplexityBot · Perplexity, search index
- Applebot-Extended · Apple, AI training
- Amazonbot · Amazon
- Brightbot
Not named in the file
- The live agents that fetch a page in the moment somebody asks a question. These are separate programs with separate names and Instagram does not list them.
So the block is not total, and it does not need to be. The live fetch is permitted and returns almost nothing, which figure 2 measures.
instagram.com/robots.txt, read 6 September 2026. The file's own header states that automated collection is prohibited without express written permission from Instagram.
Meta made that decision for you, and you cannot override it
You cannot edit that file. You cannot exempt your own profile from it. There is no toggle in a business account that says allow AI models to read my opening hours.
The same is true of a Facebook page, and of any presence that lives inside a platform rather than on a domain you control. The rules that govern whether a machine may read your business are written by the company that owns the building.
That is a different problem from the one in Instagram is where they remember you, it is not where they decide. That piece was about people. This one is about the machines that people increasingly ask first.
A profile is 939 kilobytes and about fifty characters of business fact
The blocking is only the visible half. The other half is what is actually there to read.
Fetching four large public profiles the way a model's live browser would, then stripping the scripts and styles and reading what text remains, gives this:
- NASA: 939,018 bytes downloaded, 42 characters of extractable text
- National Geographic: 939,316 bytes, 59 characters
- Starbucks: 939,481 bytes, 52 characters
- Four Seasons Hotels and Resorts: 939,831 bytes, 74 characters
In every case the extractable text is the page title and nothing else. On the NASA profile, 840,087 bytes of the download is JavaScript. Not one of the four carries any structured data.
The description tag a machine would fall back to reads, in full: 2M Followers, 333 Following, 62 Posts. That is Four Seasons. A global hotel group, and what a machine gets is a follower count.
[ 02 / Fetched the way a model would ]
Roughly 939 kilobytes in. Under 75 characters out.
Four of the largest public profiles on the platform, scripts and styles stripped, remaining text counted.
@nasa · from 939,018 bytes
42
@natgeo · from 939,316 bytes
59
@starbucks · from 939,481 bytes
52
@fourseasons · from 939,831 bytes
74
structured data records, all four
0
of the NASA page is JavaScript
840 KB
In every case the extractable text is the page title and nothing else. The description a machine falls back to reads, in full for Four Seasons: 2M Followers, 333 Following, 62 Posts. A global hotel group, and what a machine gets is a follower count.
Each profile fetched 6 September 2026 with a live agent user string, script and style blocks removed, remaining text counted. Byte counts are a snapshot of a page Meta rebuilds constantly. The ratio is the finding, not the exact figure.
Blocking is only half of it. The live fetch gets nothing either.
There is a distinction here that most coverage flattens, and it matters because it changes the fix.
The crawlers Instagram names are the ones that collect training data and build search indexes. The agents that fetch a page live, in the moment somebody asks a question, are different programs with different names, and Instagram does not list them.
So a live fetch is not blocked. It was how the measurements above were taken. It returns 939 kilobytes and about fifty characters.
Both roads end in the same place. One is closed by policy and the other is open and empty. A business that is invisible for two unrelated reasons is not less invisible.
Even when a model does find you, fewer people click
This is the part that gets oversold in both directions, so here are the numbers that actually have a study behind them.
The Pew Research Center tracked the browsing of 900 United States adults through March 2025, covering 68,879 unique Google searches, of which 12,593 produced an AI summary. On visits where a summary appeared, users clicked a traditional search result in 8% of cases. Where no summary appeared, 15%. Clicking a link inside the summary itself happened in 1% of visits.
That is United States data about Google specifically, and it should not be stretched into a claim about every model in every market. What it establishes is direction: the answer increasingly arrives without the visit.
Cloudflare, measuring its own network in the first week of August 2025, published the ratio of pages crawled to referrals sent back: roughly 50,000 to 1 for Anthropic, 887 to 1 for OpenAI, 118 to 1 for Perplexity. Content is being read at a scale that has never matched the traffic returned.
For a hotel this is not a catastrophe. It is a change in what the website is for, and it argues in the same direction as everything else: being the source the answer is built from matters more than being the eleventh blue link.
Nobody can prove that markup gets you recommended, and we will not claim it
Here is where most articles on this subject start selling something. The honest state of the evidence does not support it.
No statement from Google, OpenAI, Anthropic or Schema.org asserts that adding structured data measurably improves an AI model's odds of citing your page. Google's own position, stated publicly in July 2025, is that ordinary search practice is what applies. There is no separate lever.
The llms.txt file, proposed in 2024 and sold hard since, is worse than unproven. Ahrefs analysed 137,210 domains in May 2026 and found that 97% of published llms.txt files received zero requests from anything at all. Of the small remainder that saw any traffic, most of it was SEO audit tools rather than AI systems. Google has said publicly it does not use the file and has no plans to.
Anyone selling an AI visibility package built on those two things is selling a mechanism nobody has demonstrated.
The websites are better than the profiles, but not by as much as you would hope
Instagram is the easy target. The fair question is what a machine finds on the actual hotel websites here, so it seemed better to count than to assume.
Forty four hotel websites in North Cyprus were checked on 6 September 2026, drawn from the Cyprus Turkish Hotels Association membership and verified as northern property by property. Thirty three homepages could be fetched and read.
Thirteen of those thirty three carry any structured data at all. One describes itself to a machine as a hotel. The rest of the structured data present is the kind a content management system emits on its own: a breadcrumb trail, a site name, a search box. Useful plumbing, and it says nothing about rooms, rates, location or season.
Twenty of the thirty three carry a meta description, the one line a machine falls back to when nothing else is offered. Thirteen do not.
Then the detail that says the most about how these sites were built. Three of them declare the wrong language outright in their page code: two announce themselves as Indonesian and one as Vietnamese. Those are template defaults that were never changed, and every browser, translation tool and machine reader takes them at their word.
Thirty of the forty four publish a robots file at all. Three name any AI crawler in it, in either direction. The other twenty seven have not made a decision, which is different from having made the wrong one, and much easier to fix.
Nobody has measured this before, for this island or for any small tourism market. The numbers above are ours.
[ 03 / 44 hotel websites in the north ]
Better than a profile. Not by as much as you would hope.
Checked 6 September 2026. Thirty three of the forty four homepages could be fetched and read, and the readability figures are out of those thirty three.
of 33 carry any structured data
13
describes itself as a hotel
1
of 33 have a meta description
20
declare the wrong language entirely
3
of 44 publish a robots file
30
name any AI crawler in it
3
What the structured data actually contains
A breadcrumb trail, a site name, a search box. What a content management system emits on its own. Nothing about rooms, rates, location or season.
The three wrong language declarations
- 2 announce themselves as Indonesian
- 1 announces itself as Vietnamese
- Template defaults nobody changed. Every browser, translation tool and machine reader takes them at their word.
Twenty seven of the thirty robots files name no AI crawler at all. That is not the wrong decision, it is no decision, and it is much easier to fix than the alternative.
Hotels drawn from the Cyprus Turkish Hotels Association membership and verified as northern property by property. Homepages fetched and parsed 6 September 2026. A site counts as carrying structured data if any is present, whether or not its contents are correct, so the count is generous. No hotel is named.
The only surface whose rules you control is one you own
Strip out everything unproven and a short list of things survives, all of them boring and all of them verifiable.
A machine can read text. It cannot read a photograph of your menu, and it cannot read a profile that returns fifty characters. It can only fetch what a robots.txt permits, and the only robots.txt you are allowed to write is the one on your own domain.
So the fix is not a trick and not a package. It is that the facts about your business need to exist as text, on pages you control, at addresses that stay put: what you sell, where you are, what it costs, when you are open, in what languages, for whom.
That is not an AI strategy. It is a website, built so that a machine reading it comes away knowing what you do.
What this is actually worth doing about
The free version first, because it is real. Go and read your own site the way a machine does. Turn off images in your browser and see what is left. If the answer is a logo and three photographs, a model has nothing to work with, and neither does a customer on a slow connection.
The ceiling is where that stops being enough. You cannot restructure a site that has no pages, only sections of one long scrolling image. You cannot put your rooms and rates into text if they live inside a PDF or a photograph. You cannot control a robots file on a domain that is not yours. Those are build problems and no amount of tidying gets past them.
What we do about it is unglamorous and specific: a page per thing you sell, the facts as text rather than as pictures of text, the structured data that describes a business filled in honestly, and a robots file that says yes on purpose rather than by default.
We will not promise you a mention in an answer. Nobody who is honest can. What can be promised is that when a machine comes to read your business, there will be something there to read, which today is not the case for most of the island.