What we found
Not one of the 30 homepages scored an A. The highest was salesforce.com at 83, and only 2 sites reached a B. Of the 21 that answered with a normal page, the average was 62 out of 100 — a D on this scale.
The second finding is sharper. 9 of the 30 domains do not admit at least one of GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot or Google-Extended at their site root — 7 of them say so in robots.txt, and 2 refuse the request at the server before robots.txt is even read. Every news publisher in the sample is in that group: 6 of 6. The publishing industry has not drifted into blocking AI crawlers by accident — it has chosen to.
The third finding is about the companies selling AI. OpenAI's own homepage scores 13 and Anthropic's 62, both below the average of the sample, and OpenAI's is the sharper case: a crawler-shaped request to openai.com was refused with HTTP 403 in the same run that served a browser-shaped request to the same URL normally. Neither declares an Organization entity. The sites that score best are the ones selling infrastructure to developers, not the ones selling the answers.
Why the low scores matter more than they look
These are homepages, not articles, and homepages are the hardest page on any site to score well on: they are short, they are navigational, and they carry little of the evidence a generated answer needs. A low score here does not mean a site is invisible in AI answers — their articles may be in far better shape.
What it does mean is that the page an AI system lands on first gives it almost nothing to quote. No statistics, no quotation, usually no question-shaped heading and often no entity markup that says which organisation the domain belongs to. On 22 of the 30, a model reading the homepage has no structured statement of who owns the site.
Three different things can stop a crawler
They are easy to confuse, so the table separates them. Disallowed means robots.txt disallows an AI agent at the root — a published, deliberate policy that any crawler will obey. Refused means the server answered a crawler-shaped request with an error while serving a browser-shaped request to the same URL normally: the block happens before robots.txt is consulted, and because GPTBot, ClaudeBot, PerplexityBot and OAI-SearchBot all present as bots, a rule like that turns away every engine this page is about. 2 domains in this table are that case.
A refusal of both request shapes is a third thing and is not counted as a block. It says the request was turned away from the network this scan runs on, which is evidence about the scanner rather than about AI access. 4 domains are in that position, and the table records them as neither disallowed nor refused.
The results
How this was done
- The 30 domains were chosen by hand to cover search, AI labs, enterprise software, infrastructure, developer platforms, news and German industry. This is a convenience sample, not a random one, and it is not a ranking of the largest sites on the internet.
- Each homepage was fetched by the LLMention crawler from its own infrastructure, identifying itself honestly as a bot. It did not impersonate an AI crawler and it was not sent from an AI crawler’s address. Where that request was refused, the same URL was asked once more with a browser-shaped User-Agent, so that a rule aimed at identified bots could be told apart from an address being blocked. That second request never scores page content; it decides one check, and the rule behind it is published on the methodology page.
- robots.txt, llms.txt and sitemap.xml were requested from the same origin, and the 38 documented checks were run over the result.
- No AI engine was queried. Nothing on this page is a measurement of what ChatGPT, Claude or Perplexity actually say — that is a different experiment and this tool does not run it.
Limits of this study
- Homepages only. A site’s articles are usually in better shape than its homepage, so these scores understate most of the sites listed.
- One moment in time. robots.txt and markup change constantly; the table is a snapshot, and every row can be re-run.
- Geo-dependence. Sites that redirect by visitor location returned a regional version, and the regional version differs. stripe.com returned its Netherlands edition during one run and its Singapore edition during another.
- A refusal of both request shapes proves nothing about AI access. When the scanner is turned away as a bot and as a browser alike, the block is about the network it came from, and the row says so rather than blaming the site. What a refusal can show is the opposite case — a browser served, a bot refused — and that is a rule aimed at exactly the crawlers this page measures.
- A refusal can vary between requests. During this re-run, one site answered the browser-shaped request with 200 and, minutes later, with the same error it gave the crawler. A “refused” mark is therefore one observation at one moment, not a stable property of the site, and the row may look different on re-run.
- Not a prediction. These are readiness checks, not a measurement of citations. A site can score 12 and still be quoted tomorrow.
Reproduce it
Every row above comes from one public endpoint, and brief=1 returns just the headline numbers:
curl "https://geo-scanner.ccie13192.com/api/scan?domain=github.com&brief=1"
{"domain":"github.com","status":200,"browserStatus":null,
"scoreBasis":"homepage","score":60,"grade":"D",
"checksRun":38,"checksPassed":22,"aiCrawlersBlocked":false,
"hasJsonLdEntity":false,"hasSameAs":false,"hasRobotsTxt":true,
"topIssue":"robots-sitemap"}The rules behind every number are published on the methodology page, including the parts of the picture the score cannot see. If a result here looks wrong, that is a bug report and it is welcome — the check will be fixed rather than the number.
Questions about this study
How were the 30 sites chosen?
They are a convenience sample of large, well-known homepages, not a random one. That is enough to show a pattern and not enough to generalise from, which is why the sample is described rather than implied.
Why do so many large sites score badly?
A homepage is the hardest page on any site to score well on: it is short, it is navigational, and it carries little of the evidence a generated answer needs. Their article pages may be in far better shape.
Was any figure adjusted by hand?
No. Every number in the table came from the same scanner on one date, and each row links to a report that can be re-run against the public API.
Evidence and sources
Adding source citations produced the largest measured visibility gain for low-ranking sites, at +115%, ahead of the addition of expert quotations at +41% and statistics at +30-40%, across the strategies tested on generative engines. — Generative Engine Optimization, KDD 2024
The weightings on this site follow that measurement rather than taste, and the parts of the picture a single-URL scan cannot see are stated rather than left out.
Primary sources