Original research

ChatGPT SEO: can it actually read your website?

ChatGPT SEO starts with whether a machine can read you at all. I checked 474 UK trade websites Google already ranks. The blocker is not the one people fear.

6 min read

In the last study I found that when you ask ChatGPT to recommend a tradesperson, it reads directories rather than businesses. The obvious next question is whether it could read the businesses even if it wanted to. So on 25 August 2026 I took 474 UK trade websites that Google already ranks in the local pack and checked.

Almost nothing is blocking the crawlers. Eleven sites out of 445, that is 2.5 per cent, block an AI crawler at all, and ten of those eleven did it by flicking a switch at their host rather than deciding anything.

The real problem is duller and much worse. One site in six has no phone number a machine can read.

The short version

  • Sample: 474 trade business websites, from 200 local searches across 10 trades and 20 UK towns, checked 25 August 2026.
  • Blocking AI crawlers: 11 sites, 2.5 per cent. Ten carried a hosting provider’s signature.
  • No phone number in the text: 75 sites, 16.9 per cent.
  • No readable content in the HTML: 31 sites, 7 per cent, because the page is built by JavaScript.
  • LocalBusiness schema: 132 sites, 29.7 per cent.
  • llms.txt: 115 sites have one, and 76 of those name the plugin or platform that wrote it.

Bar chart: of 445 trade sites Google ranks, 75 have no phone number in the text, 31 have no readable content in the HTML, and 11 block an AI crawler

445 reachable sites, checked 25 August 2026.

The blocking panic is not real

There has been a lot written about businesses accidentally shutting themselves out of AI search through robots.txt. In this sample it barely happens.

86.5 per cent have a robots.txt at all, and only 5.8 per cent mention any AI crawler in it. Where a block does exist it is almost never a decision: ten of the eleven blockers carry the same fingerprint, a list including CloudflareBrowserRenderingCrawler alongside GPTBot, ClaudeBot and Google-Extended, which is what you get when the “block AI scrapers” toggle is switched on at Cloudflare. Two were hand rolled or came from a plugin.

If you are worried you have accidentally blocked ChatGPT, you almost certainly have not. Check it in thirty seconds by opening yoursite.co.uk/robots.txt and searching for GPTBot. Then stop worrying about it and go and look at your phone number.

One thing worth knowing if you do find a block. Google-Extended is not Google Search. Blocking it does not affect your ranking. It governs whether your content can be used to ground Gemini and AI Overviews. Those are different taps and people conflate them constantly.

The thing that actually breaks it

Seventy-five of 445 sites, 16.9 per cent, have no phone number in the readable text of the homepage. It is in a header image, or a graphic, or loaded by a script after the page arrives.

A person squints and finds it. A model reading the HTML cannot pass on the one piece of information the customer actually asked for. You can be ranked first by Google, cited by name, and still be the recommendation nobody can act on.

Another 31 sites, 7 per cent, return effectively nothing at all to a plain fetch. The median site in this sample returns 773 words of readable text. These returned under 120, and several returned under five, because the page is assembled by JavaScript in the browser. Google runs that JavaScript. Most other fetchers do not.

Add the six businesses whose Google listing points at a Facebook page or a directory entry instead of a website, and the shape of the problem is clear. It is not censorship, it is plumbing.

Nobody wrote their own llms.txt

Here is the finding I did not expect. 115 of these sites have an llms.txt, the emerging convention for telling a model what a site is about. That is 25.8 per cent, which would be remarkable adoption for small trade businesses.

They did not adopt it. Seventy-six of those 115 files name the tool that generated them.

Bar chart of what generated 115 llms.txt files: Wix 43, All in One SEO 15, Yoast 8, Rank Math 7, WordPress 3, and 40 naming no generator

Every llms.txt found, by what wrote it.

Generator Files
Wix 43
All in One SEO 15
Yoast and Rank Math 15
No generator named 39

Wix alone accounts for more than a third. The rest are SEO plugins that shipped the feature in an update. Most of the 39 that name nothing follow the identical house format, a heading and a one line summary, which suggests the same story.

The two AI-specific things on these websites, the llms.txt files and the crawler blocks, were both put there by somebody else’s software. On the evidence here, essentially no trade business has made a deliberate decision about AI either way.

That cuts both ways, and it is worth being honest about. It means very few have shot themselves in the foot. It also means the sites doing well out of it are doing well by accident, and anybody selling “AI optimisation” as a package of files is selling something your website builder already did for free.

What is actually worth doing

In the order I would do it, and none of it is exotic.

  • Put the phone number in the text. Not in the logo, not in a graphic. Real characters, on the page, ideally in a tel: link. This is the single highest-value fix in the whole study and it takes ten minutes.
  • Check the page has words in it with JavaScript off, or just view source and search for a sentence you know is on the page. If it is not there, a lot of things cannot read you, and that is a build problem worth fixing properly.
  • Add LocalBusiness schema. Only 29.7 per cent have it, so it is still a genuine differentiator, and it states your name, address, phone and hours in a format nothing has to guess at. It is part of the same job as getting the listing right.
  • Look at robots.txt once, confirm you are not blocking GPTBot, and then leave it alone.
  • Do not buy an llms.txt. If you are on Wix or running a modern SEO plugin you already have one.

None of that guarantees you get named by anything. The last study found the models mostly read directories anyway. What this does is remove the reasons you would be skipped when something does come looking, which is the part you control.

Method, and what it does not show

200 searches, 10 trades across 20 UK towns, top 20 organic plus the local pack, on 25 August 2026. That produced 587 local pack businesses, 515 with a website, 480 unique domains. Six pointed at a social or directory page and were counted separately, leaving 474 own websites of which 445 responded.

Each site got its homepage, its robots.txt and its llms.txt fetched, with the crawler identified and robots.txt obeyed. A site counted as having no readable content below 120 words of text after stripping scripts and tags. Phone numbers were matched on UK formats including the +44 form.

Three limits.

  • A homepage is not a website. A business might carry its phone number on a contact page and not the front one. That is better than nothing and still worse than having it on both.
  • “Blocked” means a full Disallow for that named crawler. Partial disallows were recorded separately and are not counted as blocks.
  • This is what a simple fetcher sees. Some AI systems render JavaScript, and the ones that do will see more than this study did. The gap between the two is exactly the risk being measured.

Every figure here was checked by hand before it went in. The llms.txt number looked implausible at 25 per cent for small trade businesses, which is why I refetched all 115 files and found the generators, and that turned a boring adoption statistic into the most interesting thing in the study.

I will run it again in twelve months against the same 200 searches. The number I will be watching is the phone number one, because it is the one that costs a business real work today, not in some predicted future.

Where you stand

Find out where you actually rank.

Send me your business name and the areas you cover, and I’ll check where you’re showing up across your area and where you’re not. There’s no charge for it, and if you’re already doing well I’ll tell you that.