Notes from the work

ChatGPT almost never opens your page. Here is what it reads instead

New log research maps ChatGPT's retrieval stack: its own index, a shared read cache, and rare live page opens. What it means for AI search visibility.

John KyprianouJohn Kyprianou9 min readUpdated

Most advice about getting cited in ChatGPT assumes the model reads your page. New log research says that is the rarest thing it does.

In a sample of 61,332 URLs that ChatGPT retrieved, only 759 were actually opened and read (Search Engine Land). Everything else was judged on a title and roughly 200 characters of text.

Diagram of ChatGPT's three retrieval layers, its own index, a shared read cache and rare live page opens, by SEO Turtle

The three layers, and why the distinction matters

The research, published on 17 August 2026 by Olivier de Segonzac and based on log analysis by the Resoneo team, splits ChatGPT's retrieval into three separate systems that most people treat as one thing (Resoneo).

The index. OpenAI's own search index, which the researchers found internally referred to as "labrador". It holds a URL, the full page title, and a snippet of roughly 200 characters anchored on your H1.

The read cache. A shared store of full pages converted to Markdown. When ChatGPT needs the body of your page, this is usually what it gets, not your live server.

The live open. The model sends a fetch to your actual URL. This is rare, and it is the layer that decides citations.

Those are different systems with different crawlers, different refresh behaviour and different content. Treating them as one pipeline is why so much GEO advice misfires.

Layer one: ChatGPT's index is not Bing's index

This is the finding that should retire a tired piece of advice.

The researchers compared the URLs in OpenAI's index against Bing results for the same queries and found that only 1.5 percent of them appeared in Bing's top 20. OpenAI is running its own crawl and its own index, not passing your site through Microsoft.

We flagged the softer version of this recently when we wrote about the second AI answer network running on Microsoft's index. That post still holds: Copilot, DuckDuckGo and Yahoo Scout genuinely do depend on Bing. What it does not buy you is ChatGPT. Fixing Bing Webmaster Tools is worth an afternoon, and it is not a ChatGPT strategy.

The practical detail buried in the crawl data is more useful anyway. OAI-SearchBot, the crawler OpenAI says "is used to surface websites in search results in ChatGPT's search features," discovers pages by following links (OpenAI). URLs that exist only in your XML sitemap were not reliably picked up.

If a page on your site is reachable only from the sitemap, or sits four clicks deep behind a paginated archive, it is not in the conversation. Internal linking is doing the discovery work here, not your sitemap submission.

Layer two: the cache serves an old copy of your page, on purpose

The read cache is where the mechanics get genuinely strange, and where most people's mental model breaks.

A cached copy is treated as fresh for about 30 minutes. Past that, ChatGPT still serves the stale copy to the next user, then refreshes in the background so the following request gets the new version. Retention runs for months rather than days, with the researchers seeing a copy from 11 July still served on 23 July.

Two consequences worth sitting with.

Your fix is not live when you think it is. You rewrite a page, you ask ChatGPT about it, and you get the old text back. That is not the model being wrong. That is you reading a cached copy. Any before-and-after test you run on the same day is close to meaningless.

Blocking directives that work everywhere else do not work here. The cache ignores Cache-Control: no-store, and pages carrying meta robots: noindex were still cached and served. If you have a staging URL, a thin location page or an old pricing page you quietly noindexed rather than removed, assume it is readable. Removing it or gating it properly is the only reliable answer, which is a slightly different job from the one covered in our AI crawlers and robots.txt guide.

There is also a hard ceiling: pages over 4 MB were rejected outright. That is not a theoretical limit for a modern site with hero video posters, uncompressed images and a fat client bundle.

Layer three: opens are rare, and they decide everything

Here is the asymmetry that reframes the whole exercise.

Pages that ChatGPT actually opened were cited 74 percent of the time. Pages that were merely retrieved from the index were cited 7 percent of the time. Opens happened for roughly one page in eighty.

So there are two completely different games. Winning the retrieval game gets you into a pool where nine out of ten candidates get discarded, which lines up with earlier work showing 85 percent of retrieved sources never make the final answer. Winning the open is close to winning the citation.

What triggers an open is mostly outside your control. Free instant responses lean entirely on index snippets and never open a page. Paid thinking mode does the real fetching, and the Think button free users got in August opens well under one page per conversation in the same logs. You cannot make a user upgrade their plan.

What you can control is whether the snippet that represents you is worth opening, and whether the page survives being opened.

Our take: this is the least glamorous GEO advice you will ever read

We will state the opinion plainly, because it cuts against a lot of what is being sold.

If the model's view of your page is a full title and about 200 characters after your H1, then the highest-leverage work in generative engine optimisation is writing a good title and a good opening paragraph. That is it. That is the lever.

No schema markup fixes a vague opening sentence. No llms.txt file rewrites your first 200 characters, which is part of why we called that standard mostly SEO theatre. An entity optimisation subscription does not help when the snippet representing your page says "We are passionate about delivering solutions tailored to your needs."

One specific habit worth breaking. SEO has spent a decade truncating title tags at around 60 characters because that is what Google displays. The index stores the full title. A title that reads as a complete, specific statement is now working harder than one trimmed to fit a SERP snippet preview. That does not mean writing 140-character titles for Google. It means the useful words should be in there, not cut for display reasons.

The second habit: stop burying the answer. Plenty of pages open with a paragraph of scene setting before the substance arrives in section three. In classic search Google would find the good passage anyway. Here, the 200 characters after your H1 are the whole audition. This is the same principle behind structuring content so LLMs cite you, applied to a much smaller window than most people realise.

What we would actually check on a site this week

Six things, in rough order of return.

  1. Read your own first 200 characters after each H1. Out loud. If it does not state what the page is and who it is for, rewrite it. Do this on your money pages first.
  2. Write titles that are complete statements. Specific, unhedged, with the thing you do and the place or audience you do it for. Stop trimming useful words purely for display. If you want variants to compare, our meta title and description generator writes them at SERP length.
  3. Check internal links to your important pages. If a page is only reachable via the sitemap or deep pagination, add real links from pages that already get crawled.
  4. Check page weight. Anything approaching 4 MB is at risk of being dropped entirely. This is usually images and an oversized JavaScript bundle.
  5. View source, not the rendered page. If your service detail, pricing or FAQ answers only appear after a client-side fetch, they are invisible at this layer. The cache stores what it receives, and it is not running your app for you.
  6. Audit anything you noindexed instead of deleting. Assume it is cached and readable. Decide whether you are comfortable with that.

None of that is new technique. It is technical SEO and clear writing, aimed at a narrower window than we are used to.

The caveat we would attach to all of it

This is one research team's log analysis of a sampled crawl, on a system that OpenAI changes without announcement, as Reddit found out when its ChatGPT citations collapsed in four days. The 30-minute figure, the 4 MB limit and the codename are observations, not published specification. They will move.

What we would not expect to move is the shape of it. A retrieval system that judges most candidates on a title and an opening snippet, reads a cached copy when it needs more, and rarely opens the live page is a reasonable design for anyone serving answers at that volume. The specific numbers are this month's. The incentive to be legible in a very small window is structural.

It also pairs badly with the other half of the picture. A large share of ChatGPT answers never run a web search at all and come straight from model memory, which we covered in two-thirds of ChatGPT answers never search the web. Between memory on one side and a 200-character snippet on the other, the space where a long, carefully argued page gets read in full is genuinely thin.

That is not an argument for writing worse pages. Depth is still what earns the open, and the open is what earns the citation. It is an argument for making sure the first thing anyone reads, human or model, does its job.

If you want a second pair of eyes on how your pages are structured for this, our AI search optimisation work starts with exactly these checks, and the free SEO review covers the technical side of it.

More on technical seo
John Kyprianou

John Kyprianou

Founder & SEO Strategist

John brings over a decade of experience in SEO and digital marketing. With expertise in technical SEO, content strategy, and data analytics, he helps businesses achieve sustainable growth through search.

Keep reading

More to think about.

All insights
Diagram of a small website's internal links showing an orphan page, a service page buried several clicks deep and a link pointing at a redirect, by SEO Turtle

Technical SEO

Internal linking audit: orphan pages, click depth and wasted links

Most internal linking advice is written for six-figure e-commerce sites and does nothing for a 200-page business site. The problems there are smaller and more specific: orphan pages, money pages buried five clicks deep, navigation anchors that all say the same thing, and links still pointing at old URLs. Here is how we find each one and what we fix first.

September 24, 2026
Diagram of an ecommerce filter sidebar producing hundreds of URL combinations, with a small set marked index and the rest marked block, by SEO Turtle

Technical SEO

Faceted navigation SEO: which filter pages to index and which to block

A filter sidebar can generate more URLs than your store has products, and the two standard answers (index everything, block everything) are both wrong. Here is how to decide facet by facet, why rel=canonical does not fix the crawl problem, and what to check in Search Console after you ship the change.

September 22, 2026
Old website URLs mapped to new URLs through 301 redirects during a site redesign

Technical SEO

Website redesign SEO: the checks that decide whether you keep your traffic

Most redesigns that lose traffic fail for one of a handful of reasons, and every one of them is checkable before launch. Here is the order we work through, why mass-redirecting to the homepage is worse than a 404, and what changes now that ChatGPT and Perplexity crawl on their own schedule.

September 17, 2026
Diagram of a server access log with lines sorted into Googlebot, AI training crawlers and user-triggered retrieval bots, by SEO Turtle

Technical SEO

Log file analysis for SEO: what your server actually saw

A crawl tool tells you what a bot could do. A server log tells you what it did. Most log file guides are written for sites with millions of URLs, then handed to people with four hundred pages and no SSH access. Here is the honest version: why crawl budget is probably not your problem, why the bot population is, and where the logs actually live on Cloudflare, Vercel, cPanel and Shopify.

September 15, 2026
Diagram of a Google Search Console page indexing report with URLs sorted into technical faults and quality judgements, by SEO Turtle

Technical SEO

Crawled, currently not indexed: how to work out what is actually wrong

Google has looked at the page and decided not to keep it. Most advice treats that as a technical fault and sends you off to hammer Request Indexing. Here is the order that actually finds the cause: an afternoon of cheap plumbing checks, a look at what Google chose as the canonical, and then the harder question of whether the page deserves to exist.

September 8, 2026