Notes from the work

Apple quietly tripled Applebot's crawl capacity weeks before Siri AI ships

Apple added 4,656 IP addresses to Applebot with no announcement, right before Siri AI ships. What it signals, and what to check in your robots.txt now.

John KyprianouJohn Kyprianou8 min readUpdated

Apple does not announce crawler changes. It edits a JSON file and lets you find out.

Late in July, the file listing Applebot's published IP ranges grew by 21 new CIDR prefixes: eighteen /24 blocks and three /28 blocks, all inside 17.166.0.0/16. That is 4,656 new addresses. The published pool went from 2,400 to 7,056, an increase of roughly 194% in a single edit, detected by monitoring on 17 August (PPC Land, Search Engine Roundtable).

No blog post. No developer note. No explanation of any kind.

Diagram showing Applebot's published IP pool growing from 2,400 to 7,056 addresses ahead of the Siri AI launch, by SEO Turtle

Why the timing matters

Apple confirmed the rebuilt Siri at WWDC in June. It runs on the next generation of Apple Foundation Models, it gets its own app, and it can pull current information from the web to answer open ended questions (Apple Newsroom, TechCrunch). iOS 27 ships this autumn as a free update.

You do not triple the addresses you crawl from because you plan to crawl the same amount from more places. Infrastructure gets provisioned ahead of load.

We are reading this as capacity being staged for a product launch, and we will say plainly that this is inference rather than fact. Apple has said nothing, and a large IP allocation can also mean a routing change or a data centre migration. But the size of it, the concentration in one /16, and the four weeks between the file edit and the iOS 27 release window all point the same direction.

The mistake we see most: blocking the wrong user agent

Here is the part worth acting on regardless of what Apple's plans turn out to be.

Apple runs two user agents and they do completely different jobs. Most sites treat them as one thing.

Applebot is the crawler. It fetches your pages and feeds Spotlight, Siri and Safari search. It is the reason your site can appear in an Apple search surface at all.

Applebot-Extended does not crawl anything. It is a control token. Its only purpose is to tell Apple whether the content it already fetched can be used to train foundation models (Apple).

So these two rules are not variations on a theme:

User-agent: Applebot-Extended
Disallow: /

That says: index me for search, do not train on me. It is the publisher opt out.

User-agent: Applebot
Disallow: /

That says: remove me from Apple search entirely. Different decision, much bigger consequence, and we run into sites that made it by accident.

The accident usually happens one of two ways. Someone read a 2024 thread about blocking AI crawlers and pasted a list of user agents into robots.txt without checking what each one did. Building the file deliberately, by hand or with a robots.txt generator that writes per-bot rules, avoids that class of mistake entirely. Or the site sits behind a managed bot rule that treats every non-Google crawler as scraping. We wrote about the second pattern in Cloudflare's AI crawler blocking, and Applebot is one of the agents that quietly gets caught in it.

The Googlebot fallback almost nobody knows about

This one surprises people, and it is in Apple's own documentation.

If your robots.txt has no rules addressing Applebot specifically, Applebot follows your Googlebot rules instead.

Read that again with your own robots.txt in mind. Every restriction you ever wrote for Googlebot, including the crawl trap exclusions and the faceted navigation blocks and the staging path you disallowed in 2021, is currently governing what Apple can see. If you have a tight Googlebot block anywhere, you inherited it for Siri without deciding to.

That is defensible when your Googlebot rules are current and deliberate. It is a problem when they are archaeology. Most robots.txt files we audit are archaeology.

Your IP allowlist just went stale

If you verify crawlers by matching against a hardcoded list of Apple IP ranges in your WAF or your log pipeline, 4,656 addresses now fail that check.

Depending on how the rule is written, real Applebot traffic from those addresses gets rate limited, served a challenge, or logged as a spoofed bot. None of those outcomes are visible unless someone goes looking.

Apple documents two verification methods, and only one of them survives an event like this. Reverse DNS resolves genuine Applebot traffic to the *.applebot.apple.com domain, and a forward confirm on that hostname is the check that keeps working when the address pool changes. Matching against the published CIDR list only works if something re-fetches that file on a schedule.

Our position on this is not Apple specific. Verify crawlers by reverse DNS, not by a list you copied into a config file. Every major operator expands ranges without warning, and a static allowlist is a slow leak you will not notice until traffic is already gone. The same logic applies to the AI crawlers we covered in the robots.txt guide.

Our take: Apple is the surface nobody is optimising for, and that has been rational until now

We will frame this clearly as opinion.

Almost no business has ever done a single piece of work aimed at Apple's search surfaces, and until this year that was the correct allocation of effort. Siri could not answer an open web question. Spotlight was a launcher. There was nothing to optimise for, so nobody optimised.

What changes this autumn is that the assistant on roughly six in ten phones in the US and almost four in ten in Cyprus (StatCounter, August 2026, Cyprus figures) becomes something that reads the web and answers from it. Combine that with Assistant being retired from Android on 4 September and both default phone assistants become language models within the same few weeks.

The honest version of the opportunity is smaller than the headline. Nobody is going to build an Apple specific content strategy, and they should not. Apple crawls the open web with an ordinary crawler and reads ordinary pages. The work that makes you retrievable by ChatGPT and AI Mode is the same work that makes you retrievable by Siri.

The difference is the failure mode. With Google you have Search Console telling you when something breaks. With Apple you have nothing. There is no report, no coverage view, no notification. If Applebot cannot reach you, the only place that fact exists is in your server logs, and it will sit there for a year unless someone reads them.

That asymmetry is the whole argument for spending twenty minutes on this. The upside is ordinary. The downside is invisible.

What to check this week

Five things, none of which take long.

Grep your logs for Applebot. Confirm it is actually reaching you and getting 200s. If you have never seen it in your logs, that is the finding.

Read your robots.txt properly. Find every rule that touches Applebot or Applebot-Extended, and if there are none, go read your Googlebot rules instead, because those are the ones in force. A robots.txt rule tester will show you which line a given URL actually matches rather than which line you think it matches.

Decide the training question deliberately. Blocking Applebot-Extended is a legitimate choice and it costs you nothing in search visibility. Just make it on purpose rather than inheriting it from a pasted list.

Check your WAF and CDN rules. Managed bot protection is where good crawlers go to die quietly. Look at what your challenge and rate limit rules do to a request from an Apple range. If you are on Cloudflare, note that its 1 July announcement names Applebot alongside Googlebot and BingBot as a multi-purpose crawler that gets blocked wherever a customer has chosen to block Training, and that Training and Agent bots are blocked by default on ad-carrying pages for domains joining the network from 15 September (Cloudflare).

Switch crawler verification to reverse DNS. If anything in your stack allowlists Apple by hardcoded IP, that list is now missing 4,656 addresses.

None of this is glamorous, and that is roughly the point. The sites that show up in new answer surfaces are usually not the ones with a clever strategy for them. They are the ones that were reachable when the crawler arrived.

If you want a second pair of eyes on how your site handles AI and search crawlers, that is standard scope in our technical SEO and AI search optimisation work, and it is one of the first things we look at in a free SEO review.

More on technical seo
John Kyprianou

John Kyprianou

Founder & SEO Strategist

John brings over a decade of experience in SEO and digital marketing. With expertise in technical SEO, content strategy, and data analytics, he helps businesses achieve sustainable growth through search.

Keep reading

More to think about.

All insights
Diagram of a website sending Google contradictory canonical signals from internal links, sitemap and redirects, with Google selecting a different URL, by SEO Turtle

Technical SEO

Why Google ignores your canonical tag, and how to fix each case

The ticket usually says the canonical is correct and Google is still indexing the wrong URL. Most of the time the tag is fine and the rest of the site is arguing with it. Here is how to find out which URL Google actually picked, what each Search Console status really means, and which contradictions to go and remove.

October 7, 2026
Diagram of a small website's internal links showing an orphan page, a service page buried several clicks deep and a link pointing at a redirect, by SEO Turtle

Technical SEO

Internal linking audit: orphan pages, click depth and wasted links

Most internal linking advice is written for six-figure e-commerce sites and does nothing for a 200-page business site. The problems there are smaller and more specific: orphan pages, money pages buried five clicks deep, navigation anchors that all say the same thing, and links still pointing at old URLs. Here is how we find each one and what we fix first.

September 24, 2026
Diagram of an ecommerce filter sidebar producing hundreds of URL combinations, with a small set marked index and the rest marked block, by SEO Turtle

Technical SEO

Faceted navigation SEO: which filter pages to index and which to block

A filter sidebar can generate more URLs than your store has products, and the two standard answers (index everything, block everything) are both wrong. Here is how to decide facet by facet, why rel=canonical does not fix the crawl problem, and what to check in Search Console after you ship the change.

September 22, 2026
Old website URLs mapped to new URLs through 301 redirects during a site redesign

Technical SEO

Website redesign SEO: the checks that decide whether you keep your traffic

Most redesigns that lose traffic fail for one of a handful of reasons, and every one of them is checkable before launch. Here is the order we work through, why mass-redirecting to the homepage is worse than a 404, and what changes now that ChatGPT and Perplexity crawl on their own schedule.

September 17, 2026
Diagram of a server access log with lines sorted into Googlebot, AI training crawlers and user-triggered retrieval bots, by SEO Turtle

Technical SEO

Log file analysis for SEO: what your server actually saw

A crawl tool tells you what a bot could do. A server log tells you what it did. Most log file guides are written for sites with millions of URLs, then handed to people with four hundred pages and no SSH access. Here is the honest version: why crawl budget is probably not your problem, why the bot population is, and where the logs actually live on Cloudflare, Vercel, cPanel and Shopify.

September 15, 2026