Skip to main content

AI Crawlers and Robots.txt: What to Block, What to Allow

Summarize with ChatGPT
JK
John Kyprianou
February 13, 2026
4 min read
AI crawlers and robots.txt policy setup

AI crawlers are now a normal part of technical SEO. Alongside classic search bots (Googlebot, Bingbot), sites are seeing more traffic from model-related agents such as ClaudeBot, GPTBot, OAI-SearchBot, and PerplexityBot.

The challenge is simple: do you want these systems to access your content, or not?

This guide explains how to decide, and how to implement a clean policy in robots.txt without accidentally harming your core search traffic.

What AI Crawlers Actually Do

Not all AI-related user agents do the same thing. In practice, they usually fall into three buckets:

  1. Crawlers that fetch pages for indexing or training-related pipelines.
  2. Crawlers that support live retrieval for AI answers and citations.
  3. Extended-control directives that signal usage preferences for AI features.

For most websites, the operational question is still binary: allow or disallow.

Should You Block AI Crawlers?

There is no universal answer. Use business goals:

Usually block if:

  • You publish proprietary content that you do not want reused in AI-generated outputs.
  • You run a subscription or paywalled content business and want tighter control.
  • You are seeing heavy crawl load with little referral value from AI platforms.

Usually allow if:

  • You want visibility and citations in AI assistants.
  • You publish educational content, comparison pages, or thought leadership.
  • You are actively investing in AI search optimization.

For many brands, allowing selective access is now part of demand generation. If your content cannot be fetched, it cannot be cited.

Safe Robots.txt Pattern

If your policy is to block common AI crawlers, add explicit blocks like:

User-agent: ClaudeBot
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

Then keep your normal baseline section for everything else:

User-agent: *
Disallow:

This structure is clear, auditable, and easy to update as policies change.

Common Mistakes to Avoid

  1. Blocking User-agent: * by accident.
    If you set Disallow: / under the wildcard group, you can remove your entire site from standard crawling.

  2. Mixing contradictory rules without intent.
    Set your specific AI-agent rules explicitly, then define a clear wildcard baseline.

  3. Forgetting crawl monitoring after policy changes.
    Use logs and crawl stats to confirm your directives are being respected.

  4. Treating robots.txt as security.
    robots.txt is a crawl directive, not an access control mechanism. Sensitive data should be protected at the server/application layer.

A Practical Policy Framework

Use a simple review framework every quarter:

  1. Business value: Are AI platforms sending qualified traffic or branded searches?
  2. Content risk: Is your content highly sensitive or commercially unique?
  3. Infrastructure cost: Is crawler load causing measurable performance issues?
  4. Brand strategy: Do you want broader AI mention visibility this quarter?

If value > risk, allow selectively.
If risk > value, block aggressively and revisit later.

How We Handle It in the Robots.txt Generator

In our robots.txt generator, you can now use the Block AI crawlers toggle to automatically add disallow rules for common AI agents while keeping your default crawler policy intact.

That gives non-technical teams a safer way to apply policy without manually editing syntax.

Final Recommendation

Treat AI crawler policy like any other SEO control: test, measure, iterate.

Do not choose a permanent stance once and forget it. AI referral patterns, crawler behavior, and model ecosystems are still changing quickly in 2026.

If you want help deciding whether to block or allow specific AI agents for your site, request a free SEO review and we can map policy to your goals.

John Kyprianou

John Kyprianou

Founder & SEO Strategist

John brings over a decade of experience in SEO and digital marketing. With expertise in technical SEO, content strategy, and data analytics, he helps businesses achieve sustainable growth through search.

Related Articles

Bar chart comparing AI traffic conversion uplift claims: 23x from Ahrefs, 4.4x from Semrush, 1.54x from Adobe
AI Search

AI Traffic Converts 23x Better Than Google. That Number Is Wrecking Budgets

One number has done more to move marketing budgets in the last year than any Google update: AI search traffic converts 23x better than organic. It came from a single site over 30 days, and the company that published it said so at the time. Bigger datasets tell a very different story.

August 13, 2026
Diagram contrasting a wide shallow spread of AI citations with a narrow deep column of AI brand recommendations, by SEO Turtle
AI Search

AI Will Cite You Anywhere. It Only Recommends You Where You Go Deep

Your AI visibility tool says you appear in thirty categories. Your sales pipeline says two of them are real. New research from Search Engine Land explains the gap: AI engines will cite you across almost any topic, but they only recommend you inside the category you actually own. That changes what a content plan should look like.

August 12, 2026
Google routing search queries to lighter, cheaper AI models, explained by SEO Turtle
AI Search

Google Is Reading Your Site With a Cheaper AI Model. Here Is What That Changes

Everyone assumed AI search would get smarter at reading nuanced content. The models Google actually puts in front of your pages are getting lighter, because at Google's query volume cost wins. That flips the content advice: stop writing for a reader that infers.

August 11, 2026
Diagram of a user starring a source in Google settings and that choice feeding into an AI answer as a ranking signal, by SEO Turtle
AI Search

Google Is Turning User Choice Into a Ranking Signal

Preferred Sources moved into AI Overviews and AI Mode this summer, and Google has said it wants to use those picks as a ranking signal. That turns a settings toggle into search infrastructure. Our read on what it rewards, why the adoption numbers matter more than the feature, and what a business should actually take from it.

August 10, 2026
Split diagram showing ChatGPT pulling English pages while Google AI Overviews pulls local language pages, by SEO Turtle
AI Search

ChatGPT Prefers English. Google's AI Prefers Your Customer's Language.

Two studies published this summer found the same split from opposite directions: ChatGPT over-retrieves English pages, and Google AI Overviews almost never leaves the language of the question. If you run a Greek and English site, or English and Spanish, you have two different blind spots and one bad reflex to avoid.

August 7, 2026

Continue Your SEO Journey

Explore more expert insights and take action on your SEO strategy