robots.txt for AI Bots: Block, Allow, or Something In Between?
If you want to control AI access to your site, robots.txt ai bots policy is the first lever most teams reach for, and the easiest to get wrong. A single wildcard can block search traffic. A vague rule can be ignored by the exact training bot you meant to stop. A missing test can leave paywalled archives open for months. This guide gives you precise, copy ready patterns for common AI agents plus a decision method so you choose block, allow, or selective limits with confidence.
This guide is for site owners, SEOs, and developers who manage crawling and indexing. You will learn how robots.txt actually works, which AI tokens you can control, three starter policies with exact syntax, path level controls, testing steps, and maintenance routines. You will also learn when robots.txt is not enough and what to pair it with for licensed or member only content. By the end, you will have a tested file, a rollout checklist, and a review cadence that keeps policy aligned with business goals.
Key takeaways
- robots.txt is a voluntary per user agent signal. Compliant AI crawlers honor it. Bad actors ignore it, so sensitive paths need access controls too.
- Write separate rules per AI token, such as GPTBot, ClaudeBot, and PerplexityBot, instead of one broad vendor block.
- Three practical defaults cover most sites. Allow all for open marketing content, block training for licensed archives, or selective limits by path.
- Test every change by fetching the file, validating syntax, and monitoring Search Console coverage for 7 days.
- Review quarterly. Bot names, docs, and crawl patterns change, and stale rules cause silent over blocking or under blocking.
- How robots.txt actually works
- Which AI bots you can control
- Three starter policies
- Writing precise rules for major AI agents
- Wildcards paths and sitemap hints
- Testing and validating your file
- Beyond robots.txt
- Common syntax errors
- Documenting your robots.txt ai bots policy
- Rolling out changes safely
- Review cadence and maintenance
- FAQ
- Sources
- Further reading
<!-- IMAGE-PROMPT cover: 1200x630, DependsIt brand, deep charcoal #121212 background, vibrant mint #22E3B0 accent glow, thin node-network line art, Clash Display style bold heading space on left, General Sans clean labels, subject: robots.txt document with User-agent Allow Disallow lines gating AI bot icons from website paths, flat vector, high contrast, accessible, no photorealistic faces, no text smaller than 24px, no em dash in rendered text, export PNG then cwebp -q 82 to WEBP -->
How robotstxt actually works
robots.txt is a plain text file served from the site root, for example yourdomain.com/robots.txt. Crawlers fetch it before broad crawling and look for the group that matches their user agent token. Each group contains Allow and Disallow path rules plus optional Sitemap hints and Crawl-delay style directives where supported. When no group matches, the wildcard group applies. When no file exists or the file allows a path, compliant crawlers proceed. When a path is disallowed for their token, compliant crawlers skip it, though they may still index a URL if links point to it without fetching content.
Matching is case sensitive for paths and tokens, and order matters in predictable ways. The crawler finds the most specific matching group for its token, then applies the longest matching Allow or Disallow rule for the requested path. A Disallow of slash blocks everything under root for that token. An Allow of slash plus Disallow of slash private blocks only the private section. Comments start with hash and are ignored. Blank lines separate groups. A missing colon, a stray space in the User-agent line, or a BOM character at file start can invalidate parsing, so keep the file simple and ASCII clean.
For AI control, three facts matter most. First, robots.txt controls fetching, not indexing by reference and not reuse of already collected data. If a page was crawled before you added a block, that copy may remain in existing datasets. Second, robots.txt is voluntary. Major AI providers state that documented crawlers honor it, which covers GPTBot, ClaudeBot, PerplexityBot, and similar named agents. Scrapers that fake tokens or ignore the file need firewall, CDN, and auth controls instead. Third, robots.txt cannot authenticate. It does not hide a URL, it only asks polite bots not to fetch it. Member only PDFs, client exports, and staging hosts need login, IP allow lists, or signed URLs regardless of robots content.
A minimal valid file looks like this. It allows search crawlers broadly while keeping an admin path closed. Later sections add AI specific groups with the same structure.
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /cart/
Sitemap: https://www.example.com/sitemap.xml
Keep the file under commonly supported size limits, serve it with a 200 status and text plain content type, and update it through version control rather than manual server edits. For a full audit method that protects search traffic while you edit, see is robots.txt blocking your indexing and how to audit it. That audit pairs well with every AI rule change in this guide.
Which AI bots you can control with robotstxt
You can control any crawler that identifies with a stable user agent token and states that it honors robots.txt. In practice, that includes the major documented AI agents. GPTBot for OpenAI training, OAI-SearchBot for OpenAI search, ChatGPT-User for on demand browsing, ClaudeBot for Anthropic training, PerplexityBot for Perplexity collection and cited retrieval, Google-Extended for AI training opt out without blocking Googlebot search, Bytespider for ByteDance collection, PetalBot for Huawei search and AI tasks, CCBot for Common Crawl archives, and Meta-ExternalAgent for Meta AI and preview tasks. Each token gets its own group so you can allow, limit, or block per purpose.
You cannot reliably control unnamed or rotating agents with robots.txt alone. Research scrapers, SEO tools that ignore the file, and malicious bots that spoof well known tokens will fetch regardless of rules. Logs reveal these cases. If a token shows no robots.txt fetch before broad crawling, ignores Crawl-delay style pacing, hammers sequential IDs, or fails reverse DNS validation, treat it as non compliant and handle it at the CDN or WAF layer. Robots edits alone will not stop it, and repeated robots tightening may accidentally catch good bots while the bad bot continues.
New tokens appear regularly. Providers add image, video, or enterprise variants, rename agents during product changes, and test limited crawls before updating docs. Rather than hardcoding a once and done list, maintain a dated inventory with columns for token, operator, purpose, first seen in your logs, last seen, request volume, and current policy. Check vendor docs quarterly and compare with your own top 50 user agents monthly. When a new token enters your top 30 or exceeds 500 requests per day, assign policy promptly instead of waiting for a headline.
For background on crawler directives and header behavior, see the MDN guide to robots handling. Pair that reference with your log data so rules reflect both documented behavior and observed behavior on your own site. Documented purpose tells you what the bot should do. Observed paths, rates, and response codes tell you what it actually does for you.
Three starter policies allow block or selective
Most sites fit one of three starting points. Choose the one that matches your content mix, then refine per path in later sections. The allow leaning default suits open marketing sites, docs, and blogs that benefit from broad discovery and answer citations. The block leaning default suits licensed archives, paid research, stock media, and member databases where bulk reuse harms the business. The selective default suits mixed sites with both public guides and restricted assets, which describes most publishers, marketplaces, and SaaS knowledge bases. Use these robots.txt ai rules as a baseline when you need to block ai crawlers for archives while keeping marketing pages open.
Policy one allows all documented AI bots for public content while keeping admin and utility paths closed. It is simple to maintain and maximizes eligibility for citations and discovery. Use it when your content is original but not licensed for resale, when attribution benefits outweigh reuse concerns, and when bandwidth is healthy. The file stays short, monitoring stays light, and future AI products can discover your pages without manual allow listing. The trade off is broad training reuse to the extent operators honor the allow.
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /account/
Disallow: /search/
Policy two blocks training and pipeline bots while allowing search and on demand answer fetchers for public paths. It reflects a common middle ground. Training collection is limited, live retrieval that can attribute and send visits remains available, and search crawling is untouched. Use it when you want visibility in answers but prefer to opt out of bulk model training where supported. Note that blocks affect future collection, not copies already collected, and that live citation behavior still depends on product retrieval at answer time.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Bytespider
Disallow: /
Policy three applies different rules by path. Public guides stay open. Licensed, member only, and high cost media paths stay closed to training and pipeline bots while remaining available to authenticated humans and, where appropriate, to search crawlers that drive subscriptions. This is the most work to maintain but the best fit for mixed businesses. It requires a clear path map, consistent URL structure, and discipline about placing new content in the correct bucket. Without that structure, selective rules drift into confusion within months.
Pick one starter, deploy it to staging first, validate, then promote to production with monitoring. Record the choice, the date, the owner, and the reason in your policy notes. That record prevents repeated debates each time a new bot makes news. It also makes quarterly reviews fast, because you compare new evidence against a stated baseline instead of starting over.
Writing precise rules for major AI agents
Precision beats brevity. Write one group per token with an explicit purpose comment, so future editors understand intent without guessing. Use exact token spelling as documented, because User-agent matching is substring based and case sensitive in practice across major parsers. Keep each group narrow. If you intend to block training but allow search, block GPTBot while explicitly allowing OAI-SearchBot and Googlebot in separate groups. Never rely on a single vendor level assumption when the vendor operates multiple agents with different roles.
An example selective file for a publisher with open guides and a licensed research archive shows the pattern. Public paths stay open to answer fetchers and search. Training and pipeline bots are closed for the archive but open for marketing guides. Comments state the reason per group, which helps legal and editorial review without reading server docs.
User-agent: GPTBot
Allow: /guides/
Allow: /blog/
Disallow: /research/
Disallow: /archive/
User-agent: ClaudeBot
Allow: /guides/
Allow: /blog/
Disallow: /research/
Disallow: /archive/
User-agent: PerplexityBot
Allow: /guides/
Allow: /blog/
Disallow: /research/
User-agent: Googlebot
Allow: /
User-agent: Google-Extended
Disallow: /research/
Disallow: /archive/
Note the Google pattern. Googlebot remains allowed for search across the site, while Google-Extended is disallowed only for sensitive paths. This preserves search traffic and AI Overviews eligibility built on search indexing, while signaling a training opt out for the archive where Google honors it. For background on how Google crawls and evaluates pages for search features, see the Google documentation on crawling and indexing. Keep Google rules minimal and test coverage after every edit, because wildcard mistakes here harm search directly.
For OpenAI, keep training and search separate. Disallow GPTBot where you want to limit training reuse, allow OAI-SearchBot where you want fresh discovery for OpenAI search, and allow ChatGPT-User where you want users to browse specific pages from chat. For example, you can allow ai training for public guides in the user-agent GPTBot group while you block ai training for research archives in the same file. For Anthropic, disallow ClaudeBot for paths you want out of training pipelines. For Perplexity, decide based on citation value versus archive sensitivity. For pipeline bots such as CCBot and Bytespider, apply the same path logic at broader scope, since one rule can affect many downstream reusers at once.
Wildcards paths and sitemap hints
Path rules use prefix matching with simple wildcards that vary slightly by parser, so keep patterns conservative. Disallow slash private blocks slash private, slash private dash reports, and slash private slash team, because all start with that prefix. To block only exact file types, use a dollar end anchor where supported, such as Disallow slash star pdf dollar to block PDF URLs. To allow a single file inside a blocked directory, place an Allow for that file above the broader Disallow and rely on longest match behavior. Test each pattern with real URLs from your logs rather than assuming.
Common path buckets for AI policy include slash admin, slash account, slash checkout, slash search, slash api, slash internal, slash staging, plus content specific buckets such as slash research, slash archive, slash members, slash downloads, and slash media originals. Map your own URL structure to these buckets before writing rules. If sensitive content lives under mixed paths with public content, restructure URLs first. Robots rules cannot reliably separate intermixed content without fragile patterns that break on the next redesign.
Sitemap hints in robots.txt help compliant crawlers discover canonical URLs efficiently. List only clean index sitemaps with canonical public URLs, accurate lastmod dates, and no more than supported URL counts per file. Do not list staging, parameter variant, or expired listing sitemaps. A focused sitemap reduces wasted AI crawler hits on duplicates and helps search crawlers spend budget on pages that matter. For sitemap hygiene that benefits every crawler, see crawl budget explained and how Google decides what to index. The same pruning that helps Googlebot also calms aggressive AI collectors.
Avoid exposing internal search result URLs, infinite calendar views, and faceted filter combinations to bulk crawlers. These paths multiply URL count without adding citable value. Block them for AI training and pipeline tokens, add canonical or noindex handling for search, and fix internal links that point into the trap. Path control plus discovery fixes beat repeated rate limit firefighting after each spike.
<!-- IMAGE-PROMPT diagram-01: 1600px max, DependsIt brand mint #22E3B0 on charcoal #121212, node-network line art, subject: robots.txt path control diagram showing public guides open and licensed archive blocked for AI crawlers, flat vector, accessible, no em dash, Clash Display style headings, General Sans clean labels -->
Testing and validating your file
Every robots change deserves the same validation sequence, no matter how small the edit looks. First, check syntax locally. Confirm each User-agent line has a colon, each group has at least one rule, comments use hash, the file is UTF-8 without BOM, and the size is within reasonable limits. Second, fetch the live file as a bot would. Request slash robots.txt from production, confirm a 200 status, text plain content type, and identical content across CDN edges and origins. Third, test representative URLs against each AI group. Verify that public guides are allowed, licensed paths are disallowed for training tokens, and search crawlers remain allowed for search paths.
Use Search Console robots.txt tester and URL Inspection for Google behavior, plus CDN logs for AI tokens. After deployment, watch for the robots.txt fetch from each major AI bot followed by the expected allow or block pattern on subsequent page requests. If a blocked bot continues to fetch disallowed paths at full rate, check token spelling, group placement, caching of an old file at the edge, and whether the requests come from spoofed agents that ignore the file. Edge caching of a stale robots.txt is a frequent cause of confusion. Purge CDN cache for the file on deploy and set a short TTL so updates propagate quickly.
Monitor for 7 days after each change. Track Search Console coverage for indexed versus excluded URLs, server error rates, sitemap indexation, and organic traffic for affected sections. Track AI bot request counts and bandwidth per token to confirm the intended reduction or continued access. Any unexpected coverage drop, traffic dip in allowed sections, or continued crawling of blocked paths triggers immediate review. Keep the previous file version in version control so rollback takes minutes, not hours.
Document the test. Save the file diff, the fetch headers, three example allowed URLs and three example blocked URLs per major token, and the 7 day before and after metrics. That packet makes audits fast and gives legal and leadership confidence that controls work as stated. For audit steps that catch accidental search blocks early, see wildcard rules in robots.txt that accidentally block indexing. That checklist pairs directly with AI rule testing.
Beyond robotstxt when you need more than a signal
robots.txt asks politely. It does not enforce. For licensed, paywalled, or personal content, pair it with real controls. Require authentication for member paths, serve licensed PDFs through signed URLs with expiry, add noindex to pages that should never appear in search, and use firewall or CDN bot management for abuse. For AI Overviews style citation control where search visibility matters, remember that noindex removes search eligibility as well as AI citation eligibility. Use it only when exclusion from search is acceptable.
Meta robots tags and X-Robots-Tag headers provide page level noindex, nofollow, nosnippet, and noarchive style controls that complement robots.txt path rules. Use them for individual pages that need indexing restraint without blocking the fetch itself. For example, allow a bot to fetch a page for link discovery but add noindex to keep it out of search results. Keep tags consistent between HTML and headers, and avoid sending conflicting signals such as Disallow in robots plus noindex on the page for the same crawler, since the Disallow prevents the crawler from seeing the noindex. Choose one mechanism per goal and test.
Paywalls and flexible sampling need structured signals. Use JSON-LD paywall markup, clear subscription messaging, and consistent allow rules for search crawlers that support paywalled content discovery, while restricting training and pipeline bots for the same paths where policy requires. For media, control originals separately from previews. Allow compressed previews for discovery while keeping high resolution originals behind auth or signed URLs. This layered approach protects revenue without removing all visibility.
For background on crawler versus browser handling of directives, the MDN reference explains headers and meta behavior in vendor neutral terms. Combine that with CDN features such as token authentication, rate limiting by user agent and path, and validated bot allow lists based on published IP ranges. Layers beat single file reliance. robots.txt states intent for compliant bots. Auth, headers, and edge controls enforce boundaries for everyone else.
Common syntax errors that break everything
The most damaging error is an overbroad wildcard. A single Disallow slash under User-agent star blocks all compliant crawling, including search, and can deindex sections within weeks. A Disallow slash star pdf dollar placed under the wrong group can block product manuals that drive conversions. A missing slash, such as Disallow private without the leading slash, may not match any URL on strict parsers while appearing correct to humans. Always use leading slashes, test with real URLs, and keep wildcard use minimal and explicit.
The second group of errors involves group placement. Rules placed before any User-agent line are ignored. Two groups for the same token in different parts of the file can interact in surprising ways on some parsers. Comments without hash, colons replaced by similar Unicode characters after copy paste from docs, and invisible BOM characters at file start also break parsing silently. Edit in plain text, review diffs character by character, and generate the file from a template or script rather than hand editing on the server.
The third group involves caching and serving. Serving robots.txt with a 404, a soft 404 HTML page, a 500 during deploys, or different content per edge node causes inconsistent behavior across bots. Some crawlers treat prolonged 500s as a signal to slow or stop crawling. Others cache a 404 as allow all longer than expected. Serve a stable 200, keep deploys atomic, purge CDN cache for the file on change, and monitor status codes for slash robots.txt as a synthetic check alongside uptime.
The fourth group involves conflicting signals. Disallowing a path in robots while also requesting noindex for URLs under that path prevents compliant crawlers from fetching the noindex and can leave stale indexed copies visible longer. Canonical tags on disallowed URLs are similarly unseen. Decide per goal. Use robots Disallow to save crawl budget and limit fetching. Use noindex and canonicals for indexed URL management where fetching remains allowed. For audit patterns that reveal these conflicts, see how to check if a page is indexed beyond site search. Fixing conflicts often restores both search health and intended AI limits.
Documenting your robots.txt ai bots policy
A one page policy prevents repeated debates and accidental edits. Start with a path map that lists each URL bucket, example URLs, owner, and sensitivity. Public guides, blog, docs, and marketing pages usually form the open bucket. Research, archives, member areas, client data, and high resolution media form restricted buckets with different AI rules. Keep the map short and tied to URL prefixes so engineering can implement it without interpretation.
Next, add a bot matrix. Rows are tokens such as GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Bytespider, PetalBot, CCBot, and Meta-ExternalAgent. Columns are path buckets plus the chosen action of allow, disallow, or limit with rate notes. Include the date, the owner, and the reason in plain language, such as training opt out for licensed research or allow retrieval for public guides to support citations. Include an ai crawling control note per row so reviewers see why each token is allowed or limited. Store the matrix with version history so changes are traceable.
Include the generated robots.txt content or a link to the source template, plus validation evidence. Save the fetch headers, the tester screenshots or logs, and three allowed plus three blocked example URLs per major token. Note CDN rate limits, WAF rules, and auth controls that complement the file, because robots.txt alone rarely tells the full story for restricted content. When leadership asks what happens if a new bot appears, point to the matrix process rather than promising a static file that never changes.
Share the policy with editorial, legal, engineering, and SEO. Editorial needs to know where to place new content so it inherits the correct bucket. Legal needs to confirm licensing posture. Engineering needs to own generation, deployment, and edge caching. SEO needs to monitor search impact and citation trends. A 30 minute quarterly review with these roles keeps the policy current without heavy process.
Rolling out changes without breaking SEO
Roll out in stages. Start in staging with identical URL structure and a staging robots file that mirrors production intent without affecting live crawlers. Validate syntax, test allowed and blocked URLs per token, and confirm sitemap references point to staging sitemaps, not production. Promote to production during low traffic hours, purge CDN cache for slash robots.txt, and fetch the live file from multiple regions to confirm consistency. Announce the deploy internally with the diff and the reason so support and marketing are not surprised by log or traffic shifts.
Monitor search health closely for 7 to 14 days. Watch Search Console coverage for unexpected excluded spikes, check indexed counts for allowed sections, and compare organic traffic week over week with seasonality in mind. Watch server logs for robots.txt fetch volume per AI token followed by the expected page fetch pattern. A correct block shows robots fetch plus sharply reduced page fetches for that token on disallowed paths. Continued full rate fetching after a correct file suggests spoofing or edge caching of an old file, not a syntax error. Investigate IPs and reverse DNS before tightening further.
Keep rollback ready. Store the previous production file, the deploy timestamp, and the purge command in the release notes. If coverage drops, if allowed sections lose traffic, or if a partner integration that relies on preview fetching breaks, revert first and diagnose second. Most incidents come from overly broad patterns, group misplacement, or CDN inconsistency rather than AI specific logic. Narrow rules and fast rollback limit blast radius.
Coordinate with content launches. When you publish a guide that should earn citations, confirm it lives under an allowed path, appears in the sitemap with current lastmod, receives internal links from hubs, and returns 200 to both search and permitted AI fetchers. Request a fresh crawl through URL Inspection for important URLs and sample AI citation status over the next four weeks. A clean rollout plus prompt discovery gives new content the best chance in both classic results and AI answers.
<!-- IMAGE-PROMPT workflow-02: 1600px max, DependsIt brand, deep charcoal #121212 background, vibrant mint #22E3B0 accent glow, thin node-network line art, subject: safe rollout workflow from staging test to production deploy to coverage monitoring to rollback plan, flat vector, accessible, no em dash, Clash Display style headings, General Sans clean labels -->
Review cadence and maintenance checklist
Monthly tasks take under an hour once the system is set. Export the top 50 user agents by request count, note any new AI tokens in the top 30, and compare per token bandwidth with the prior month. Confirm that slash robots.txt serves 200 with consistent content across edges. Scan Search Console coverage for new excluded patterns that suggest over blocking. Log the snapshot with dates so trends are visible without digging through raw logs each time.
Quarterly tasks go deeper. Re read vendor docs for renamed or new tokens, update the bot matrix with observed counts and policy decisions, re validate reverse DNS for high volume agents, and test three allowed plus three blocked URLs per major token. Review the path map for new sections, redesigns, or paywall changes that require bucket updates. Regenerate robots.txt from the template if you use scripted generation, then run the full validation sequence before deploy. This is also the right time to review CDN rate limits and WAF rules for impersonation and abuse.
After migrations, template changes, or auth updates, test immediately. Confirm that staging hosts remain closed to all compliant bots, that production canonicals are consistent, and that licensed paths still enforce auth plus robots limits. Check that sitemaps list only canonical public URLs and that lastmod reflects real changes. A short post deploy checklist prevents long silent failures where a new template accidentally exposes member PDFs or blocks product guides.
Keep everything in one folder with dates. Store the path map, the bot matrix, the generated file versions, validation evidence, and monthly snapshots together. When a new AI bot makes headlines, you will answer with your own data in minutes. That steady routine matters more than any single rule. Bots change names and behaviors, but a maintained policy keeps control with you rather than with the latest announcement.
FAQ
Should I block all AI bots in robots.txt?
Only if all your content is licensed, private, or otherwise unsuitable for any AI reuse and you accept reduced discovery in AI answers. Most mixed sites do better with selective rules that keep public guides open while restricting training and pipeline bots for sensitive paths. Most mixed sites do better when they block ai crawlers for sensitive archives but still allow ai training for open guides where attribution helps.
Will blocking GPTBot remove me from ChatGPT answers?
Blocking GPTBot limits training collection. It does not always remove live citations that come from on demand retrieval or search grounded features. If you want to limit answer visibility for a specific engine, address its retrieval agents and access controls for those paths as well.
Does Google-Extended affect Google Search rankings?
Google states that Google-Extended controls certain AI training uses without blocking Googlebot search crawling. Disallowing Google-Extended alone should not remove you from search or from search based features. Always monitor coverage after changes to confirm expected behavior on your site.
How do I test robots.txt for a specific AI bot?
Fetch slash robots.txt and confirm 200 plus consistent content. Then test representative URLs against that token's group using a tester or manual longest match review. Check the user-agent GPTBot group first, then confirm your other robots.txt ai rules still allow search crawlers. Check logs for the bot's robots fetch followed by the expected allow or block pattern on page requests.
What if a bot ignores my robots.txt?
First verify token spelling, group placement, and CDN caching of an old file. Check reverse DNS to spot spoofing. For truly non compliant actors, use CDN rate limits, WAF rules, IP controls, or authentication for sensitive paths, since voluntary signals alone will not stop them.
How often should I update AI robots rules?
Review logs monthly and vendor docs quarterly. Update immediately after migrations, paywall changes, or coverage anomalies. Keep a dated matrix so each change ties to evidence and ownership rather than reacting to headlines. Review ai crawling control notes each time so block ai training decisions stay tied to current contracts and observed crawl volume.
Sources
- https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Robots
- https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag