The Complete List of AI Crawlers (And What They Do)
If you review server logs today, you will likely see new visitors with names like GPTBot, ClaudeBot, PerplexityBot, and Bytespider alongside familiar Googlebot and Bingbot hits. The ai crawlers list grows every quarter as model providers, answer engines, and data vendors expand collection. Each bot has a different owner, a different purpose, and a different impact on bandwidth, privacy, and future AI answers. This guide gives you a complete, plain language map of the bots that matter, what each one actually does, and how to set a sensible policy.
This guide is for site owners, SEOs, developers, and content teams who want facts without hype. You will learn the difference between training crawlers and live answer fetchers, the exact user agents for OpenAI, Anthropic, Google, Perplexity, Meta, Cohere, ByteDance, and others, how to spot them in logs, what they cost in crawl budget and performance, and how to decide your access policy before you edit robots.txt. By the end, you will have a maintained bot inventory and a monitoring routine you can run monthly.
Key takeaways
- AI crawlers fall into three groups. Training crawlers collect data for future models, search fetchers retrieve pages for live answers, and data pipeline bots feed third party datasets.
- OpenAI, Anthropic, Google, Perplexity, Meta, Cohere, ByteDance, and Common Crawl operate the most frequently seen AI user agents.
- Each bot publishes a user agent string and usually respects robots.txt, but behavior, frequency, and documentation quality vary by provider.
- Bandwidth and log noise add up. Unmanaged AI crawlers can increase origin hits, CDN costs, and crawl budget pressure on large sites.
- Start with inventory and measurement, then set policy. Block, allow, or limit selectively based on value, not headlines.
- What counts as an AI crawler
- Training crawlers versus live answer fetchers
- OpenAI crawlers explained
- Anthropic and Claude crawlers
- Google AI related crawlers
- Perplexity Cohere Meta and Mistral
- ByteDance Huawei Common Crawl and data pipelines
- How to identify AI crawlers in logs
- Bandwidth cost and performance impact
- Deciding your crawl policy
- Monitoring and maintaining your ai crawlers list
- FAQ
- Sources
- Further reading
<!-- IMAGE-PROMPT cover: 1200x630, DependsIt brand, deep charcoal #121212 background, vibrant mint #22E3B0 accent glow, thin node-network line art, Clash Display style bold heading space on left, General Sans clean labels, subject: network map of AI crawler bots with names visiting a central website node, flat vector, high contrast, accessible, no photorealistic faces, no text smaller than 24px, no em dash in rendered text, export PNG then cwebp -q 82 to WEBP -->
What counts as an AI crawler
An AI crawler is any automated agent that fetches web pages to support model training, retrieval augmented answers, or dataset creation. The label covers a wide range. Some bots crawl broadly across the web on a schedule, much like classic search crawlers. Others fetch a single page only after a user asks a question in a chat product. Others run inside research pipelines that repackage content for resale. Grouping them together without distinction leads to poor decisions, such as blocking a bot that sends referral traffic while allowing one that only trains competitors.
A practical definition uses three tests. First, the operator states an AI related purpose, such as model training, AI search grounding, or dataset building. Second, the agent identifies itself with a distinct user agent string or IP range that can be logged. Third, the fetch behavior serves an AI product rather than classic web search indexing alone. Googlebot still indexes for search even when its data also informs AI features, so most teams do not class core Googlebot as an AI crawler. Google-Extended is different, because its stated purpose is to control AI training use. That distinction matters when you write rules.
You will also see hybrid agents. For example, a search bot may fetch for both an index and a live answer feature. A chat product may use one user agent for batch training and another for on demand retrieval when a user pastes your URL into a prompt. Treating each user agent separately gives you finer control than blocking an entire vendor at once. Logs help here. Before you set policy, collect two to four weeks of user agent data so you know which bots actually visit your site, how often, and which paths they request.
Finally, remember that new names appear regularly. Providers rename bots, add variants for images or video, and test limited crawls without broad notice. Maintain a table with columns for bot name, operator, user agent token, stated purpose, observed frequency on your site, and current policy. That table is the backbone of this guide. Every later section adds rows to it, and the final section shows how to keep it current without weekly research. If you search for crawler list ai roundups, you will see overlapping names, so keep your own ai bot list tied to observed logs rather than headlines.
Training crawlers versus live answer fetchers
Training crawlers collect pages in bulk to build or improve language models. They run on a schedule, request many URLs across many sites, and do not send immediate traffic in return. Examples include GPTBot for OpenAI training, ClaudeBot for Anthropic training, and CCBot for Common Crawl datasets that many labs reuse. Allowing them means your content may influence future model behavior. Blocking them means your content is excluded from those training runs to the extent the operator honors your rules. Neither choice produces an instant traffic change, which is why this decision is about long term rights and positioning rather than short term clicks.
Live answer fetchers work differently. They retrieve pages on demand to ground a specific answer. Examples include ChatGPT-User, which fetches a page when a user asks ChatGPT to browse it, and PerplexityBot plus related Perplexity user agents that retrieve sources for cited answers. These bots can send referral visits when users click the citation. They also create the familiar trade off of AI search. Your content informs the answer, the user may or may not click through, and the citation still builds visibility. Blocking live fetchers removes you from those answers but also removes any chance of referral clicks from that engine.
Data pipeline bots form a third group. Common Crawl is the best known example. It publishes open web archives that startups, researchers, and model teams download instead of crawling themselves. Bytespider and similar collectors also feed large content pipelines for search and AI products. These bots rarely show up in product citation lists, yet their data can flow into many downstream models. If you care about broad reuse, pipeline bots deserve explicit policy, not just the well known brand names.
The practical takeaway is to set policy per purpose, not per headline. Many sites allow on demand answer fetchers that attribute and send visits, limit high volume training crawlers during peak hours or to selected paths, and decide on pipeline bots based on licensing posture. Document the reason for each choice so future debates focus on evidence. A simple policy matrix with allow, limit, and block columns per bot purpose keeps engineering, legal, and editorial aligned. Most llm crawlers fall clearly into one of these three purposes, and any public ai spider list you consult should label each entry the same way before you copy its rules.
OpenAI crawlers explained
OpenAI operates several distinct agents, and confusing them is a common error. GPTBot is the training crawler. Its user agent contains GPTBot and it collects pages for future model training. OpenAI states that GPTBot respects robots.txt and that site owners can disallow it without affecting other OpenAI fetch behavior. Blocking GPTBot is therefore a training data decision, not a ChatGPT visibility switch on its own. Many publishers block GPTBot while still allowing on demand browsing, which preserves the option to appear in cited answers that users request.
OAI-SearchBot is the search oriented crawler that supports search features inside OpenAI products. Its user agent contains OAI-SearchBot. It behaves more like a classic search crawler with discovery and refresh patterns. If you block it, your pages are less likely to appear as fresh sources in OpenAI search experiences. If you allow it, expect regular crawling of important pages plus linked discovery. Most content sites that want visibility in OpenAI search allow OAI-SearchBot while deciding on GPTBot separately based on training preferences.
ChatGPT-User is the on demand fetcher. Its user agent contains ChatGPT-User and it retrieves a specific URL when a user asks ChatGPT to look at that page. Volume is low and tied to user actions, not bulk schedules. Blocking it has little bandwidth benefit and removes the direct browse path users explicitly request. Most sites allow ChatGPT-User unless they have a strict no AI access rule for paywalled or licensed content. In that case, access controls and authentication matter more than robots rules alone, because robots.txt is a voluntary signal, not enforcement.
To configure OpenAI correctly, write separate rules per token rather than a single block for the vendor. Test with log sampling after each change. Look for GPTBot, OAI-SearchBot, and ChatGPT-User strings in the user agent field, confirm the request path and response code, and verify that classic search crawlers are unaffected. Keep notes on dates, because OpenAI occasionally updates documentation and crawl patterns. A per agent approach avoids the common mistake of blocking training and search with one broad rule when you only intended to block training.
Anthropic and Claude crawlers
Anthropic operates ClaudeBot as its primary crawler for training and data collection supporting Claude models. The user agent contains ClaudeBot, and Anthropic publishes documentation describing its purpose and control options. Like GPTBot, ClaudeBot respects robots.txt when configured correctly. Site owners who want to exclude their content from Anthropic training pipelines disallow ClaudeBot explicitly. Those who want to remain available for AI related discovery allow it and monitor volume.
Anthropic also uses Claude-User style fetching for on demand tasks in some contexts, where the agent retrieves a page at a user's request. Volume for on demand fetch is much lower than bulk training crawl. If your policy distinguishes training from attribution bearing retrieval, reflect that distinction in separate rules and notes. The same principle that applies to OpenAI applies here. Do not assume one rule covers all Anthropic activity. Check current documentation and your own logs for the exact tokens you see.
Observed behavior for ClaudeBot varies by site size and link profile. Large content sites with broad internal linking often see steady crawling across sections, while smaller sites see infrequent bursts. If you notice spikes, check whether the bot follows sitemap URLs, paginated archives, or faceted navigation that multiplies URL variants. Limiting crawl waste with clean sitemaps, canonical tags, and parameter controls helps every crawler, not just AI bots. For a method to audit waste, see how to check if a page is indexed beyond site search. The same hygiene that improves index accuracy also reduces unnecessary AI crawler load.
For teams concerned about attribution, note that allowing ClaudeBot does not guarantee citations in Claude answers. Citation behavior depends on the product, the query, and retrieval at answer time. Treat training access and answer visibility as related but separate outcomes. Your robots rules influence collection, while content clarity, structure, and topical authority influence whether you are cited when answers are generated. Both matter, but they operate on different timelines.
Google AI related crawlers
Google's setup confuses many teams because classic search crawling and AI training controls use different signals. Googlebot and Googlebot-Image remain the core crawlers for search indexing. Blocking them removes you from Google Search, which most sites do not want. Google-Extended is the separate token that lets you control whether your content can be used to improve certain AI models without blocking search crawling. Allowing Googlebot while disallowing Google-Extended is a common middle path for sites that want search traffic but prefer to opt out of AI training use where Google honors that signal.
Google also operates other specialized fetchers for products such as Bard era experiments, Vertex AI data connectors, and Notebook style tools, but documentation and naming have shifted over time. Rather than hardcoding a stale list, check current Google Search Central crawler documentation and your own logs each quarter. Look for Google-Extended plus any additional AI labeled tokens your logs show. Record what you find in your bot inventory with dates, because Google updates pages without broad announcements.
A frequent question is whether Google-Extended affects AI Overviews citations. Google states that AI Overviews build on search systems, so pages that are crawled by Googlebot and indexed for search remain eligible for AI Overviews even when Google-Extended is disallowed. In other words, the training opt out and the search answer layer are not the same switch. If AI Overviews visibility is your goal, focus on crawlability, index eligibility, content clarity, and topical coverage rather than assuming a single AI token controls everything. For background on how Google crawls and evaluates pages, see the Google documentation on crawling and indexing.
For implementation, keep Google rules minimal and explicit. Allow Googlebot for search, decide on Google-Extended based on training preference, and avoid wildcard blocks that accidentally catch Googlebot-Image, AdsBot, or verification fetchers. Test in Search Console URL Inspection after changes and monitor coverage for unexpected drops. A cautious, narrow rule beats a broad block that harms search traffic.
<!-- IMAGE-PROMPT diagram-01: 1600px max, DependsIt brand mint #22E3B0 on charcoal #121212, node-network line art, subject: AI crawler categories diagram showing training crawlers live answer fetchers and data pipelines connecting to website, flat vector, accessible, no em dash, Clash Display style headings, General Sans clean labels -->
Perplexity Cohere Meta and Mistral
Perplexity operates PerplexityBot for broad collection plus on demand retrieval that supports cited answers. The user agent contains PerplexityBot, and Perplexity publishes details on its behavior and control options. Sites that value answer engine visibility often allow Perplexity retrieval because citations can send qualified visits for research heavy queries. Sites with strict licensing rules sometimes limit PerplexityBot crawl rate or restrict it to public sections while keeping paywalled paths closed. Either way, separate bulk collection from on demand citation fetching in your notes so future reviews are clear.
Cohere operates Cohere-ai or similarly named agents for data collection supporting its models and enterprise retrieval products. Volume is lower than the largest consumer bots on most sites, but it appears regularly in logs for content heavy domains. Meta operates Meta-ExternalAgent and related Facebook external hit agents for AI training and link preview tasks. Meta documents these agents for AI use cases alongside classic Facebook crawler roles. Mistral AI operates MistralAI-User and related tokens at lower volume, primarily for model data and evaluation. None of these bots send large referral traffic on their own, so decisions hinge on training posture and operational load rather than click expectations.
For each of these vendors, the same workflow applies. Confirm the exact user agent token in current docs and in your logs. Note the stated purpose and the observed request rate on your site. Decide allow, limit, or block per token, not per vendor rumor. For example, you might allow PerplexityBot for public guides, allow Meta-ExternalAgent for preview rendering, and disallow bulk training agents for licensed archives. Write the reason in your inventory so the choice survives staff changes.
Watch for impersonation. Smaller or unknown agents sometimes copy well known tokens. Validate surprising spikes by checking reverse DNS, IP ranges published by the vendor, and request patterns. A genuine Perplexity or Meta agent will align with published ranges and behave consistently. A scraper using a fake token will often show odd paths, missing reverse DNS, or aggressive rates. Rate limiting and bot management at the CDN layer handle impersonation better than robots.txt alone, because bad actors ignore voluntary signals.
ByteDance Huawei Common Crawl and data pipelines
Bytespider, operated by ByteDance, collects pages for search, AI training, and content understanding across TikTok era products and related services. It is one of the higher volume AI related crawlers on many sites. Its user agent contains Bytespider. Site owners with global audiences see it frequently, sometimes at rates that rival mid tier search crawlers. If bandwidth or origin load is a concern, Bytespider deserves explicit attention in rate limits and crawl scheduling, not just a line in robots.txt.
Huawei operates PetalBot for search and AI related collection supporting Petal Search and related services. Its user agent contains PetalBot, with variants for different resource types. Common Crawl operates CCBot for its open web archive. CCBot collects broadly and publishes datasets that many research teams and startups reuse instead of crawling directly. Blocking CCBot reduces downstream reuse through that specific pipeline, though it does not block every training pipeline that may have already copied the data. Data growers and model teams also use additional pipeline bots that appear under varied names, so quarterly log review matters more than a one time list.
The policy question for pipelines is about breadth. A single robots rule for CCBot can affect many downstream users at once, which is efficient if your goal is to limit broad reuse. Conversely, allowing Common Crawl keeps your content available to researchers, non profits, and small teams that cannot run their own crawl. There is no single right answer. Publishers with licensed photography, paid research, or member only data often disallow pipeline bots for those paths while allowing them for public marketing pages. That selective approach balances openness with business constraints.
Operationally, pipelines reward the same hygiene as search. Keep sitemaps focused on canonical public URLs, avoid exposing infinite calendar, filter, and search result URLs to bulk crawlers, and use caching plus CDN rules to absorb spikes. For background on how crawl budget pressure builds on large sites, see crawl budget explained and how Google decides what to index. Although that guide focuses on Google, the logic of prioritizing important URLs and pruning waste applies to every high volume bot.
How to identify AI crawlers in logs
Start with the user agent field in your CDN or origin logs. Export the last 14 to 30 days, normalize case, and group by agent substring. Search for tokens such as GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Bytespider, PetalBot, CCBot, Meta-ExternalAgent, Cohere, and MistralAI. Count requests, distinct IPs, requested paths, response codes, and bytes transferred per token. This table tells you which bots actually matter for your site, rather than which bots matter in headlines.
Next, validate identity for high volume tokens. Check reverse DNS and forward DNS for the IP, compare against IP ranges or validation methods published by the vendor, and look at behavior. Genuine crawlers usually respect robots.txt, space requests, identify consistently, and fetch robots.txt itself before broad crawling. Scrapers using fake tokens often ignore robots.txt, hammer sequential IDs, request admin or API paths, and rotate through data center IPs without matching reverse DNS. Flag mismatches for WAF or rate limit action rather than robots edits alone.
Segment by path to find waste. AI crawlers often discover paginated archives, tag pages, calendar views, internal search results, and faceted filters that multiply URL count. If 60 percent of AI bot hits go to low value variants, fix discovery rather than blaming the bot. Update internal links, add noindex or canonical controls to variants, prune the sitemap to canonical URLs, and block parameter paths that should never be crawled. For robots audit steps that apply to all crawlers, see is robots.txt blocking your indexing and how to audit it. The same audit that protects search traffic also reveals AI crawl waste.
Finally, create a monthly log snapshot. Save the top 50 user agents by request count, the top AI tokens by bandwidth, and three example request lines per AI token with timestamps. Store the snapshot with your bot inventory. When someone asks why a new block was added or why traffic shifted, you will have evidence instead of memory. Consistent snapshots also make it easy to spot new tokens early, before they become cost or policy surprises.
Bandwidth cost and performance impact
AI crawler cost shows up in four places. Origin compute rises when bots request uncached HTML, especially faceted or search result pages that bypass cache. Egress and CDN billed requests rise with bytes transferred, including images, PDFs, and video posters when bots fetch media variants. Log volume grows, which increases storage and analysis cost. Most importantly, crawl pressure can delay fresh content discovery when servers respond slowly during bot spikes. Small sites rarely notice, but sites with hundreds of thousands of URLs or fragile origins feel it quickly.
Quantify before you act. For each AI token, calculate requests per day, average response time, error rate, cache hit rate, and estimated bytes per day. Multiply by 30 to estimate monthly load. Compare AI bot totals with Googlebot and Bingbot totals for context. On many content sites, all AI bots combined still request less than Googlebot, but a single aggressive collector such as Bytespider can exceed Bingbot on some weeks. That context prevents overreaction. You do not need to block every AI bot because one bot spiked on one day.
Mitigation starts with caching and path control, not broad blocks. Cache public HTML at the CDN where possible, set sensible TTLs for stable guides, and bypass cache only for truly dynamic paths. Prune crawler access to low value paths with robots rules and parameter handling. Use CDN rate limits or crawl delay style throttling where supported to smooth spikes without hard blocks. Reserve hard blocks for licensing or abuse cases. This order matters because caching helps human users and all bots at once, while blocks only affect compliant bots and do nothing against bad actors.
Also consider render cost. If your pages require heavy JavaScript rendering for core text, every bot fetch costs more compute and is more likely to time out or capture incomplete content. Server rendering core answers reduces cost and improves citability at the same time. For background on crawler directives and caching behavior, see the MDN guide to robots handling. Pair that with CDN analytics to confirm that cache hit rate improves after each fix.
<!-- IMAGE-PROMPT workflow-02: 1600px max, DependsIt brand, deep charcoal #121212 background, vibrant mint #22E3B0 accent glow, thin node-network line art, subject: workflow from log analysis to bot inventory to robots policy to monitoring dashboard, flat vector, accessible, no em dash, Clash Display style headings, General Sans clean labels -->
Deciding your crawl policy before you touch robotstxt
A good policy answers three questions. What content is public and reusable. What content is public for humans but not for bulk AI reuse. What content is restricted for everyone except authenticated users. Public marketing guides often fall in the first bucket. Licensed images, paid research, member directories, and client data fall in the second or third. Mapping paths to buckets before you write rules prevents accidental over blocking of search or under blocking of sensitive archives.
Next, map bots to purposes. Allow on demand answer fetchers that attribute and can send visits, unless a path is licensed. Decide on training crawlers based on long term positioning and rights, not weekly traffic. Decide on pipeline bots based on tolerance for broad downstream reuse. For most standard blogs and SaaS marketing sites, a common starting point is to allow answer fetchers and search crawlers, allow or limit training crawlers for public guides, and restrict training and pipeline bots for paywalled or licensed paths. Publishers, marketplaces, and stock media sites often choose stricter defaults for archives while keeping marketing pages open.
Involve the right owners. Editorial owns content value judgments. Legal owns licensing and rights posture. Engineering owns robots implementation, CDN limits, and log monitoring. SEO owns measurement of search impact and citation tracking. A 30 minute review with these four roles produces a clearer policy than weeks of async debate. Record decisions in a one page matrix that lists each bot token, each path bucket, and the chosen action with a date and owner.
Finally, plan rollout and rollback. Publish robots changes during low traffic hours, fetch the file as each major bot would, validate syntax, and monitor Search Console coverage plus origin error rates for 7 days. Keep the previous file version so you can revert quickly. Announce the change internally with the reason, because unexplained blocks create confusion when traffic or citations shift later. Policy first, syntax second, monitoring always.
Monitoring and maintaining your ai crawlers list
Bot lists decay quickly, so build a light maintenance loop. Monthly, refresh your top user agent snapshot, check vendor docs for renamed or new tokens, and update your inventory table with observed counts. Quarterly, re validate reverse DNS for high volume tokens, review bandwidth per bot, and confirm that policy still matches business goals. After any major site migration, template change, or paywall update, re test robots.txt, meta robots, and CDN rules immediately rather than waiting for the next cycle.
Watch for four warning signs. A new unknown token appears in the top 20 by volume. A known token suddenly increases tenfold week over week. Image or PDF bandwidth from bots spikes without new content launches. Search Console coverage drops after a robots edit. Each signal has a different fix. New tokens need identification and policy assignment. Sudden spikes need path analysis and rate smoothing. Media spikes need attachment and variant controls. Coverage drops need immediate rollback and narrow rule correction. For wildcard errors that often cause coverage drops, see wildcard rules in robots.txt that accidentally block indexing.
Automation helps at scale. Maintain your bot inventory as a simple CSV or sheet with columns for token, operator, purpose, first seen, last seen, requests per 30 days, policy, and owner. Generate robots.txt from that sheet with a script so manual edits do not drift. Alert on new tokens above a threshold, such as 500 requests per day or top 30 rank, so review happens only when warranted. Keep alerts quiet otherwise to avoid fatigue.
Document everything for the next person. Store the inventory, the generated robots file, the validation results, and the monthly snapshots in one folder with dates. When leadership asks whether to block a newly headlined bot, you can answer with your own data in minutes. That calm, evidence based routine is the real value of this guide, more than any static list. Bots will change names and owners, but a maintained inventory keeps you in control.
FAQ
Which AI crawlers visit most sites?
Most sites see GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Bytespider, CCBot, and Google-Extended plus Meta agents in logs. Exact mix depends on niche, size, and link profile. Export 30 days of user agents to see your own ranking before setting policy. Compare your logs with a public ai spider list or another crawler list ai roundup only after you rank your own top tokens.
What is the difference between GPTBot and ChatGPT-User?
GPTBot collects pages for model training in bulk. ChatGPT-User fetches a specific page when a user asks ChatGPT to browse it. Blocking GPTBot affects training collection. Blocking ChatGPT-User affects on demand browsing of your pages from chat sessions.
Does blocking an AI bot remove me from AI answers?
It depends on the bot. Blocking a training crawler reduces future training use but does not always remove live citations. Blocking a live answer fetcher is more likely to remove you from that engine's cited answers. Check whether the token is used for training, live retrieval, or both.
Do AI crawlers respect robots.txt?
Major providers state that their documented AI crawlers respect robots.txt, and most observed behavior matches that claim. Bad actors can fake user agents and ignore rules. Use robots.txt for compliant bots plus CDN rate limits and access controls for abuse and for sensitive paths.
How much bandwidth do AI crawlers use?
On many small sites, all AI bots combined use less than Googlebot. On large or media heavy sites, one aggressive collector can rival mid tier search crawlers for weeks. Measure requests, bytes, and cache hit rate per token for 30 days before deciding that cost requires action.
How often should I update my AI crawler list?
Review logs monthly and vendor docs quarterly. Update immediately after migrations, paywall changes, orulcoverage drops. Keep a dated inventory so new tokens are easy to spot and policy stays tied to evidence. Keep your ai bot list in the same sheet as other llm crawlers so new entries are easy to compare during each review.
Sources
- https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Robots
- https://developers.google.com/search/docs/crawling-indexing/overview