Indexer by DependsiT

llms.txt: The Complete Guide to the AI Crawler File

llms.txt file guiding AI crawlers to key site content on dark background

This guide explains llms.txt for site owners who want AI systems to find their most useful pages faster. It is for SEOs, developers, and content teams who already manage robots.txt and sitemaps and who want to know whether an additional AI discovery file is worth adding. You will learn what llms.txt contains, how to write it, how it differs from robots.txt, where adoption stands, and how to maintain it without creating extra work.

Key takeaways

  • llms.txt is a proposed Markdown file at the site root that lists key pages with short descriptions to guide AI models and assistants.
  • It does not replace robots.txt for permission or sitemaps for inventory, and most crawlers treat it as optional guidance rather than a rule.
  • A useful llms.txt is short, accurate, and kept in sync with canonical URLs, with links to full HTML pages rather than duplicated content.
  • Small sites can start with ten to twenty links, while docs and blogs should group links by task and keep descriptions factual.
  • Measure impact through crawl logs, AI referrals, and citation checks, and remove or fix the file if it drifts out of date.

llms.txt file guiding AI crawlers to key site content on dark background <!-- IMAGE-PROMPT cover: 1200x630, DependsIt brand, deep charcoal #121212 or clean white background, vibrant mint #22E3B0 accent glow, thin node-network line art, Clash Display style bold heading space on left, General Sans clean labels, subject: llms.txt document pointing to organized website pages for AI crawlers, flat vector, high contrast, accessible, no photorealistic faces, no text smaller than 24px, no em dash in rendered text, export PNG then cwebp -q 82 to WEBP -->

What llms.txt is and what it is not

llms.txt is a proposal for a plain Markdown file that helps language models discover the most relevant pages on a site. The idea is simple. Many sites have thousands of URLs, but only a few dozen carry the core facts that assistants need. A short curated list with clear titles and one line descriptions lets a model prioritize those pages when context is limited. The file lives at the root, uses Markdown headings and links, and points to full HTML pages where the complete content remains. It is guidance, not a feed of the content itself.

What it is not matters more. llms.txt is not a permission mechanism. It does not grant or deny crawling the way robots.txt does. A crawler that respects robots.txt will still respect robots.txt even if llms.txt lists a blocked URL, and listing a URL in llms.txt does not override a disallow. It is also not an index submission channel. Adding a URL to llms.txt does not place the page into ChatGPT or Perplexity answers by itself. The page must still be crawlable, parseable, and useful for specific prompts. Think of llms.txt as a reading list you leave on the table, not a key to the library.

The proposal comes from practitioner discussion about context limits and site structure. Models can only read a limited amount of text per task, and site navigation built for humans with menus, filters, and scripts is not always efficient for machines. A flat list of canonical guides, references, and FAQs in Markdown removes guesswork. It also gives small publishers a way to say which pages represent them best, rather than letting crawlers infer priority only from link counts. For docs sites and blogs with clear hubs, that signal is genuinely helpful.

Scope stays narrow by design. llms.txt should list public, stable, indexable pages that you want assistants to consult. It should not list admin pages, checkout flows, private accounts, internal search results, staging URLs, or parameter variants. It should not duplicate entire articles into the file. Long file bodies with full text make maintenance painful and drift quickly. Link to the canonical HTML and keep the description to what the page covers and who it is for. If a page stops being representative, remove it rather than leaving a stale pointer.

A common misunderstanding is that llms.txt replaces the need for clean HTML. It does not. Assistants still fetch and parse the linked pages, so those pages must render main content in HTML, carry accurate titles and dates, and avoid blocks for AI agents where visibility is wanted. The broader AI indexing workflow still applies, as described in the overview of AI search indexing for ChatGPT and Perplexity. Treat llms.txt as a small addition on top of solid crawl foundations, not as a shortcut around them.

Practical expectations help teams decide how much effort to invest. For most sites, a first version takes one to two hours: inventory top guides, write one line descriptions, publish at the root, and verify fetch. Maintenance is minutes per month if tied to the publish checklist. Impact is indirect and gradual, through better prioritization rather than instant citations. Sites with messy navigation or many weak pages benefit most, because the file compensates for poor discovery. Sites with a handful of strong, well linked pages may see little change, and that is fine. The file is cheap to keep when it is short and honest.

A concise llms txt file lists your best canonicals with one line descriptions so models can prioritize hubs. Follow this llms.txt guide pattern for small sites first, with five to fifteen links grouped by task before expanding.

Where llms.txt lives and how crawlers find it

By convention, llms.txt lives at the site root next to robots.txt and the main sitemap. For a site at example.com, the file is fetched at the root path with the name llms.txt. Some proposals also describe an llms-full.txt variant with more detail, but the core file stays concise. Keep the filename lowercase exactly, serve it with a plain text or Markdown content type, and allow unauthenticated GET from any user agent you want to guide. If the file requires login, blocks bots at the edge, or returns a redirect chain, most consumers will ignore it.

Discovery is straightforward but not standardized. There is no single registry where you submit the file. AI systems that support the convention fetch it directly when they already know your domain, often after finding you through classic crawl or a user provided link. Others may never fetch it. That is why classic discovery still matters. Keep sitemaps accurate, keep internal links clean, and keep robots.txt permissive on public paths so assistants find your domain in the first place. llms.txt then helps them choose among your pages, but it does not introduce your domain to systems that have never seen it.

Hosting details affect reliability. Serve the file over HTTPS with a 200 status, set a reasonable cache lifetime such as one hour to one day, and ensure CDN rules do not serve a stale copy for weeks after edits. Avoid redirecting the root file to another host or to a JavaScript app route that returns HTML shells. Test from outside your network with a simple fetch and confirm the body matches what you published. Log requests for the path by user agent to see which systems actually retrieve it. Early on, expect sparse hits. That is normal and does not mean the file is broken.

Subdomains and international variants need explicit decisions. If docs live on a subdomain and the blog on the main domain, each host that should guide models needs its own file scoped to its own canonical URLs. Do not point the main domain file at staging hosts or at URLs that canonicalize elsewhere, because that teaches consumers to distrust the list. For multilingual sites, either keep one file with clearly labeled language sections or keep per locale files on each locale host. Match the approach to how your sitemaps and hreflang are organized so all discovery signals agree.

Access control should mirror robots.txt intent. If a section is public and you want AI answers to use it, list only URLs that are allowed for the agents you target. If a section is public for humans but you prefer AI systems not to train on it, use robots.txt and edge policy for enforcement, not llms.txt omission alone, because omission is not a block. Document the policy next to the file in your runbook: which paths are listed, which agents are allowed, who approves additions, and how removals are handled. Clear ownership prevents the file from becoming an unreviewed dump of links.

Finally, link the file where humans can find it. Reference it from your docs contribute guide and your SEO checklist so new pages are considered for inclusion at publish time. There is no need to link it prominently in site navigation for users, but keeping it discoverable in your repo and deploy pipeline ensures updates ship with content changes rather than lagging behind. A file that updates with deploys stays trustworthy. A file edited once and forgotten becomes another stale signal that crawlers learn to ignore.

Diagram showing llms.txt at site root linking to canonical guides <!-- IMAGE-PROMPT diagram-01: 1600px max, DependsIt brand mint #22E3B0 on charcoal #121212 or white, node-network line art, subject: site root with llms.txt file branching to canonical guide pages for AI consumption, flat vector, accessible, no em dash -->

llms.txt syntax block by block

The syntax stays close to plain Markdown so it is easy to write and parse. Start with an H1 naming the site or project, followed by a short paragraph stating what the site covers and who the content is for. Then add an H2 for primary guides with a bullet list of links, each link followed by a colon and a one sentence description. Optionally add H2s for reference docs, FAQs, and data or changelogs. Keep line lengths reasonable, use standard Markdown link format with absolute or root relative URLs that resolve to canonical pages, and avoid embedded HTML, scripts, or images. Parsers expect simple structure.

A typical skeleton has four parts. The header block names the project and its scope in two to three sentences. The core guides block lists five to fifteen essential how to and explainer pages. The reference block lists API docs, schemas, limits, and error catalogs where relevant. The updates block links to the changelog, status page, or version notes so models can check freshness. Not every site needs all four. A small business site might have only header plus core guides. A docs site should include reference and updates. Choose blocks that match how users actually seek help.

Descriptions carry the weight. Each description should state the task the page covers, its scope, and any version or audience note in under 25 words. Good descriptions name verbs and objects, such as setup, limits, errors, or migration, plus the product or topic they apply to. Vague descriptions such as useful information or various tips give models nothing to match against prompts. Write descriptions for retrieval, not for marketing. If you cannot describe the page plainly in one line, the page itself may need a clearer focus before it belongs in the file.

Link choice matters more than link count. Prefer canonical URLs that return 200, have self referencing canonicals, and appear in your XML sitemap. Avoid redirects, parameter variants, print views, and paginated fragments. For guides split across several URLs, either list the hub plus each part with clear labels or consolidate into one stronger URL and list that. Keep trailing slash behavior consistent with your canonical rules so the file does not teach two forms of the same address. After drafting, click every link from outside the CMS preview and confirm it resolves to the intended live page.

Formatting rules keep the file machine friendly. Use ATX headings with hash signs, unordered lists with dashes, and inline links with parentheses. Avoid tables for the main list, because simple bullets parse more reliably across consumers. Avoid code fences for the link list itself, though a short fenced example elsewhere in docs is fine. Keep the file under a reasonable size, often under 100 lines for small sites and under 500 lines for large docs. If the file grows beyond that, split detail into linked hub pages rather than stuffing everything into one file.

Validation is manual but quick. Confirm the file parses as Markdown, that every link returns 200, that no listed URL is blocked by robots.txt for the agents you target, and that no listed URL carries noindex. Check that descriptions match the live page titles and scopes. Revalidate after template changes, because a CMS update can alter canonicals or add parameters that break the correspondence between the file and the live site. A monthly link check plus a check on every deploy keeps the file trustworthy with little effort.

Minimal example for a small site

Small sites need only a short file that points to the pages that answer common questions. Imagine a ten page business site with a home page, three service pages, pricing, about, contact, and three guides. The llms.txt file should list the three guides plus pricing and one service overview, not every page. Five to eight links with plain descriptions are enough. The goal is to tell an assistant where the substantive answers live, not to mirror the full navigation.

A minimal file starts with the site name and a two sentence scope note. For example, state that the site covers a specific service in a specific region, plus how to guides for common tasks. Then list the guides under a Primary guides heading, each with a link and a description naming the task. Add pricing and contact only if they contain facts assistants need, such as price ranges, service areas, or hours. Omit legal pages, thank you pages, and thin category pages that add no unique information. Brevity here is a feature.

Keep descriptions concrete. Instead of writing pricing information, write what the pricing page actually contains, such as plan names, included items, and how quotes work. Instead of writing guide to our process, write the steps the guide covers and who it is for. Specific descriptions help the model decide whether the page matches a prompt about cost, timing, or requirements. They also force you to notice when a page is too vague to be useful in answers, which is itself a useful audit.

Publish the file at the root and verify in three ways. First, fetch it anonymously and confirm 200 plus the exact body you wrote. Second, click each link and confirm the live page matches the description. Third, confirm each listed page is indexable: no noindex, allowed in robots.txt, canonical self referencing, and present in the sitemap. Fix any mismatch before announcing the file internally. A minimal file that is fully accurate beats a longer file with two broken links, because broken links teach consumers to deprioritize the whole list.

Maintenance for a small site is light. Review the file quarterly or whenever you add, remove, or rewrite a core page. Update descriptions when page scopes change. Remove pages you retire with redirects, and add the replacement if it carries the same task. Keep a short note in your publish checklist: if a page is important enough to be in llms.txt, it is important enough to keep fast, accurate, and clearly dated. That single rule keeps the file and the pages it points to aligned without extra process.

A minimal llms txt example has a title block, three to five grouped links, and short task focused descriptions. Copy that shape first, then test that linked URLs are canonical, crawlable, and allowed by robots.

Full example for docs or a large blog

Docs sites and large blogs need more structure, but the same principles apply. Group links by task so models can pick the right section for a prompt. Common groups include getting started, how to guides, reference, troubleshooting, and updates. Under each group, list canonical hubs first, then key child pages. Aim for clarity over completeness. A docs site with 400 pages might list 40 to 80 in llms.txt, focusing on installation, authentication, core workflows, limits, errors, and migration. A blog with 300 posts might list 20 to 40 pillar guides plus a link to the full archive or sitemap for the rest.

The header should state product name, covered versions, and audience in plain terms. For versioned docs, note the current stable version and where older versions live. For blogs, note the topics covered and the update cadence. This context prevents models from mixing versions or treating a five year old post as current. Under getting started, link the install or setup hub plus prerequisites. Under how to guides, link task pages with descriptions that name inputs, outputs, and preconditions. Under reference, link endpoint catalogs, schemas, and limit pages. Under troubleshooting, link error guides with the exact codes they cover.

Descriptions for docs must include scope and version where relevant. State which version a page applies to, which plan or role it requires, and what changed in the last update. For example, note that an auth guide covers a specific token flow for a specific API version, not all auth for all products. These scope notes prevent misquotation where a model applies version specific steps to the wrong version. They also help human readers who land on the page from an AI citation and need to confirm it matches their setup.

Link hygiene is critical at scale. Use canonical URLs only, keep trailing slashes consistent, and avoid linking to paginated fragments or anchor only variations as separate entries. If a guide spans multiple pages, list the series hub with a description of the full task, then list each part with its step range. Ensure each listed URL appears in the appropriate sitemap section with an honest lastmod. Run automated link checks on every deploy and on a weekly schedule, because docs churn breaks links faster than teams expect. A docs file with even a small share of dead links loses trust quickly.

Governance keeps a large file accurate. Assign ownership by section so each team maintains its own links and descriptions. Require that new canonical guides be considered for inclusion at publish time, with a default of no unless the page is truly core. Require that rewrites update the description in the same pull request. Keep the file in version control next to docs source so reviews catch drift. Render a human readable copy in your contribute guide that explains what belongs and what does not, with three good and three bad examples. Clear rules prevent the file from growing into an uncurated dump.

llms.txt vs robots.txt vs sitemap

These three files solve different problems and work best together. Robots.txt controls permission: which crawlers may fetch which paths. Sitemaps provide inventory: which canonical URLs exist and when they last changed. llms.txt provides prioritization: which public pages best represent the site for AI tasks. Confusing their roles causes most implementation mistakes, such as listing blocked URLs in llms.txt, omitting key URLs from sitemaps because they are in llms.txt, or treating llms.txt as a block.

A healthy setup aligns all three. Robots.txt allows AI search agents on public content paths and blocks private, duplicate, or wasteful paths such as account areas, cart flows, and internal search results. Sitemaps list all indexable canonical URLs in those allowed paths with accurate lastmod values. llms.txt highlights a subset of those sitemap URLs with human written descriptions of what each covers. When a crawler compares the three, it sees consistent permission, complete inventory, and clear priority without contradictions. Inconsistent setups, such as llms.txt pointing at URLs missing from sitemaps or blocked by robots.txt, force crawlers to guess and reduce trust in all signals.

Caching and update cadence differ. Robots.txt changes should propagate quickly and are fetched often. Sitemaps update on publish and reflect lastmod for changed URLs. llms.txt can cache longer, such as a day, because it changes less often, but it must still update when core pages move or retire. Keep each file served with appropriate cache headers and verify after deploys. Log fetches by path and agent to confirm consumers see fresh copies. Stale robots.txt causes crawl outages. Stale sitemaps cause delayed discovery. Stale llms.txt causes misdirected prioritization. All three deserve monitoring, with robots.txt monitored most closely.

Security boundaries stay with robots.txt and edge policy, not with llms.txt. Do not rely on omitting a URL from llms.txt to keep it private. If a URL must not be crawled or used, block it properly and require auth where needed. Conversely, do not list URLs in llms.txt that require login or that return different content to bots than to users in ways that mislead. The file is public by design. Assume any URL and description you publish will be read by many systems, including ones you did not target. Publish accordingly.

For teams that already automate IndexNow or sitemap pings for classic search, keep that automation and add llms.txt as a separate lightweight step. On publish, update the sitemap and lastmod, queue classic submission where appropriate, and consider whether the new URL belongs in llms.txt. Most new posts do not. Only promote a new URL to llms.txt when it becomes a core answer for a recurring task. This restraint keeps the file short and the signal strong. The interaction with classic submission is covered further in the beginners guide to generative engine optimization.

Comparison of robots.txt permission sitemap inventory and llms.txt priority <!-- IMAGE-PROMPT workflow-02: 1600px max, DependsIt brand, deep charcoal #121212 or clean white background, vibrant mint #22E3B0 accent glow, thin node-network line art, Clash Display style bold headings, General Sans clean labels, subject: three file roles diagram comparing robots sitemap and llms files, flat vector, accessible, no em dash -->

The llms txt vs robots question is common, and the roles differ clearly. Robots controls permission and sitemaps provide inventory, while llms.txt suggests priority, so keep all three consistent on every deploy.

Adoption status across AI systems

llms.txt is a proposal, not a standard, and adoption varies. Some AI tooling and assistants fetch the file when present, especially developer focused products that already work with Markdown. Others ignore it and rely on classic crawl plus retrieval. No major assistant publishes a guarantee that listing a URL ensures citation. Treat adoption as emerging and uneven. The file is cheap to maintain when short, so early adoption is reasonable, but it should not divert effort from crawl access, HTML quality, and content clarity that all systems use.

Differences in behavior are worth noting. Systems that fetch the file tend to use it for prioritization and context building rather than as an index. They still fetch the linked HTML to verify facts before citing. Systems that ignore the file are unaffected by its presence, good or bad, provided the file does not leak private URLs or create confusion. There is no penalty for publishing an accurate file, but there is cost to publishing a sloppy one if humans or tools follow stale links. Accuracy matters more than enthusiasm.

Versioning adds uncertainty. Proposals mention variants such as a fuller file for detailed context, but naming and parsing differ across consumers. Stick to the widely described root llms.txt with simple Markdown and canonical links. Avoid exotic extensions, custom directives, or large embedded content that only one tool understands. Standard simplicity maximizes the share of consumers that can use the file and minimizes maintenance when proposals evolve. If the ecosystem converges on a stricter spec, a simple file will be easy to adapt.

Publisher policies also shape adoption. Some organizations prefer AI systems not to use their content for training while still wanting citations in answers. llms.txt does not encode that distinction. Use robots.txt, terms, and licensing signals for policy, and use llms.txt only for guidance within allowed scope. Keep legal and editorial aligned on what listing implies. Listing a page suggests you consider it representative and suitable for reference. Do not list pages under rights review or with uncertain sourcing.

Practical guidance given this status is measured. Publish a short accurate file if you maintain docs, guides, or reference content that assistants already cite or should cite. Skip it for now if your site is tiny with clear navigation and no AI referral opportunity, and revisit when content grows. In both cases, keep classic foundations strong, because every AI system benefits from them regardless of llms.txt support. Reassess quarterly by checking fetch logs, citation tests, and any new consumer documentation. Adopt incrementally rather than betting heavily on one mechanism.

Current llms.txt adoption is early and uneven across AI systems, so treat the file as optional guidance. Foundations like fast HTML, clear headings, and sitemaps still decide whether you earn citations.

Should you add llms.txt now

The decision depends on site shape, content value, and maintenance capacity. Add it now if you have ten or more guides or docs pages that answer recurring tasks, if navigation buries those pages more than three clicks deep, or if AI referrals already appear in analytics and you want to steer them toward canonical hubs. These are signs that prioritization guidance will help. A focused file can surface the right hubs faster than link graph changes alone, especially for new sections that have not yet earned many internal links.

Delay if the foundations are broken. If core pages block AI agents, render main content only after heavy JavaScript, carry conflicting canonicals, or show stale dates, fix those first. An llms.txt file that points to blocked or broken pages does more harm than good. Similarly, delay if you cannot keep the file in sync. A file that lists retired URLs or mismatched descriptions teaches consumers to ignore it. In that case, invest the same hour in sitemap accuracy and internal linking, which help all crawlers immediately, then add llms.txt once publishing discipline is steady.

Consider content type. Docs, how to blogs, references, and FAQs map well to llms.txt because they answer bounded tasks with stable URLs. News, user generated forums, and rapidly expiring listings map less well, because prioritization shifts daily and a static list ages fast. For those sites, a small file pointing to evergreen explainers plus topic hubs is still useful, but do not try to list every fresh item. Let sitemaps and feeds handle freshness while llms.txt handles orientation.

Weigh maintenance realistically. A good file needs ownership, review on publish, link checks on deploy, and quarterly pruning. For a small team, that is minutes per month if the file stays under 50 lines. For a large docs team, it needs section owners and automated checks, which most docs pipelines already have for sitemaps. If your team already maintains clean sitemaps and changelogs, adding llms.txt is a small increment. If sitemaps are already stale, adding another file will likely repeat the same neglect. Fix the existing loop first.

A simple decision rule works. If you can name ten canonical URLs that should represent your site in answers and you can keep their descriptions accurate, publish the file. If you cannot name ten without including weak or duplicate pages, improve content and structure first. Either path leads to better AI visibility, because the exercise forces clarity about which pages deserve citations. The broader prioritization tactics appear alongside other answer focused methods in the beginners guide to generative engine optimization.

Publishing and maintaining llms.txt

Publishing takes four steps. First, draft the file in version control with header, core guides, and optional reference and updates sections. Keep lines short and descriptions factual. Second, deploy to the root path and verify anonymous fetch returns 200 with the exact body. Third, validate every link for status, robots access, canonical consistency, and sitemap presence. Fourth, add the file to your publish checklist and monitoring so future changes keep it in sync. Announce internally with clear ownership so edits go through review rather than ad hoc patches.

Integrate with your stack concretely. For static sites, keep llms.txt next to robots.txt in the public folder and generate the link list from the same content collection that builds sitemaps, so canonicals cannot drift. For CMS sites, create a template or shortcode that renders the file from curated hub entries, with caching of one day and manual review for additions. For docs frameworks, generate sections from nav metadata plus hand written descriptions stored in frontmatter. Automation helps with link accuracy, but descriptions should stay hand written to keep retrieval quality high.

Monitoring covers both the file and its targets. Alert if the file returns non 200, if its size changes unexpectedly, or if weekly link checks find broken or redirected entries. Alert if a listed URL gains noindex, becomes blocked in robots.txt, or drops from the sitemap. Track fetch counts by user agent to see which systems retrieve the file, without over interpreting sparse data. Pair file health with page health: time to first byte, text only content presence, and canonical stability for listed URLs. Most issues are caught by these simple checks before citations suffer.

Change management keeps trust. When you retire a listed page, update llms.txt in the same deploy as the redirect, pointing to the replacement with a revised description. When you rewrite a guide, update its description if scope changed. When you launch a new core guide, consider it for inclusion after it has stable URLs and internal links, not on day one before it is proven. Keep a short changelog for the file itself in your repo so later audits can see what was listed when. These habits mirror good sitemap discipline and take little extra time.

Access and security review belongs in the same routine. Confirm the file lists only public URLs, contains no internal hostnames, tokens, or private paths, and does not describe restricted content in ways that invite probing. Review edge rules to ensure the file is served consistently across regions and that bot management does not challenge its path for the agents you want to guide. Keep the file minimal to reduce exposure. A short list of public hubs is easy to secure. A sprawling dump of deep links is harder to review and more likely to include something it should not.

For llms txt wordpress setups, generate the file from the same source as your sitemap or with a small deploy script. For llms txt seo maintenance, check links on deploy, prune retired URLs with redirects in the same release, and review quarterly for accuracy.

Measuring impact and avoiding pitfalls

Impact from llms.txt is indirect, so measure with multiple signals rather than one metric. Watch server logs for fetches of the file by AI agents to confirm it is discovered. Watch crawl frequency for listed URLs to see if prioritization improves. Run fixed prompt tests in ChatGPT and Perplexity to track whether listed hubs earn citations for their target tasks. Segment AI referrals by landing page to see if listed URLs gain share relative to unlisted pages. No single signal proves causation, but consistent movement across several signals after a clean publish suggests the file helps.

Avoid common pitfalls that erase gains. The most frequent is drift: listed URLs redirect, descriptions no longer match titles, or new canonicals replace old ones without updating the file. The second is bloat: adding every new post until the file is hundreds of lines with no priority signal. The third is duplication: listing parameter variants or print views alongside canonicals, which teaches consumers that the list is noisy. The fourth is permission mismatch: listing URLs blocked by robots.txt or carrying noindex, which wastes the consumer effort spent following the file. Quarterly pruning plus deploy time checks prevents all four.

Do not expect instant citations. Even with a perfect file, assistants cite pages per query based on usefulness, freshness, and competition. New files need weeks for fetchers to notice and for retrieval to reflect the prioritized hubs. Judge after at least two crawl cycles and several prompt test rounds, not after two days. If listed pages still do not appear, the issue is usually page level: weak opening answers, missing tables, stale dates, or slow rendering. Improve the pages first, then reassess the file. The file amplifies strong pages. It does not rescue weak ones.

Compare costs honestly. For most sites, llms.txt costs an hour to launch and minutes to maintain, with modest upside in orientation and prioritization. That is a good trade when foundations are solid. It is a poor trade when the same hour could fix blocking robots rules, broken sitemaps, or thin core guides that affect every crawler. Prioritize universal fixes first, then add the file as a small enhancement. Report internally in those terms so stakeholders understand what the file does and does not promise.

Finally, keep the file honest as the ecosystem evolves. If consumers publish clearer specs, align formatting promptly. If your analytics show no fetches and no citation movement after several months despite strong pages, consider sunsetting the file rather than maintaining noise. Either decision is defensible when based on logs and tests. What matters is the underlying discipline: canonical URLs, accurate sitemaps, clean internal links, and quotable pages. Those assets help every current and future AI consumer, with or without llms.txt.

FAQ

Where should llms.txt live on my site?

Place it at the site root with the exact lowercase name llms.txt, served over HTTPS with a 200 status to unauthenticated GET. Keep it next to robots.txt and reference only canonical URLs on the same host where possible. For subdomains with distinct content, maintain a separate file per host. Verify from outside your network that the body matches what you published and that caching does not serve stale copies for weeks. Place it at the site root with the exact lowercase name and HTTPS 200 status. A clean llms txt file next to robots.txt should list only canonicals. This llms.txt guide setup keeps discovery simple and caching predictable for crawlers.

Does llms.txt replace robots.txt or sitemaps?

No. Robots.txt controls permission, sitemaps provide inventory with lastmod, and llms.txt suggests priority with descriptions. All three should agree: listed URLs must be allowed by robots.txt and present in sitemaps as canonicals. Treating llms.txt as a block or as a submission channel causes mismatches that reduce trust. Maintain each file for its own role and check consistency on every deploy. No, each file has its own role and they must agree. The llms txt vs robots distinction matters because permission stays with robots while priority hints stay with llms.txt. Check consistency on every deploy to protect trust.

What should I include in llms.txt?

Include public, stable, canonical pages that best answer recurring tasks: core guides, references, FAQs, and update logs. Write one line descriptions that name the task, scope, and version where relevant. Exclude private pages, duplicates, parameter variants, staging URLs, and thin pages with no unique value. For most small sites five to fifteen links are enough. For large docs, group by task and keep the total focused rather than exhaustive. Include stable public canonicals that answer recurring tasks with one line descriptions. A focused llms txt example with five to fifteen links beats a long stale list. Prune thin pages and keep grouping by task for clarity.

How long should descriptions be?

Keep each description under about 25 words, stating what the page covers and who or which version it is for. Name verbs and objects such as setup, limits, errors, or migration plus the product or topic. Avoid vague phrases that give retrieval nothing to match. If a page needs more than one sentence to describe, consider whether the page itself needs a clearer focus before it belongs in the file. Keep each under about twenty five words with verbs and scope. For llms txt seo value, name the task, version, and audience so retrieval can match. Rewrite the page itself if one sentence cannot describe it clearly.

Will adding llms.txt get me cited in ChatGPT or Perplexity?

Not by itself. Citations depend on crawlable HTML, clear structure, useful content, and query level competition. llms.txt can help models prioritize your best hubs once they crawl your site, but the linked pages must still be fast, accurate, and quotable. Treat the file as optional guidance after foundations are solid, then measure with prompt tests and referral segments over several weeks. Not alone, since citations need crawlable HTML and quotable structure. Early llms.txt adoption helps prioritization once crawling works. Measure with prompt tests and referral segments over several weeks after foundations are solid.

How do I maintain llms.txt without extra overhead?

Store it in version control, generate links from the same source as sitemaps where possible, and require description updates in the same pull request as page scope changes. Add link checks to deploy and weekly jobs, review quarterly for pruning, and update the file in the same deploy as redirects when pages retire. Assign section owners for large sites so accuracy stays high with little coordination cost. Store in version control and generate from sitemap sources where possible. For llms txt wordpress sites, require description updates in the same pull request as page changes. Add link checks to deploy and weekly jobs to keep the ai discovery file accurate.

Sources

  • https://developer.mozilla.org/en-US/docs/Glossary/Robots.txt
  • https://schema.org/Article

Further reading

Put this into practice. Indexer submits URLs to the Google Indexing API and IndexNow, audits coverage with Search Console, and shows exactly which pages are indexed. Start free or see how it works.