A short file, a lot of noise. Here is what llms.txt actually is, a complete worked example you can copy, and a straight answer on whether it does anything for your visibility. The answer is smaller than the marketing suggests.
llms.txt is a markdown file at the root of a website that lists your most useful pages with a line of description each, so a language model reading your site gets an orientation rather than a navigation menu. It is a community convention, not a standard. No major assistant has committed to using it. Publish one because it costs an hour, not because it will lift your citations.
llms.txt was proposed in September 2024 by Jeremy Howard as a way to solve a narrow and real problem. A language model reading a web page receives a great deal of material that is not content: navigation, cookie notices, footers, promotional interruptions, script tags. Context windows are finite and attention is not free. The proposal was that a site could publish a small markdown file at a known location, listing the pages a human maintainer considers most useful, each with a sentence explaining what it contains, so a model arriving at the site has a curated map rather than a homepage.
That is the entire idea. It is a table of contents written for a machine reader, in markdown because markdown is close to the format models are most comfortable with, at a fixed path so it can be found without being linked. There is nothing more sophisticated underneath it, and the simplicity is the appeal.
It is worth being precise about what it does not do, because most of the confusion in the market comes from assuming it does more. It does not grant or withhold permission to crawl. It does not control training use. It does not carry directives, only links and prose. It has no version negotiation, no conditional logic and no enforcement. If a model ignores it, nothing happens and you have no way to know.
The convention is per host, like robots.txt. A site with a documentation subdomain and a marketing domain needs two files, and they should describe different things, because the audiences and the useful pages are different.
The format is plain markdown with a light structure. A single H1 with the name of the site or the organisation. An optional blockquote immediately underneath giving a short summary. Then H2 sections, each containing a bulleted list of links, where every bullet is a markdown link followed by a colon and one line of description. A section headed Optional carries material that can be dropped if the reader is short of space, which is the only piece of conditional meaning the format has.
Here is a complete example for a consultancy. This is close to what we publish, and it is short enough that a person can read the whole thing and confirm it is accurate, which is the property that matters most.
# Green Arrow Consultancy
> A UK AI agency and digital consultancy, founded 2012 in Cardiff. We build production AI
> systems and run privacy, AI governance, AI security and accessibility programmes for
> global consumer brands. Company number 12491770.
Contact: info@greenarrowconsultancy.com · +44 29 2117 5902
Registered office: Sophia House, 28 Cathedral Road, Cardiff, CF11 9LJ, Wales, UK
## Services
- [AI consulting](https://greenarrow.app/services/ai-consulting/): Strategy, build and
operation of production AI systems, including retrieval assistants and enterprise search.
- [AI search optimisation](https://greenarrow.app/services/ai-search-optimisation/): Making a
site reachable, extractable and citable by ChatGPT, Claude, Gemini and Perplexity.
- [AI governance](https://greenarrow.app/services/ai-governance/): EU AI Act, ISO/IEC 42001
and the NIST AI Risk Management Framework turned into controls a team can operate.
- [Privacy consulting](https://greenarrow.app/services/privacy-consulting/): UK GDPR and
global privacy programmes, consent management, records of processing.
- [Accessibility](https://greenarrow.app/services/accessibility/): WCAG 2.2, EN 301 549 and
European Accessibility Act conformance for websites and documents.
## Reference
- [Glossary](https://greenarrow.app/glossary/): Definitions for AI, privacy, accessibility
and AI search terms used across this site.
- [Frequently asked questions](https://greenarrow.app/faq/): Pricing, ways of working,
data handling and contracting.
- [About](https://greenarrow.app/about/): Company history, registrations and the team.
## Evidence
- [Case studies](https://greenarrow.app/case-studies/): Named client engagements.
- [Insights](https://greenarrow.app/insights/): Long-form articles on AI search, crawler
access, governance and privacy engineering.
## Optional
- [Neuro demonstrations](https://greenarrow.app/neuro/): Seven production AI systems
generalised into live demos running on sample data.Note what is happening in that file. Every link carries a description that says what the reader will find, not what we would like them to think. The identifying details appear once, near the top, in a form that can be checked against a public register. The sections are organised by what someone might want, not by our internal org chart. And it is short, which is the point of the exercise. A file of two hundred links is a sitemap with extra steps.
You can compare it against the live version at /llms.txt on this site.
Four files, four different jobs. Only one of them is permission, only one of them is inventory, and only one of them has thin evidence behind it.
| File | What it declares | Who reads it | Strength of evidence that it matters |
|---|---|---|---|
| llms.txt | Orientation. Which pages a maintainer thinks are worth reading, and why | Some developer documentation tools and coding assistants. No confirmed use by major assistants | Weak. No operator has committed to it as an input, and the correlational claims are confounded |
| robots.txt | Permission. Which agents may fetch which paths, plus content signals where used | Every compliant crawler, including all the major AI operators | Strong. A disallow demonstrably removes you from compliant crawlers |
| sitemap.xml | Inventory. Every canonical URL you want discovered, with change and priority hints | Search engine crawlers, and indirectly the AI search indexes built on them | Strong for discovery on large or poorly linked sites. Weak as a ranking input, which it was never meant to be |
| Schema.org JSON-LD | Meaning. What the entities on this page are, and how they relate to external records | Search engines, knowledge graph builders, and retrieval systems doing entity resolution | Moderate. Not a ranking factor, but it demonstrably reduces the ambiguity that causes a source to be skipped |
Evidence ratings are our assessment based on operator documentation and our own testing, not a published benchmark.
The companion convention is llms-full.txt, and it inverts the idea. Instead of links to pages, it contains the pages: the full text of the documentation, converted to markdown and concatenated into one file, so a model can ingest the entire corpus in a single fetch without following twenty links and stripping twenty sets of navigation.
For technical documentation this is genuinely useful, and it is where the convention has real traction. A developer pasting a library's llms-full.txt into a coding assistant gets a working reference in one action, and several documentation platforms now generate the file automatically as part of the build.
For a marketing site it is close to pointless. Your services pages concatenated into a single markdown document are a few hundred kilobytes of persuasive prose with no reference value, and no assistant is going to load them to answer a question. If your content is reference material that someone would want to consult in full, publish it. If it is content designed to be read one page at a time, do not.
There is also a maintenance cost that gets skipped in the enthusiasm. A generated llms-full.txt drifts from the site the moment a page changes, and a stale full-text copy of your documentation is a genuine liability if anyone acts on it. Only generate it as part of the build, never by hand.
The honest map of adoption looks like this. Developer tooling and documentation platforms lead, because the convention came out of that world and solves a problem those teams feel every day. A number of AI companies and API providers publish one for their own documentation, which is unsurprising given the audience. A growing set of software companies have added one because a developer suggested it in a sprint.
Outside that, adoption thins quickly. Most large consumer brands do not publish one. Most news organisations do not. Most professional services firms do not. Where you do see it on a marketing site, it is frequently auto-generated from the sitemap by a plugin, which produces a file with none of the curation that gave the idea its value in the first place.
The pattern is worth reading carefully, because it tells you something about who the file is for. It is adopted where the content is reference material consumed by machines as part of a working process, and it is largely ignored where the content is persuasion consumed by people. That is not a criticism of either. It just means the case for publishing one is stronger if you have documentation than if you have brochures.
We publish one. We do not count it in any forecast, and we would not put it above crawler access, extractable structure or third-party corroboration on anybody's list.
If you are going to publish one, publish a good one. The difference between a curated file and a generated one is the entire value of the exercise.
The pages that answer questions rather than the pages that sell. Your definitions and glossary, because definitions are the most useful thing a model can find on a site. Your pricing or engagement model if it is public, because it is the most-asked question and the least-published answer. Your identifying details, once, in a checkable form. Your reference documentation, your FAQs and your genuinely substantive articles. A one-line description per link that would still be accurate if a sceptical reader clicked through.
Everything you would not defend in person. Campaign landing pages, gated assets, thin category pages, duplicate regional variants of the same content, and anything behind a login. Marketing adjectives, which cost you space and credibility at the same time. Keyword stuffing, which reads exactly as badly to a model as it does to a person. And any page you would not want quoted, because a curated list of pages is an invitation to quote them.
Give it an owner and a review date. A file that points at three pages which now return 404 is a live signal that the site is not maintained, and that is a worse outcome than never having published it. If nobody will own it, generate it from the build or skip it.
Yes, and then stop thinking about it. That is the whole recommendation.
It takes an hour, it cannot hurt you if it is accurate, the exercise of writing it is quietly useful, and if the convention does gain traction you are already there. Those are good enough reasons. They are not the reasons being given in most of the content written about it, which promise a visibility improvement nobody has demonstrated.
What we would resist is the reordering of priorities that often comes with it. We have been in rooms where a team was discussing llms.txt for forty minutes while their web application firewall was returning 403 to every AI crawler that mattered, which is a much larger and much less interesting problem. The cheap thing gets attention because it is cheap. The expensive thing gets deferred because it involves another department.
The ordered version of this work is in our citation checklist, and llms.txt does not appear on it. That is deliberate. If you want the difference between this kind of tactic and the underlying discipline spelled out, the piece on GEO and SEO covers it, and the service page describes what we actually do about it.
Short answers, including the ones that are less flattering to the file.
A markdown file at the root of a website that lists the pages an author considers most useful to a language model, with a short description of each, so a model reading the site has an orientation instead of a navigation menu.
At the root of the domain, served as plain text or markdown at https://example.com/llms.txt. Same location convention as robots.txt and security.txt. It is per host, so a subdomain needs its own.
No. It is a community proposal, published in 2024 and adopted by a slice of the developer tooling world. It has no standards body behind it, no registry, and no formal specification process. That is not a criticism, plenty of useful web conventions started the same way, but it is worth knowing before anyone describes it as a requirement.
There is no public commitment from OpenAI, Anthropic, Google or Perplexity that llms.txt is used as a retrieval or ranking input. Some developer documentation platforms and coding tools do consume it, which is where the genuine adoption sits. Treat any claim that a named assistant ranks by it as unverified.
Probably not on its own, and anyone promising otherwise is ahead of the evidence. The observable correlations are confounded, because the sites that publish an llms.txt are also the sites doing clean structure, honest dates and answer-first writing. Publish it for the low cost, not for a projected uplift.
A companion file containing the actual content rather than links to it, concatenated into one long markdown document so a model can ingest an entire documentation set in a single fetch. It is genuinely useful for technical documentation and mostly pointless for a marketing site, where the same content is a few hundred kilobytes of prose nobody wants in a context window.
No, and this is the most common misunderstanding. It grants nothing and forbids nothing. Permission belongs to robots.txt, your terms and, in a more expressive form, a content signals declaration. llms.txt is a map, not a lock.
That defeats the point. A sitemap dumped into markdown is a list of every URL with no judgement applied, which is exactly what a model already gets from crawling. The value, if there is any, comes from a human deciding which twenty pages matter and writing one accurate line about each.
When the structure of the site changes, which for most businesses is two or three times a year. A stale llms.txt pointing at pages that no longer exist is worse than none, because it is an active signal that nobody is maintaining the site.
We audit crawler access, extractability and entity resolution across whole estates and tell you which of it is worth your budget. The markdown file is not the part we charge for.