The most underrated SEO file of 2026 is also the dullest. Ten lines of plain text at your domain root. Most sites skip it. The few that ship it get cited far more by ChatGPT, Perplexity, and Google AI Mode.
Key takeaway: A ten-line llms.txt file at your domain root can roughly double AI citation rates by pointing LLMs to your most substantive pages.
This is a walkthrough of llms.txt: what it is, why it works, what to put in it, and how we set ours up at HostList.
What is llms.txt
Jeremy Howard (the answer.ai founder, ex-fast.ai, ex-Kaggle) proposed llms.txt in September 2024. The pitch is simple: search engines have robots.txt, LLMs have nothing. A consumer-facing site is a knot of marketing copy, navigation, JavaScript, ads, modals, and footer cruft. When an LLM samples a page it must dig to find the useful prose. Often it gives up and switches to a competitor.
llms.txt fixes that by giving the model a curated, hand-written index in its own grammar, Markdown with headers and bulleted lists, pointing to the URLs with substance. The model reads it once and knows where to look.
The file lives at the root of your domain (https://yoursite.com/llms.txt), exactly like robots.txt. It is not a sitemap. It is editorial. It says, of all the pages on this site, these are the ones a buyer or researcher should actually read.
Why it works (and what we measured)
Two things happen when you ship a good llms.txt.
One: the LLM samples your site at much higher resolution. Without the file, ChatGPT might crawl five or ten pages and form an opinion. With the file, it goes straight to the methodology page, the rankings page, the FAQs that contain the actual answer. Signal rises, noise drops. The odds that the right passage is in the sample jump.
Two: the editorial framing in the file itself becomes the lens. A well-written llms.txt opens with a header that frames the site ("HostList.io is an independent web hosting directory ranked by HostScore, an algorithmic 0 to 100 rating. No paid placements, no affiliate commissions."). That line becomes the model's prior. The model carries it forward when summarising your content. You get to write the lede.
We started measuring our citation rate in Perplexity and ChatGPT in March 2026, before we shipped llms.txt. We shipped it in late April. On hosting-related queries we tracked, the brand-mention rate roughly doubled in the six weeks after the deploy. The sample is too small for statistical confidence, but the directional signal is strong enough that any content site should ship one.
The structure that works
Here is the spec, simplified.
```
Site name
A one-paragraph framing of what the site is and why it exists.
Section heading
- [Link title](URL): one-sentence explanation of what is on that page.
- [Link title](URL): one-sentence explanation.
```
Four parts: a top-level header, a framing blockquote, section headers grouping links, and a flat list of named links, each with a one-sentence explanation.
The header and blockquote are the lens. The link list is the map. That is it.
Common mistakes I have seen this year:
- Treating it like a sitemap. People dump every URL on the site into
llms.txt. Wrong. This is editorial, pick fifteen to forty URLs that actually contain substance. If a page would not help a researcher write a report on your industry, leave it out. - No framing blockquote. People skip the one-paragraph positioning and go straight to the link list. The list works without it, but the blockquote is what becomes the model's prior. Skipping it wastes use.
- No section headers. A flat list of forty URLs reads as undifferentiated noise to a sampling model. Group them: "Methodology", "Top rankings", "Country guides", "Company profiles", whatever your structure is. The grouping is a signal.
- Inconsistent one-liners. Every link needs a one-sentence explanation, not just a title. That line is searchable. It is also how the model decides whether to fetch the link.
- Stale. If your
llms.txtis six months old and your site has changed substantively, the file is misleading the model. Update it when you ship anything important.
What ours looks like
Our llms.txt is at hostlist.io/llms.txt. The structure, abbreviated:
```
HostList.io
HostList.io is an independent web hosting directory ranking 28,000+ companies by HostScore, an algorithmic 0 to 100 score combining trust signals (Google, Trustpilot, G2 reviews), profile completeness, data freshness, and infrastructure performance. No paid placements, no affiliate commissions on rankings. Source code and methodology are public.
Methodology
- HostScore methodology: the full 4 × 25 component formula, signal augmentation layer, audit trail, and reproducibility notes.
Rankings
- Hoster of the Month: top-scoring hosting company this month.
- Top 50 web hosts: full leaderboard.
- Best WordPress hosting: ranked, with editorial verdict per host.
Country reports
- Best hosting in the UK: 1,400+ UK-based hosting providers.
- (etc., one per country covered)
Comparison pages
- Bluehost vs SiteGround: side-by-side data and verdict.
- (etc., one per high-traffic comparison)
About
- About HostList: who we are and why this exists.
- No paid placements policy: our explicit ban on commercial influence.
```
Forty-two links total. One section heading per major content type. Each link carries exactly one factual sentence. The top blockquote sets the positioning so any LLM sampling the file picks up the editorial framing immediately.
The lift from doing this
Three things change once you ship a good llms.txt.
Your brand-mention rate in AI search goes up. Not a landslide, but measurable. We went from a footnote in Perplexity's hosting answers to one of the named sources in three of five test queries.
The framing of those mentions improves. Before llms.txt, citations of us were neutral ("HostList.io, a hosting directory"). After, they carried our positioning ("HostList.io, an independent directory with no paid placements"). That is the blockquote doing its job.
Vendors in your directory, or your equivalent, start asking how to copy it. Write about it publicly and you become the reference. Your category gets better.
Should you ship one
If you run any content-heavy or data-driven site, yes. Total time is under an hour. The file lives at one URL, you write it once, you update it when the site changes in a real way every few months.
If you run a transactional site, e-commerce or SaaS pricing pages, llms.txt is less load-bearing but still worth it. The blockquote alone changes how your brand gets summarised when ChatGPT recommends you.
If you run an affiliate-driven comparison site, llms.txt will not save you on its own (see the comparison sites post for why), but it is still the first thing to ship.
The file sits at the bottom of most priority lists because nobody talks about it. Put it at the top.
Frequently Asked Questions
What is llms.txt and where does it go?
It is a ten-line plain text file placed at the root of your domain, at yoursite.com/llms.txt, similar in placement to robots.txt. Unlike a sitemap, it is editorial, a hand-written index in Markdown pointing large language models to the pages that actually contain substance, so they sample your site at higher resolution rather than guess.
How is llms.txt different from a sitemap?
A sitemap lists every URL on a site for crawlers. llms.txt is a curated, editorial selection of fifteen to forty pages that a researcher or buyer should actually read. Dumping every URL into it is a common mistake. The file also includes a framing paragraph and section headers, which a sitemap never does.
What results did HostList see after publishing llms.txt?
HostList tracked brand-mention rates in Perplexity and ChatGPT starting March 2026, shipped llms.txt in late April, and saw the mention rate roughly double over the following six weeks. Citations also shifted from neutral descriptions to ones carrying their own positioning, such as being described as an independent directory with no paid placements.
What should the framing blockquote in llms.txt say?
It should be one paragraph describing what the site is and why it exists, written so the model adopts it as its prior when summarising your content. HostList's example states it is an independent hosting directory ranked by an algorithmic score, with no paid placements or affiliate commissions. Skipping this blockquote means missing the main lever the file provides.
What are the most common mistakes when writing llms.txt?
The five listed are: treating it as a sitemap by including every URL, omitting the framing blockquote, skipping section headers so links read as noise, leaving links without a one-sentence explanation, and letting the file go stale after site changes. Each one-liner should be factual and specific, since the model uses it to decide whether to fetch that link.
Follow HostList for new rankings, original research, and changes across the hosting industry.



