Guide

Does a Practice Site Need llms.txt? No Engine Says It Reads One

Summary

No. No answer engine says its crawler reads an llms.txt file. That is what their own documents for site owners show, and Google says outright that Search ignores one. The file is a 2024 proposal by one author for a markdown summary of a site. The crawlers that exist read robots.txt and the ordinary HTML page a person asked about. Ten minutes on a plain facts page, with name, credential, NPI, address, hours, fees, plans and how to book, changes what an assistant can read. A paid llms.txt does not.

By Gale Editorial · Updated 2026-09-15. Every figure cited to a dated source. How we write.

What llms.txt is, and who proposed it

llms.txt is a proposal, published at llmstxt.org by Jeremy Howard on September 3, 2024 and revised since, for a plain markdown file at the root of a website that hands an AI agent a short overview and a curated list of links to the pages that matter 1. It is one author's format. No standards body has adopted it, and the crawler pages the engines publish do not name it. They are read below one by one.

The format is short. The file needs an H1 with the site's name, and that is the only required part. After it comes a blockquote summary, then H2 sections that hold lists of links, and by convention a final section headed Optional for links an agent can skip 1. The proposal places the file beside robots.txt, which governs access, and sitemap.xml, which lists every URL, and says it replaces neither.

The page also carries the author's own adoption claims: that thousands of sites publish one, that documentation platforms generate it automatically, that Chrome's Lighthouse audits for it, and that OpenAI, Anthropic and Gemini publish llms.txt files for their developer documentation 1. Every one of those is a claim about who publishes the file. Publishing a file and reading one from somebody else's site are different acts, and a practice is paying for the second.

Which answer engines say they read one?

None of them, in the documents they publish for site owners. Google's guide to its generative AI features says a site does not need llms.txt files, AI text files, Markdown or special markup, that Search itself does not use them, and that keeping one for some other service neither helps nor harms visibility 2. Google's documentation changelog dates that note to June 15, 2026 3.

OpenAI's crawler page names four robots.txt tokens and what each does. OAI-SearchBot decides whether a site can be shown in ChatGPT search answers, GPTBot collects content for its models and respects robots.txt, and ChatGPT-User fetches a page when a person asks for it 4. The only llms.txt on that page is a link to OpenAI's own documentation index, a table of contents for its developer docs.

Anthropic's help-center page describes three robots. ClaudeBot collects content that may contribute to training, Claude-User fetches pages when a person asks Claude a question, and Claude-SearchBot navigates the web to improve search results; all three honor robots.txt directives and the Crawl-delay extension 5. The page does not mention llms.txt. Perplexity's crawler page has the same shape as OpenAI's: PerplexityBot surfaces and links sites in Perplexity results, the page recommends allowing it in robots.txt, and Perplexity-User fetches a page a user asked about 6. Its one llms.txt is, again, a link to Perplexity's own documentation index.

Bing's webmaster guidelines say one crawling, indexing and ranking foundation serves Bing search, Copilot and its grounding API, and they name four ways a URL becomes discoverable for it: IndexNow, XML sitemaps, crawlable internal links and external links from relevant sites 7.

EngineThe documentWhat it says the crawler readsNames llms.txt as an input?
Google (AI Overviews, AI Mode)Guide to optimizing for generative AI featuresAn indexed page eligible for a snippetNo; Search does not use the file
ChatGPT searchOverview of OpenAI Crawlersrobots.txt; the HTML page a user asked forNo; the page links only OpenAI's own docs index
ClaudeAnthropic help center, crawler articlerobots.txt, including Crawl-delayNo
PerplexityPerplexity Crawlersrobots.txt and published IP rangesNo; the page links only Perplexity's own docs index
Bing and CopilotBing Webmaster GuidelinesIndexNow, XML sitemaps, internal and external linksNot among the routes it lists

That is the published record as of September 2026. It can change, and the plan below watches for the day an engine names the file.

What the assistants read instead

They read the page itself. Every fetcher in the documents above, ChatGPT-User, Claude-User and Perplexity-User, retrieves the ordinary HTML page a person asked about, at the moment they ask 456. Google adds a condition on the search side: to appear in AI Overviews or AI Mode, a page must be indexed and eligible to show with a snippet, and the site must be included in Search generative AI features in Search Console 2.

So the summary an llms.txt would carry has a better home: the page the assistant will fetch anyway. Bing's guidelines describe what makes that page usable for a citation: content that surfaces key information early, defines its entities clearly and can be verified on the page itself 7. Google's guide adds that details from a Business Profile can surface in generative responses 2. Neither document asks for a new file.

Robots.txt decides whether that page reaches an assistant at all, and the crawler allowlist walks through which user agent to allow for each engine. The page's own index and snippet directives decide the rest: Bing says NOINDEX removes a URL from Copilot and grounding, and NOARCHIVE blocks its use in Copilot responses 7. The snippet controls cover the same decisions for Google.

The ten-minute change: a facts page in plain sentences

Write, or fix, one page on the site that states the practice's facts in plain sentences: the practice name, the clinician's name and credential, the NPI, the street address, the hours, the fee or the fee range, the plans accepted, the states served, and how to book, with the date the page was last checked at the bottom. That is the summary an llms.txt would hold, placed on the page the fetchers already read.

Plain means a sentence a fetcher can read without running anything. "The practice is at 12 Main Street, Springfield, and is open Tuesday through Friday, 9 to 4" reads the same to a person and to a robot. A fee that appears only after a click, hours inside a widget that loads later, or an address drawn inside an image may not be there when the page is fetched. The raw-HTML test shows what the fetcher sees. Keep the page in the navigation and in the sitemap, so the crawlers that walk links and read sitemaps find it 7.

The facts page is the about page of the five-page website, and its facts should match the practice's Business Profile, its NPPES record and its directory listings word for word. When the site and the listings disagree, an assistant reading several sources has no way to know which one is current, and the basics of local SEO rest on the same match. The site is the one you control most directly among the five sources an assistant can draw on.

What to ask an agency that quotes llms.txt

Ask one question, in writing: which engine's own documentation names llms.txt as an input its crawler reads? Then read the reply against the five documents in the table above and keep the answer. A reply that points to the proposal page, to a platform that generates the file, or to a Lighthouse audit is naming a publisher of the file, and the question was about a reader.

Google's guide also says no third-party tool has access to its internal ranking or AI systems, and that seeking inauthentic mentions does not help 2. Carry that line into the conversation. The file itself is harmless by Google's own account. What a practice buys with the line item is the vendor's time, and the same time on the facts page changes what every fetcher in the table can read.

If you still want one: a template that costs nothing

A practice that wants the file anyway can write it in five minutes with the facts the about page states, and Google's position is that doing so neither helps nor harms visibility 2. The proposal asks for a markdown file at the site root with one required H1, a blockquote summary, and H2 sections holding link lists, each link written as a page's name and its address 1. Nothing about it can be measured afterwards, so treat it as housekeeping.

``` # Example Practice > Solo practice at 12 Main Street, Springfield. In person and by video. New patients accepted. Facts last checked 2026-09-15.

## Facts - About: the facts page, with credential, NPI, hours, fee, plans and states - Book: the booking page

## Optional - News: the practice's own posts ```

When the facts page changes, tell the engines that accept a ping. IndexNow is one of the four discovery routes Bing lists 7, and the IndexNow setup takes a key file and one request per changed URL. The gale.care profile page gets the same ping from Gale the day its facts change. A ping buys discovery; the engine still decides what to do with the page.

Whether any of this moved anything is a question for the monthly prompt panel: the same questions, asked of the same assistants, once a month, with each answer and any cited page written down. Google's Search Console also carries a Generative AI performance report for its own features 2. Both are records of what an engine did, and neither promises what it will do next month.

Common questions

No. Google's guide to its generative AI features says a site does not need llms.txt files, AI text files, Markdown or special markup, that Search does not use them, and that keeping one for another service neither helps nor harms visibility. The documentation changelog dates the llms.txt note to June 15, 2026. What Google asks for is an indexed page that is eligible to show with a snippet.

By Google's own account, no. The guide says creating and maintaining one for other services neither harms nor helps a site's visibility in Search. None of the other crawler documents describe reading it, so it sits at the root doing nothing. Leave it, or update it so its facts match the about page. What matters is that the about page carries the same facts in plain sentences.

The two are different acts. Their developer-documentation sites link to an llms.txt as a table of contents for their own docs, which is what the proposal was written for. Their crawler pages, on the same sites, describe robots and fetchers that read robots.txt and the HTML page a user asked about, and name no llms.txt as an input. Publishing the file for your docs says nothing about reading it from someone else's site.

The practice name, the clinician's name and credential, the NPI, the street address, the hours, the fee or fee range, the plans accepted, the states served, how to book, and the date the page was last checked. Each in a plain sentence, not inside an image, a widget or a click-to-reveal block. Keep it in the navigation and in the sitemap, and keep every fact identical to the Business Profile and the NPPES record.

View the page's source in the browser and confirm every fact appears as text there, since that is what a fetcher receives. Check that robots.txt does not disallow the crawlers named in the engines' documents. In Search Console, inspect the URL and confirm it is indexed. Then ask the assistants the questions a patient would ask, once a month, and write down what each one answers and cites.

Then the template takes five minutes and the facts already exist on the about page. The plan carries a standing check for exactly that: if an engine publishes crawler documentation naming llms.txt as an input, a note lands in the review pen and the plan changes. Until a document like that exists, the file is housekeeping, and the page the fetchers read today is the page worth the ten minutes.

Run your practice on Gale

The software is free. Gale earns one flat 3.5% all-in per paid transaction — only on transactions that actually pay. No subscription, no setup fee, no network cut.

Start or manage a practice →

References

  1. 1.Jeremy Howard (2024). The /llms.txt file, v2. llmstxt.org. linkThat llms.txt is a September 2024 proposal by one author for a markdown file with a required H1, a blockquote summary, H2 link lists and an Optional section; that it sits beside robots.txt and sitemap.xml; and the author's own adoption claims (thousands of sites, documentation platforms, Lighthouse, OpenAI, Anthropic and Gemini publishing one for their docs).
  2. 2.Google Search Central (2026). Google's Guide to Optimizing for Generative AI Features on Google Search. Google Search Central documentation (developers.google.com). linkGoogle's statement that a site does not need llms.txt, AI text files, Markdown or special markup, that Search does not use them and they neither help nor harm; that a page must be indexed and snippet-eligible and the site included in Search generative AI features; that Business Profile details can surface in AI responses; that no third-party tool has access to Google's systems; and the Search Console Generative AI performance report.
  3. 3.Google Search Central (2026). Latest documentation updates. Google Search Central, What's new (developers.google.com). linkThe date, June 15, 2026, on which Google added the note that llms.txt files are not used by Google Search and neither help nor harm visibility.
  4. 4.OpenAI (2026). Overview of OpenAI Crawlers. OpenAI Developer Platform documentation. linkOpenAI's four robots.txt tokens and what each does: OAI-SearchBot governs appearance in ChatGPT search answers, GPTBot collects content for its models and respects robots.txt, ChatGPT-User fetches a page for a user-initiated action.
  5. 5.Anthropic (2026). Does Anthropic crawl data from the web, and how can site owners block the crawler?. Claude Help Center (Anthropic). linkAnthropic's three robots (ClaudeBot, Claude-User, Claude-SearchBot), what each does, and that all three honor robots.txt directives and Crawl-delay.
  6. 6.Perplexity (2026). Perplexity Crawlers. Perplexity Docs. linkThat PerplexityBot surfaces and links sites in Perplexity results, that Perplexity recommends allowing it in robots.txt with published IP ranges, and that Perplexity-User fetches a page a user asked about.
  7. 7.Microsoft Bing Webmaster Tools (2026). Webmaster Guidelines - Bing Webmaster Tools. Bing Webmaster Tools Help Center. linkThat one crawling, indexing and ranking foundation serves Bing, Copilot and grounding; that a URL is discovered through IndexNow, XML sitemaps, crawlable internal links and external links from relevant sites; that citable content surfaces key information early, defines entities and can be verified on the page; and that NOINDEX removes a URL from Copilot and grounding while NOARCHIVE blocks use in Copilot responses.

https://www.gale.care/for-providers/aeo-llms-txt-worth-it · 7 sources. Competitor details are cited to dated public sources and maintained as they change; figures are estimates, not commitments. Synthetic demonstration.

Findability, by specialty

How practices like yours get found in local search and AI answers — the honest playbook, per specialty.

SEO for private practices · SEO for AI search / answer engines (all verticals)