What the checker looks at
Think of it as a quick AI visibility checker for the technical side of a site: can AI crawlers get in, and is there a guide waiting for them when they do? For any address you enter it checks three things:
- robots.txt: whether each of 14 AI crawlers may fetch the home page, and whether that comes from a group that names the crawler or from the shared
User-agent: *rules. It works as a robots.txt checker built for AI bots. - llms.txt: whether the site already serves one at
/llms.txt, so you can use it as an llms.txt checker as well as a generator. - Indexing: whether the home page carries a noindex in its robots meta tag or X-Robots-Tag header, which keeps it out of search and out of the AI answers built on search.
Then it reads the sitemap and drafts an llms.txt you can edit, copy or download.
What is llms.txt?
llms.txt is a proposed standard published at llmstxt.org in September 2024 by Jeremy Howard of Answer.AI. It is a Markdown file at the root of a website that tells an AI assistant what the site is and which pages are worth reading. A sitemap lists every URL for crawlers; llms.txt is a short, curated reading list written for language models.
The format is simple: a heading with the site's name, a one-line summary in a quote block, then sections, each with a heading and a list of links with a short note on what each page covers. Here is a small example:
# Example Co > Example Co makes invoicing software for small agencies. ## Product - [Features](https://example.com/features): What the app does, page by page - [Integrations](https://example.com/integrations): The apps it connects to ## Pricing - [Pricing](https://example.com/pricing): Plans and what each includes ## Docs - [Getting started](https://example.com/docs/start): Setting up a first invoice
The proposal also reserves a section called Optional for links an assistant can skip when it is short on space. You can see a real file on our own site at gravity.fast/llms.txt.
Does llms.txt help you show up in AI answers?
Be careful with the claims you read about it. Google has said it does not use llms.txt for Search, and its guidance for AI Overviews says no special optimization is needed to appear in them. Elsewhere, support varies: some coding assistants and documentation tools read llms.txt files, while most AI companies have not said whether their crawlers use them.
What matters more is plain. AI crawlers must be allowed to fetch your pages, and those pages must answer questions clearly, in text a machine can read without running JavaScript. An llms.txt takes ten minutes and does no harm, so it is worth having, but fix robots.txt and your key pages first. This tool checks both.
Being readable is only the first step. Whether assistants mention you depends on what the rest of the web says about you, and the gap can be large: our LLM citation statistics found Zapier drew about 28 times the AI mentions of the next automation platform in July 2026.
How robots.txt controls AI crawlers
robots.txt is a plain text file at the root of a site, made of groups. Each group starts with one or more User-agent lines naming crawlers, followed by Allow and Disallow rules for paths. The group for User-agent: * applies to every crawler that is not named anywhere else.
The detail that trips people up: a crawler obeys only the most specific group that names it and ignores all the others. If GPTBot has a group of its own, it skips the rules under * entirely. So this file lets GPTBot into /admin, because its named group holds nothing but Allow: /:
User-agent: * Disallow: /admin User-agent: GPTBot Allow: /
To keep the shared rules, list the AI crawlers in the same group as *, or repeat the Disallow lines in every named group. Our own robots.txt had exactly this mistake until October 2026; it now names every AI crawler in one group with the shared rules.
Within a group the longest matching path wins, and Allow wins a tie. A missing robots.txt means every crawler may fetch everything. And robots.txt is a request, not a lock: well-behaved crawlers follow it, while a firewall or a bot-blocking setting at your host or CDN can turn crawlers away even when robots.txt lets them in.
Training crawlers versus search and assistant crawlers
AI crawlers do different jobs, and the difference decides what you should block.
Search and assistant crawlers
These fetch pages to answer a question now and usually cite the source. Blocking them keeps you out of AI answers.
- OAI-SearchBot: ChatGPT search.
- ChatGPT-User: pages ChatGPT opens when a user asks it to.
- Claude-SearchBot: Claude's search.
- Claude-User: pages Claude opens for a user.
- PerplexityBot: Perplexity's search index.
- Googlebot: Google Search, including AI Overviews.
- Bingbot: Bing, which also feeds Copilot.
Training crawlers
These collect pages to train future models. Blocking them is a business choice: many publishers block them because their writing is what they sell, while many software companies allow them so models learn what their product does. Blocking them does not remove you from AI search.
- GPTBot: OpenAI model training.
- ClaudeBot: Anthropic model training.
- Google-Extended: controls whether Google uses your pages for Gemini training and grounding. It does not affect Google Search.
- Applebot-Extended: controls whether Apple uses your pages to train Apple Intelligence.
- Meta-ExternalAgent: Meta AI training.
- CCBot: Common Crawl, a public web archive that many models are trained on.
- Bytespider: ByteDance training.
Google-Extended and Applebot-Extended are not separate crawlers. They are names you use in robots.txt to opt out of training while Googlebot and Applebot keep crawling as before.
How the generator builds your llms.txt
- It reads robots.txt and the sitemaps it lists, or
/sitemap.xmlif it lists none, following one level of sitemap index. With no sitemap at all, it uses the links on the home page. - It keeps pages on the same domain, sorts them by URL depth so shallow pages such as product, pricing and about come first, and reads the title and meta description of the first 18.
- Jev, a decision model from TypeSafe, puts each page in a section: Product, Pricing, Docs, Company, Resources or Legal. Pages that would not help an assistant, such as logins and tag archives, are left out.
- Your browser assembles the file in the llmstxt.org format: the site name, the summary from the home page's meta description, then one list of links per section.
The result is a starting draft, not a finished file. Read every line, rewrite descriptions that were written for search results, and add the deeper pages that matter most, such as your top docs or the posts customers ask about.
What the checker cannot see
- Only 18 pages. Large sites need deeper sections added by hand.
- Raw HTML only. Sites that build their pages with JavaScript may show empty or generic titles, and pages without a title are left out of the draft.
- Home page rules only. Crawler access is checked for the path
/. A site can allow a crawler there and still block/blog/or/docs/, so read the path rules in your robots.txt too. - Firewalls are invisible to robots.txt. If the home page answers our checker with an error such as 403, AI crawlers may be turned away the same way even when robots.txt allows them, and the draft will be thin.
Settings like these change quietly during redesigns, CMS updates and CDN changes. A monthly check catches it; an agent for SEO monitoring and alerts can run that kind of recurring check for you.
Questions
What is llms.txt?
llms.txt is a proposed standard from llmstxt.org: a Markdown file at the root of a website, at /llms.txt, that gives AI assistants a one-line summary of the site and a short list of the pages worth reading, grouped under headings. It is a guide for AI tools, not a rule they must follow, and it does not replace robots.txt or your sitemap.
Does Google use llms.txt?
No. Google has said it does not use llms.txt for Search, and its guidance for AI Overviews says no special optimization is needed to appear in them. Some AI tools, mostly coding assistants and documentation tools, do read llms.txt, and support varies from tool to tool. Treat it as a cheap extra: letting AI crawlers in through robots.txt and writing clear pages matter more.
Should I block GPTBot?
Only if you do not want OpenAI to train its models on your pages. GPTBot collects pages for training. ChatGPT search uses a different crawler, OAI-SearchBot, and pages ChatGPT opens for a user come through ChatGPT-User. You can block GPTBot and still appear in ChatGPT answers as long as those two stay allowed. Many software companies allow GPTBot so models learn what their product does; many publishers block it because their writing is what they sell.
Where do I put llms.txt?
At the root of your domain, so it loads at https://yourdomain.com/llms.txt as plain text with a 200 status. A subdomain such as docs.yourdomain.com can have its own. Most hosts let you upload a file to the root; some site builders need a plugin or a setting for it.
How often should I update llms.txt?
Whenever the pages it lists change: a new product, new prices, a moved docs section or a renamed page. For most sites that means a quick look once a month and an update with every launch. A short file that is correct beats a long one that points at pages that no longer exist.
Do you store the site I check?
No. The tool fetches the site's robots.txt, llms.txt, home page and sitemap, reads up to 18 pages, and sends their titles and descriptions to Jev to sort them into sections. Gravity does not save the address or the results unless you ask us to set up a weekly check for it. We count how many times the tool runs each day, nothing more.