Is your website ready for AI agents?
Free Agent Readiness Checker
See which of 9 AI crawlers can read your site and which of 17 agent standards it supports, from llms.txt to MCP.
How to check if your website is agent-ready
One scan covers robots.txt, discovery files, .well-known endpoints, HTTP headers and DNS records. No account and no install, for your own site or any other.
Enter a domain in the field above, yours or a competitor's
Run the check: it reads robots.txt, requests each discovery file and endpoint, and looks up DNS records
Fix what fails, starting with crawler access, or copy the fix prompt into your coding agent. Then run the check again
What is agent readiness?
Agent readiness measures how easily AI agents and AI crawlers can find, read and use your website. An agent readiness checker tests whether your robots.txt lets AI crawlers in, whether you publish discovery files like llms.txt and sitemap.xml, and whether you support agent standards such as Markdown for agents, MCP server cards and OAuth discovery.
Two kinds of visitors are involved. AI crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot read pages so ChatGPT, Claude and Perplexity can find and cite them, and assistants fetch pages live when a user asks a question. AI agents go further: they call APIs, sign in on a user's behalf and complete tasks. Both start with the same questions. Am I allowed in? Where is the content? Is there a cleaner format, or an API I can use instead of the HTML?
The standards that answer those questions are new and still changing. This checker, also called an agentic readiness check, probes your domain for each of them, then shows what is missing and how to fix it.
What the agent readiness checker tests
The scan has two parts. First, it reads your robots.txt the way each of 9 AI crawlers would and reports whether that crawler is allowed. Then it checks 23 agent standards: 17 that count toward your score and 6 informational ones that only matter for some businesses. Here is every check, where it lives, who needs it and how much it matters today:
| Check | Where it lives | Who needs it | Impact today |
|---|---|---|---|
| Discoverability | |||
| robots.txt | /robots.txt | Every site | High |
| sitemap.xml | /sitemap.xml or a Sitemap: line in robots.txt | Every site | High |
| llms.txt | /llms.txt | Every site | Medium |
| llms-full.txt | /llms-full.txt | Every site | Medium |
| Link headers | Link header on the homepage response | Sites with an API or app | Emerging |
| DNS-AID | DNS records under _agents.your-domain | Sites with an API or app | Emerging |
| Content accessibility | |||
| Markdown for agents | Homepage requested with Accept: text/markdown | Every site | Medium |
| Bot access control | |||
| AI bot rules | User-agent groups in robots.txt | Every site | High |
| Content Signals | Content-Signal line in robots.txt | Every site | Medium |
| Web Bot Auth | /.well-known/http-message-signatures-directory | Sites that run bots | Informational |
| Protocol discovery | |||
| MCP server card | /.well-known/mcp/server-card.json | Sites with an API or app | Emerging |
| A2A agent card | /.well-known/agent-card.json | Sites with an API or app | Emerging |
| Agent Skills | /.well-known/agent-skills/index.json | Sites with an API or app | Emerging |
| WebMCP | Homepage JavaScript (document.modelContext) | Sites with an API or app | Emerging |
| API catalog | /.well-known/api-catalog | Sites with an API or app | Emerging |
| OAuth discovery | /.well-known/oauth-authorization-server or openid-configuration | Sites with an API or app | Emerging |
| OAuth protected resource | /.well-known/oauth-protected-resource | Sites with an API or app | Emerging |
| auth.md | /auth.md | Sites with an API or app | Emerging |
| Agentic commerce | |||
| x402 | HTTP 402 responses on paid routes | Stores and paid APIs | Informational |
| MPP | Payment metadata in /openapi.json | Stores and paid APIs | Informational |
| UCP | /.well-known/ucp | Stores and paid APIs | Informational |
| ACP | /.well-known/acp.json | Stores and paid APIs | Informational |
| AP2 | Extension declared in the A2A agent card | Stores and paid APIs | Informational |
Impact today is our own rating. High: decides whether AI systems can read your pages at all. Medium: cheap to add and useful to agents that look for it, with no proven effect on AI citations yet. Emerging: a new standard with early adoption. Informational: shown in the report, never scored.
AI crawler access
The checker evaluates your robots.txt rules for GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI), ClaudeBot and Claude-SearchBot (Anthropic), PerplexityBot and Perplexity-User (Perplexity), Google-Extended (Google) and Meta-ExternalAgent (Meta). For each one it shows whether the crawler is allowed and what decided it: a group that names the crawler, the User-agent: * group, or no rule at all. With no robots.txt, every crawler is allowed by default. This AI crawler checker is the part of the tool with the most direct effect on AI visibility today.
Discoverability
robots.txt and sitemap.xml are the basics every crawler relies on, so validate your XML sitemap if it fails. llms.txt is a Markdown file at the root of your site that gives language models a short map of who you are and which pages matter, following the llms.txt proposal. llms-full.txt is not part of that proposal, but some sites publish it next to llms.txt with the full text of their key pages in one file. Link headers (RFC 8288) let an agent find your API catalog or docs from the HTTP response of your homepage, before it parses any HTML. DNS-AID, an early IETF draft, announces agent endpoints through SVCB records under _agents.your-domain. Together, these make the tool an llms.txt checker and a discovery audit in one.
Content accessibility and bot controls
Markdown for agents checks whether your homepage answers a request sent with Accept: text/markdown with an actual Markdown response. AI bot rules checks that robots.txt has a group that applies to AI crawlers, either a named group or User-agent: *. Content Signals checks for a Content-Signal line that states whether your content may be used for search, as AI input or for AI training. Web Bot Auth is informational: it only matters if your company runs its own bots and wants other sites to verify their requests.
Protocol discovery
These checks matter for sites that offer an API, an app or an AI integration. An MCP server card describes a Model Context Protocol server so AI assistants can find and connect to it; BlogSEO's SEO MCP server publishes one. An A2A agent card describes an agent to other agents under the Agent2Agent protocol. An Agent Skills index lists ready-made instruction files that agents can load. WebMCP, a draft from the Web Machine Learning Community Group, lets a page expose its actions as tools to a browser agent. Finally, the API catalog (RFC 9727), OAuth authorization server metadata (RFC 8414 or OpenID Connect discovery), OAuth protected resource metadata (RFC 9728) and auth.md tell agents where your APIs are and how to authenticate.
Agentic commerce (not scored)
x402 and MPP let agents pay for requests over HTTP 402. UCP and ACP let shopping agents discover a store and buy from it (ACP was developed by OpenAI and Stripe). AP2, developed by Google, is a payments extension for A2A agents. They only matter if you sell to agents, so the checker reports them without scoring them.
How to read your agent readiness score
A check that passes, or passes with a warning, counts as supported. A check that could not complete, for example because a firewall blocked the request, counts as not supported, so run the scan again if you see one. The results come as three sub-scores and a total:
- Crawler access. How many of the 9 AI crawlers your robots.txt allows.
- Discoverability. How many of the 6 discovery checks pass: robots.txt, sitemap.xml, llms.txt, llms-full.txt, Link headers and DNS-AID.
- Content accessibility. How many of the 11 content, bot control and protocol checks pass.
- Agent standards. All 17 scored checks combined. Informational checks never count.
Every check weighs the same, so a missing DNS-AID record costs as much as a missing llms.txt. Not every check applies to every site, so read the score as a checklist for your type of site, not as a grade:
- Blogs and content sites. Aim for every crawler you want, robots.txt, sitemap.xml, AI bot rules, Content Signals, llms.txt, llms-full.txt and Markdown for agents. You don't need an MCP server card, OAuth discovery or auth.md, so 7 of 17 can be a complete setup.
- SaaS and API products. Add Link headers, an API catalog, OAuth discovery, OAuth protected resource metadata, auth.md, and an MCP server card if you run an MCP server. Agent Skills, WebMCP and an A2A card come next if they fit your product.
- Online stores. Start with the content site list. Commerce protocols appear in the report when your platform supports them, but they never change your score.
How to make your website agent-ready
Work through these steps in order. The first four take minutes on any site, and every failed check in your report comes with its own fix.
- Decide which AI crawlers get in. Search and answer crawlers decide whether you can appear in AI answers. Training crawlers mostly do not. A common setup lets everyone in and opts out of training only if you object to it. Each crawler follows the group that names it, so the training crawlers below ignore the
*group:# Optional: opt out of AI training User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: meta-externalagent Disallow: / # Everyone else, AI search and answer crawlers included User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Sitemap: https://www.example.com/sitemap.xmlOne exception: Google uses the Google-Extended token for both Gemini training and grounding in Gemini apps. Leave it out of the blocked list if you want Gemini to keep grounding answers in your pages.
- Declare Content Signals. The
Content-Signalline above states your preferences. According to contentsignals.org, search covers a search index with links and short excerpts (not AI summaries), ai-input covers using your pages as real-time input for AI answers, such as grounding, and ai-train covers training or fine-tuning models. A signal you leave out neither grants nor restricts that use. - Reference your sitemap. Keep the
Sitemap:line in robots.txt and validate your XML sitemap so crawlers find every page, including new and deep ones. - Publish llms.txt and llms-full.txt. Start with an H1, a one-line summary in a blockquote and sections of links to the pages that matter:
# Example Co > Example Co makes invoicing software for freelancers and small agencies. ## Product - [Features](https://www.example.com/features): what the app does - [Pricing](https://www.example.com/pricing): plans, limits and billing ## Docs - [Getting started](https://www.example.com/docs/start): set up an account - [API reference](https://www.example.com/docs/api): endpoints and authentication ## Optional - [Blog](https://www.example.com/blog): guides and product newsThen publish llms-full.txt with the full text of those pages. Our own /llms.txt is a working example, and our guide on whether you need an llms.txt file covers what to put in it.
- Serve Markdown when agents ask for it. When a request asks for Markdown, answer with Markdown and say so in the headers:
GET / HTTP/1.1 Host: www.example.com Accept: text/markdown HTTP/1.1 200 OK Content-Type: text/markdown; charset=utf-8 Vary: Accept # Example Co Invoicing software for freelancers and small agencies.On Cloudflare, Markdown for Agents does the conversion at the edge on Pro, Business and Enterprise plans. Elsewhere, add middleware that returns a Markdown version of the page when the Accept header asks for
text/markdown, and sendVary: Acceptso caches keep the two versions apart. - If you have an API, an app or an MCP server, publish the .well-known files that apply (MCP server card, API catalog, OAuth protected resource and authorization server metadata), an auth.md, and a Link header on your homepage that points to them:
Link: </.well-known/api-catalog>; rel="api-catalog" Link: </.well-known/mcp/server-card.json>; rel="service-desc"; type="application/json" Link: </docs>; rel="service-doc"; type="text/html" - Run the check again. Re-run the agent readiness checker after each change. Allow a little time for AI systems to notice: OpenAI says its search results can take about 24 hours to adjust after a robots.txt update.
Where you edit robots.txt and response headers depends on your platform and host. If you are not sure what a site runs on, detect its CMS first.
Which AI crawlers should you allow?
AI crawlers do different jobs. Some collect training data, some build the search index that AI answers cite, and some fetch a page only when a user asks for it. Blocking one has very different consequences from blocking another:
| Crawler | Company | Job | robots.txt token | If you block it |
|---|---|---|---|---|
| GPTBot | OpenAI | AI training | GPTBot | Tells OpenAI not to use your content to train its models. ChatGPT search is not affected. |
| OAI-SearchBot | OpenAI | AI search index | OAI-SearchBot | OpenAI says your pages will not be shown in ChatGPT search answers, apart from navigational links. |
| ChatGPT-User | OpenAI | User-triggered fetch | ChatGPT-User | OpenAI says robots.txt rules may not apply, because a user started the request. |
| ClaudeBot | Anthropic | AI training | ClaudeBot | Tells Anthropic to exclude your future content from its training data. |
| Claude-SearchBot | Anthropic | AI search index | Claude-SearchBot | Anthropic says it may reduce your visibility in Claude's search results. |
| PerplexityBot | Perplexity | AI search index | PerplexityBot | Perplexity recommends allowing it if you want your pages to appear in its search results. |
| Perplexity-User | Perplexity | User-triggered fetch | Perplexity-User | Perplexity says this fetcher generally ignores robots.txt, because a user asked for the page. |
| Google-Extended | AI training | Google-Extended | Opts out of Gemini model training and of grounding in Gemini apps. Google Search inclusion and ranking are not affected. | |
| Meta-ExternalAgent | Meta | AI training | meta-externalagent | Opts out of Meta's crawling for uses such as training AI models and indexing content for its products. |
Sources: the crawler documentation of OpenAI, Anthropic, Perplexity and Google.
- Blocking GPTBot does not remove you from ChatGPT search. OpenAI treats each crawler setting separately: you can allow OAI-SearchBot to appear in ChatGPT search and disallow GPTBot to opt out of training.
- Google-Extended is not a crawler. It is a robots.txt token with no user agent of its own, and Google crawls with its usual user agents. Google says it does not affect inclusion or ranking in Google Search.
- User-triggered fetchers may not follow robots.txt. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a user asked for the page. Anthropic's Claude-User, which this checker does not score, does follow robots.txt: blocking it stops Claude from fetching your pages for user questions.
For most blogs and marketing sites, the sensible default is to allow the search and user-triggered crawlers and to block training crawlers only if you object to training use. Then check your firewall: robots.txt only states rules, and a CDN or bot protection setting can still block a crawler that robots.txt allows. Once they can get in, make sure they find something worth reading: make your content easy for LLMs to crawl.
Does agent readiness affect SEO and AI visibility?
Some of it does today. Most of it is groundwork for agents that are still arriving.
- Crawler access matters now. OpenAI says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers, and Anthropic says blocking Claude-SearchBot may reduce your visibility in its search results. If the search crawlers cannot read your pages, those engines cannot cite them.
- robots.txt and sitemap.xml matter for every crawler, Googlebot included. For a deeper look at crawling and indexing, run a full technical SEO audit.
- llms.txt has no proven effect yet. It is a proposal, and we know of no public evidence that it improves rankings or AI citations. It is cheap and harmless, so add it, but do not expect traffic from it.
- The rest is forward-looking. Markdown for agents, Content Signals and the protocol files help the agents that already support them. We know of no evidence that they move Google rankings.
Agent readiness measures delivery, not the message. Whether AI engines quote you still depends on content worth citing: learn how to get cited by ChatGPT, then track whether ChatGPT, Claude and Gemini mention you.
How blogseo.io scores on its own checker
We build tools for AI search, so we publish the files this checker looks for. Our robots.txt allows every crawler and declares Content Signals, and our homepage sends a Link header that points to our API catalog and MCP server card. These files are live if you want working examples to copy:
- /llms.txt and /llms-full.txt
- /.well-known/mcp/server-card.json, which describes our SEO MCP server
- /.well-known/api-catalog
- /auth.md
We also keep an AI info page with the facts we want AI assistants to get right. Run the checker on blogseo.io to see our current result, including the checks we still fail.
Access SEO tools from any page with the BlogSEO Chrome extension
Get instant SEO insights on the site you're browsing, without leaving the tab. The same free tools, one click away in your toolbar.
- 100% free, no account required
- Works on any website you visit
- Installs in seconds from the Chrome Web Store
Free forever • No signup • Takes 10 seconds
FAQ
- What is an agent readiness checker?
- How do I know if my website is ready for AI agents?
- What is a good agent readiness score?
- Does blocking GPTBot remove my site from ChatGPT search?
- Should I block AI crawlers?
- Does llms.txt actually work? Do I need one?
- What is Markdown for agents?
- What are Content Signals in robots.txt?
- What are MCP server cards and A2A agent cards?
- How is this different from Cloudflare's isitagentready.com?

Free SEO playbook
Rank #1 on Google & ChatGPT in the AI era
Get Rank Like Crazy, the playbook with 10,000+ copies downloaded, plus the keyword research, AI citation templates and crawlability checklist that come with it. Free.