The short answer
To check whether your website is ready for LLMs and AI search, confirm that AI crawlers are allowed in and get a 200 response, that your important text is in the HTML your server returns, and that your sitemap, llms.txt and structured data are valid. Then ask the questions your buyers ask in ChatGPT, Perplexity and Google, and see whether you are named. The free LLM Scan scanner runs the technical part of this checklist in about a minute, without an account.
The checklist at a glance
| Step | What to check | Pass looks like | Scan weight |
|---|---|---|---|
| 1. Reach | Homepage status, content type, canonical | HTTP 200, HTML, canonical matches the URL | 20 |
| 2. Allow | robots.txt rules per AI crawler | No AI crawler blocked from the site or public sections | 15 |
| 3. Guide | /llms.txt | Text file, over 200 characters, a heading, absolute links | 15 |
| 4. Markdown | Accept: text/markdown on your homepage | A real Markdown version, not the same HTML | 15 |
| 5. Discover | sitemap.xml | Valid XML that lists your URLs | 10 |
| 6. Read | Semantic HTML without JavaScript | One H1, title and description, landmarks, 200+ words | 10 |
| 7. Describe | JSON-LD structured data | Valid Organization or WebSite markup | 10 |
| 8. Declare | AI usage preference | A Content-Signal line, header or meta tag | 5 |
| 9. Measure | AI answers for buyer questions | You are named, and your pages are cited | Not scored |
The weights are the ones LLM Scan uses for its 100-point score. 85 or more is rated AI-Ready, 70 or more AI-Readable, 40 or more Needs Work.
Step 1: Make sure crawlers get a 200
Request your homepage and important pages and check the status code. Anything other than a plain 200 is a problem for a crawler: 403s from a firewall, 5xx errors, or a redirect to a URL that is not the canonical one. Add a canonical link tag that matches the final URL. LLM Scan fails crawlability when the homepage does not return 200 or when robots.txt blocks it, and warns when the canonical is missing or points elsewhere.
Step 2: Allow the AI crawlers you want
Open yoursite.com/robots.txt and read the group for each AI crawler. A crawler follows the group that names it, and only falls back to User-agent: * when no group names it. The scan checks GPTBot, ChatGPT-User, ClaudeBot, Claude-Web, PerplexityBot and Google-Extended. It fails when any of them is blocked from the whole site and warns when one is blocked from public sections such as /blog, /docs, /pricing or /solutions. Blocking /admin, /api or /checkout is fine.
Two checks the scan does not do for you: OpenAI's search crawler OAI-SearchBot, which decides whether you can appear in ChatGPT search, and your CDN or firewall bot settings. See how to check if GPTBot and ClaudeBot can access your website for the firewall and log checks.
Step 3: Publish a useful llms.txt
llms.txt is a short Markdown guide to your site at /llms.txt. It is an emerging convention, not a standard every assistant reads, but it is cheap to publish. The scan passes it when the file is served as text, has more than 200 characters, at least one heading and at least one absolute URL. Link your product, pricing, docs and policy pages, and keep it current. The llms.txt checker runs the same test.
Step 4: Offer Markdown
Some AI agents ask for Markdown with the Accept: text/markdown header. The scan sends that request to your homepage and passes when the response is text/markdown and differs from the HTML. If your framework cannot do content negotiation yet, this is the step to leave for later: it is worth 15 points, but crawlability and robots.txt matter more.
Step 5: Keep a clean sitemap
Serve /sitemap.xml, or list your sitemap with a Sitemap: line in robots.txt. It must be well-formed XML with a urlset or sitemapindex root and at least one URL. List canonical public URLs only, not redirects or private pages.
Step 6: Put the text in the HTML
Most AI crawlers do not run JavaScript, so view your page source, not the rendered page. LLM Scan reads the raw HTML and checks seven things: a title of 10 to 70 characters, a meta description of 50 to 160, exactly one H1, no skipped heading levels, main, article, nav and footer landmarks, at least 200 words, and few generic links such as "click here" or "read more". If your key copy only appears after JavaScript runs, render it on the server.
Step 7: Describe your business with structured data
Add JSON-LD that matches what the page shows. The scan passes with valid JSON-LD that includes Organization or WebSite, and warns when the JSON-LD is broken or has neither type. Article, FAQPage, SoftwareApplication and Product are useful next, on the pages they describe.
Step 8: Declare how AI may use your content
A content signal tells AI systems what you allow, for example search and answers but not training. The scan passes when it finds one in a Content-Signal line in robots.txt, a header, or a meta tag. It is worth 5 points and is a declaration, not an access control.
Step 9: Check whether AI actually mentions you
A perfect technical score means AI can read your site. It does not mean ChatGPT or Perplexity recommend you. Write down 10 to 20 questions your buyers ask, such as "best [category] for [audience]", run them in the assistants your buyers use, and note who is named and which pages are cited. Repeat on a fixed schedule, because answers change. LLM Scan's paid plans do this every week and send a report with your mention rate, share of voice against competitors, cited sources, Google rank and AI Overview citations, and 3 fixes. See how to measure AI brand mentions and citations.
What the scan does not cover
The free scan analyses the page you submit, usually the homepage, plus robots.txt, llms.txt and the sitemap. It does not crawl your whole site, does not run JavaScript, and does not check noindex tags or response time. Run it on your pricing and docs pages too, and use a site-wide SEO crawler for the rest. For a comparison of tools, see the best AI SEO scanner.