Back to Blog

How to Make Your Website Visible to ChatGPT, Claude and Perplexity (2026 Checklist)

If someone asks ChatGPT, Claude or Perplexity a question your website answers, will the assistant find your page, read it and link to it? For a lot of sites the answer is no, and the owner never finds out. The cause is usually one of a handful of fixable problems: a robots.txt rule that blocks the wrong bot, content that only appears after JavaScript runs, or missing basics like a description and structured data. This checklist walks through each one, in the order that matters most. Every step can be checked in a few minutes, and you can test your site with the free [AI Visibility Score](/tools/ai-visibility-score/) as you go. ## How AI assistants actually reach your site There isn't one "AI crawler". Each company runs several bots, and they do different jobs: | Company | Bot | What it does | Obeys robots.txt? | |---|---|---|---| | OpenAI | `OAI-SearchBot` | Builds the index for ChatGPT search answers | Yes | | OpenAI | `ChatGPT-User` | Opens a page when a user's question needs it | Not necessarily (user-initiated) | | OpenAI | `GPTBot` | Collects content that may be used to train models | Yes | | Anthropic | `Claude-SearchBot` | Indexes pages for Claude's search results | Yes | | Anthropic | `Claude-User` | Opens a page when a Claude user's question needs it | Yes | | Anthropic | `ClaudeBot` | Collects content that may be used to train models | Yes | | Perplexity | `PerplexityBot` | Indexes pages to show and link in Perplexity answers | Recommended to allow | | Perplexity | `Perplexity-User` | Opens a page for a user's question | Generally no (user-initiated) | | Google | `Google-Extended` | Controls use of your content for Gemini training and grounding | Yes (doesn't affect Google Search ranking) | The important split is between **answer bots** (search and user-triggered fetchers, which decide whether you show up and get cited) and **training bots** (which collect data for future models). You can block training bots and still be visible in AI answers. Blocking the answer bots is what makes a site invisible. Sources: [OpenAI's crawler overview](https://developers.openai.com/api/docs/bots), [Anthropic's crawler FAQ](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), [Perplexity's bot guide](https://docs.perplexity.ai/guides/bots). ## Step 1: Check robots.txt for accidental blocks Open `https://yoursite.com/robots.txt`. Look for two common mistakes: 1. **A blanket block** such as `User-agent: *` followed by `Disallow: /`. That blocks every well-behaved bot, including the answer bots. 2. **Copy-pasted "block all AI" lists** that include `OAI-SearchBot`, `Claude-SearchBot` or `PerplexityBot`. Many of these lists were written to stop training and accidentally remove sites from AI search. A balanced setup that keeps you visible in AI answers but opts out of training looks like this: ```txt # AI answer bots: allowed (you can be found and cited) User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot Allow: / # Training crawlers: opted out User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: CCBot Disallow: / User-agent: * Allow: / Sitemap: https://yoursite.com/sitemap.xml ``` If you'd rather allow everything, that's fine too. The point is to decide on purpose. You can generate either version with the [AI robots.txt generator](/tools/ai-robots-txt-generator/). OpenAI notes that robots.txt changes can take about 24 hours to show up in ChatGPT search. Also check your CDN or firewall. Some bot-protection settings challenge or block AI user agents before robots.txt is ever read, so a correct robots.txt isn't enough on its own. ## Step 2: Make sure the content is in the HTML Most AI fetchers read the HTML your server sends. They usually don't run JavaScript the way a browser does. If your page is an empty `<div id="app"></div>` that fills in after scripts load, the assistant sees almost nothing. **Quick test:** run `curl -s https://yoursite.com/your-page | head -c 3000`, or use "View source" in your browser (not "Inspect"). If you can't see your main text there, neither can most AI bots. **Fix:** use server-side rendering or static generation for important pages (Next.js, Nuxt, Astro, SvelteKit and most CMSs do this by default), or pre-render the pages that matter most. ## Step 3: Give each page a clear title and description Assistants quote and link pages based on what they can understand quickly. Every important page should have: - a specific `<title>` that says what the page is (not just your brand name), - a `<meta name="description">` that answers "what will I learn or get here?" in one sentence, - a `<link rel="canonical">` pointing at the one URL you want cited, so duplicates (tracking parameters, trailing slashes) don't split the signal. ## Step 4: Add structured data JSON-LD structured data tells machines exactly what a page is: an article, a product with a price, a FAQ, an organization. It removes guesswork about names, dates and prices. A minimal article example: ```html <script type="application/ld+json"> { "@context": "https://schema.org", "@type": "Article", "headline": "How to make your website visible to ChatGPT, Claude and Perplexity", "datePublished": "2026-10-08", "dateModified": "2026-10-08", "author": { "@type": "Organization", "name": "Your Company" } } </script> ``` Use `Product` and `Offer` for things you sell, `FAQPage` for question-and-answer content, and `Organization` on your homepage. ## Step 5: Publish a sitemap and an llms.txt A **sitemap** (`/sitemap.xml`, listed in robots.txt) is still the most reliable way for any crawler to find all your pages. An **llms.txt** is a newer, optional file proposed in 2024: a short Markdown page at `/llms.txt` that lists your most useful pages with one-line descriptions, so an AI system can find the important content without crawling everything. Support varies between AI companies and it isn't a ranking signal anywhere, but it costs little to add and some agents and tools do read it. The format is simple: ```markdown # Your Company > One paragraph: what you do, who it's for, and what's most useful here. ## Docs - [Getting started](https://yoursite.com/docs/start): install and first steps - [Pricing](https://yoursite.com/pricing): plans and limits ## Optional - [Blog](https://yoursite.com/blog) ``` You can create one from your sitemap with the [llms.txt generator](/tools/llms-txt-generator/). ## Step 6: Keep pages current and say so AI assistants try to avoid stale answers. Show a visible "last updated" date on pages that change, keep `dateModified` in your structured data accurate, and actually revise the facts (prices, version numbers, steps) when they change. A page that clearly states when it was checked is easier to trust and cite. ## The 10-minute version 1. Run your homepage and one key page through the [AI Visibility Score](/tools/ai-visibility-score/). 2. Fix anything that blocks `OAI-SearchBot`, `Claude-SearchBot`, `Claude-User`, `ChatGPT-User` or `PerplexityBot`. 3. View source on your key pages and confirm the main text is there. 4. Add a description, canonical link and JSON-LD where they're missing. 5. Make sure `/sitemap.xml` exists and is listed in robots.txt, then add an `/llms.txt`. If you use AI agents yourself, the same checks are available as tools in Stormap's [free MCP server](/developers), so you can ask Claude or Cursor to audit a site for you.