Free tools · Directory
The AI bot directory.
Every major AI crawler in one place: who operates it, what it feeds, the user agent it publishes, and the robots.txt lines that allow or block it. Grouped by what each bot is actually doing on your site.
Live answer fetchers
Fetches your pages in real time while an AI model answers a question. The strongest signal AI is reading you live.
| Bot | Operator | What it does | robots.txt |
|---|---|---|---|
| ChatGPT-User | OpenAI | Fetches a page in real time when a ChatGPT user’s request references it. | Honoured |
| Claude-User | Anthropic | Fetches a page live when a Claude user’s request references it. | Honoured |
| Perplexity-User | Perplexity | Visits a page live to answer a Perplexity user’s query. | Honoured |
| DuckAssistBot | DuckDuckGo | Fetches pages for DuckDuckGo’s AI-assisted answers. | Honoured |
| MistralAI-User | Mistral | Fetches pages live for Mistral’s Le Chat. | Honoured |
AI search crawlers
Retrieves and verifies pages so an AI model can cite them as sources.
| Bot | Operator | What it does | robots.txt |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Indexes and verifies sources shown in ChatGPT Search results. | Honoured |
| Claude-SearchBot | Anthropic | Retrieves and verifies sources for Claude’s web search. | Honoured |
| PerplexityBot | Perplexity | Indexes pages so Perplexity can cite them as sources. | Honoured |
| Meta-ExternalFetcher | Meta | Fetches links on demand for Meta AI to summarise. | Honoured |
| YouBot | You.com | Indexes pages so You.com can cite them in AI answers. | Honoured |
Search index crawlers
Indexes your pages for a search engine whose results now feed AI answers.
| Bot | Operator | What it does | robots.txt |
|---|---|---|---|
| Googlebot | Indexes your pages for Google Search, which now feeds AI Overviews. | Honoured | |
| Bingbot | Microsoft | Indexes pages for Bing and Microsoft Copilot. | Honoured |
| Applebot | Apple | Indexes pages for Siri and Spotlight Suggestions. | Honoured |
| DuckDuckBot | DuckDuckGo | Indexes pages for DuckDuckGo search. | Honoured |
Training crawlers
Crawls public content to train future AI models. Blocking it is a policy choice, not a visibility fix.
| Bot | Operator | What it does | robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Crawls public web content to train future OpenAI models. | Honoured |
| ClaudeBot | Anthropic | Crawls public content to train Anthropic’s Claude models. | Honoured |
| Google-Extended | Controls whether your content trains Gemini and Vertex AI. A robots.txt token, not a separate crawler. | Honoured | |
| Meta-ExternalAgent | Meta | Crawls content to train Meta’s Llama models. | Honoured |
| Applebot-Extended | Apple | Controls whether your content trains Apple Intelligence. A robots.txt token. | Honoured |
| Amazonbot | Amazon | Crawls content for Amazon’s AI (Alexa, Rufus) and search. | Honoured |
| Bytespider | ByteDance | Crawls content to train ByteDance / TikTok AI models. Historically aggressive; verify it honours robots.txt. | Verify |
| CCBot | Common Crawl | Builds the open Common Crawl dataset used to train many LLMs. | Honoured |
| Cohere-AI | Cohere | Crawls content for Cohere’s language models. | Honoured |
On-demand fetchers
Fetches specific pages on demand for its operator.
| Bot | Operator | What it does | robots.txt |
|---|---|---|---|
| Anthropic-AI | Anthropic | Legacy Anthropic fetcher user-agent. | Honoured |
| GoogleOther | Google’s general-purpose fetcher for internal R&D and one-off crawls. | Honoured |
Which of these can actually read your site?
The free crawler access checker reads your robots.txt and fetches a page as each retrieval crawler, so you can see the policy and the reality side by side.