DocsSEO
Sitemaps, robots.txt and AI crawlers
The XML sitemap, robots.txt with rules for AI crawlers, llms.txt for AI assistants, IndexNow and verification codes for Google, Bing and others.
XML sitemap#
WordPress’s sitemap at /wp-sitemap.xml, tuned: only content types and taxonomies that are in search, without pages hidden from search or left out by hand, with each page’s last change. /sitemap_index.xml and /sitemap.xml (the addresses other SEO plugins use) redirect to it, so a sitemap already submitted keeps working. Submit it in Google Search Console and Bing Webmaster Tools.
robots.txt#
WordPress’s robots.txt, plus your own rules and the AI crawlers you block (SEO → Settings → Sitemaps & robots). A robots.txt file on the server replaces WordPress’s: the settings warn you when there is one.
AI crawlers
Choose per crawler, or use a preset: allow all, block AI training but allow AI search, or block all. Crawlers that collect training data put your pages into new AI models; AI search crawlers read pages for answers that link to you.
| Crawler | Company | What it does |
|---|---|---|
GPTBot | OpenAI | Collects pages to train AI models |
OAI-SearchBot | OpenAI | Reads pages for AI search answers that link to you |
ChatGPT-User | OpenAI | Opens a page when someone asks an assistant about it |
ClaudeBot | Anthropic | Collects pages to train AI models |
Claude-SearchBot | Anthropic | Reads pages for AI search answers that link to you |
Claude-User | Anthropic | Opens a page when someone asks an assistant about it |
Google-Extended | Google (Gemini) | Collects pages to train AI models |
Applebot-Extended | Apple | Collects pages to train AI models |
PerplexityBot | Perplexity | Reads pages for AI search answers that link to you |
CCBot | Common Crawl | Collects pages to train AI models |
Bytespider | ByteDance | Collects pages to train AI models |
meta-externalagent | Meta | Collects pages to train AI models |
Amazonbot | Amazon | Reads pages for AI search answers that link to you |
llms.txt#
A plain summary of the site for AI assistants (llmstxt.org) at /llms.txt: the site’s name and tagline, your introduction, contact details, and its main pages by kind with their descriptions — only what’s public and in search. Off until you switch it on.
IndexNow#
Tells Bing, Yandex and other search engines about new and changed pages a minute after they’re published (Google doesn’t use it). It needs a public site on a real domain.
Verification codes#
SEO → Settings → Verification: paste the code — or the whole meta tag — from Google Search Console, Bing Webmaster Tools, Yandex, Baidu or Pinterest. It’s printed on the home page.
On the features pages
Still have a question? Email us at support.