This translation has not been editorially reviewed yet. The German version is authoritative. Deutsche Fassung →
robots.txt and sitemap: the two files that control access
The check looks at whether robots.txt names a sitemap, and whether it's reachable. robots.txt controls what may be crawled; the sitemap says what exists at all — two different jobs that often get confused.
- The most common mistake is a reference to a sitemap that no longer exists, or that sits on a different host. The check reports both, because a dead reference has the same effect as no reference at all.
- The second most common is an exclusion that blocks too much — a directory with stylesheets and scripts, say. Search engines then can't render the page properly, and rate it accordingly.
- The audit service itself abides by this file: it doesn't crawl addresses it blocks for its product token — and explicitly states the exclusion in the result, instead of presenting it as missing content.




