robots.txt Checker with Sitemap Validation
Enter any page URL and this robots.txt checker tells you whether Googlebot, Bingbot and AI crawlers such as GPTBot or ClaudeBot may fetch it – along with the rule that decides. It also flags syntax problems and checks the sitemap: number of URLs, foreign hosts and lastmod format.
Source: RFC 9309 – Robots Exclusion Protocol. Updated: .
How it is calculated
How crawlers read robots.txt
robots.txt always lives in the root of a host (https://example.com/robots.txt) and applies to exactly that host and protocol. Since 2022 the format is standardized as RFC 9309. The core rules:
- Groups: one or more
User-agentlines followed byAllowandDisallowrules. A crawler obeys only the group with its own name – if there is none, theUser-agent: *group applies. - Longest match wins: with
Disallow: /shopandAllow: /shop/deals, /shop/deals/1 is allowed. On a tie, Allow wins. - Wildcards:
*matches any characters,$marks the end of the URL:Disallow: /*.pdf$blocks all PDF files. - No file (404) means everything may be crawled. If the server answers with 5xx, Google pauses crawling.
Google reads only user-agent, allow, disallow and sitemap, and processes at most 500 KiB. Crawl-delay and Noindex are ignored by Google – the checker points that out.
robots.txt blocks crawling, not indexing
A blocked URL can still show up in search results if other pages link to it – just without a description. To keep a page out of the index reliably, give it noindex (meta tag or X-Robots-Tag) and do not block it in robots.txt, otherwise Google never sees the noindex.
Blocking AI crawlers
Many sites want search engines but not AI training crawlers. That is why the checker also tests GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (use for Google’s AI models) and CCBot (Common Crawl). Example:
User-agent: GPTBot Disallow: / User-agent: Google-Extended Disallow: /
Sitemap validation
If robots.txt lists a sitemap, the first one is checked; otherwise the default location /sitemap.xml is tried. Under the sitemap protocol a file may hold at most 50,000 URLs and 50 MB, all from the same host, and lastmod must use W3C datetime format (e.g. 2026-09-30). Need a new one? Try the sitemap XML generator.
Limitations
Gzip-compressed sitemaps (.xml.gz) are not unpacked, and for sitemap indexes only the index itself is counted. Very large files are analyzed up to 2 MB.
Frequently asked questions
How do I test my robots.txt file?
Enter the address of a page on your site and click “Check robots.txt”. The table shows for Googlebot, Bingbot, AI crawlers and everyone else whether that exact URL is allowed, and which rule decides. Google’s own robots.txt report is in Search Console.
What does “Disallow: /” mean?
The slash stands for every path: the crawlers in that group may not fetch anything from the site. Under User-agent: * it blocks the whole site for everyone – only sensible on test or staging environments.
Where does robots.txt have to be?
In the root of the host, i.e. https://example.com/robots.txt. Every subdomain (such as shop.example.com) and each protocol, http and https, has its own file.
Should the sitemap be listed in robots.txt?
It is optional but handy: a line like Sitemap: https://example.com/sitemap.xml helps every search engine find it. The URL must be absolute, including https:// and the domain.
Why is my page on Google even though robots.txt blocks it?
robots.txt only stops fetching. If other pages link to the URL, Google can still list it without a description. To remove it from the index, use noindex and allow crawling so Google can see it.
Sources and legal basis
- RFC 9309 – Robots Exclusion Protocol
- Google Search Central – How Google interprets the robots.txt specification (500 KiB, supported fields, 5xx handling)
- sitemaps.org – Sitemaps XML format (50,000 URLs, 50 MB, single host, W3C Datetime)
- W3C – Date and Time Formats (W3C Datetime)
- OpenAI – Overview of OpenAI crawlers (GPTBot)
- Google – Overview of Google common crawlers (Google-Extended)
As of:
Related tools
- Sitemap XML generator: build sitemap.xml from your URL listSitemap XML generator without signup: paste URLs, get a valid sitemap.xml per the Sitemaps protocol, check for errors and download it. Runs in your browser.
- SEO Checker: Analyze Any Web Page for FreeAudit any URL: title, meta description, H1, canonical, hreflang, noindex, Open Graph, structured data and alt text – with an SEO score and clear fixes.
- Redirect Checker: Trace the Full Redirect ChainThis redirect checker traces every hop of a URL with its status code (301, 302, 307, 308), target and timing – and flags chains, loops and HTTPS mistakes.
- HTTP Header Checker with Security GradeSee every HTTP response header of a URL and grade its security headers – HSTS, Content-Security-Policy, X-Frame-Options, Referrer-Policy – from A to F.
- HTTP Status CodesLook up HTTP status codes: the meaning of 200, 301, 404, 500 and over 40 more codes explained, with search by code or keyword and a filter by class.
- ASCII table with decimal, hex, octal, binary and Unicode lookupComplete ASCII table (0–127) plus Latin-1 (128–255): decimal, hex, octal, binary, HTML entity, control codes. Look up any Unicode character, UTF-8 included.