What Robots.txt Checker checks
- File availability — response code, redirects, size (Google reads only the first 500 KiB) and encoding.
- Syntax — rules before the first User-agent, lines without a colon, unknown directives, paths without a slash.
- Full blocking — whether the site is blocked from Googlebot through
Disallow: /. - Rendering resources — whether the CSS, JS and images Google needs to "see" the page are blocked.
- Sitemap — whether there is a
Sitemap:line and whether the address is absolute. - URL check — whether the selected bot may crawl a specific page, which rule applied and on which line.
How robots.txt works
The file lives in the site root (https://example.com/robots.txt) and consists of groups: a User-agent line names the bot, and below it come Allow and Disallow rules with paths. A bot picks the group that fits it best, and among the rules the one with the longer path wins; if the length is equal, Allow wins. The characters * (any characters) and $ (end of the address) are supported.
Common robots.txt errors
Does Disallow hide a page from Google
No. Disallow forbids a bot to crawl a page, but does not forbid showing it in search: if there are links to it, the address can remain in the results without a description. To remove a page from the index, it must be left open to crawling and have noindex added. Whether a page has got into Google is shown by the page indexing check.
Limits of the check
We download the file from a public address with one request, so sites that block third-party bots (Cloudflare, etc.) may return an error to us even though Googlebot sees the file normally. Each subdomain and protocol has its own robots.txt, so check the host you need separately. The sitemap file that the Sitemap: line points to can be checked by Sitemap Checker.
How a page opens after redirects is shown by Redirect Checker.
What else StayIndexed can do
This page is a free tool from the StayIndexed service. In your account you can save a list of your sites, refresh the check in one click and see the history of robots.txt changes (also free), while the main job of the service is to track how your pages are doing in Google.
How much it costs
The check is free, both on this page and in your account. Tokens are used to pay for the other services: indexing, speed and backlinks cost 3 tokens per check. 30 tokens are credited at sign-up, no card required.
Frequently asked questions
What is robots.txt?
It is a text file in the site root that tells search bots which sections may be crawled and which may not, and where the sitemap is. It is a recommendation for bots, not protection against access.
Where should robots.txt be located?
Only in the root of the host: https://example.com/robots.txt. Each subdomain and protocol (http and https) has its own file, and a file in a subfolder is not read by bots.
Does Disallow hide a page from Google?
No. Disallow forbids crawling, but the address can still get into search without a description. To remove a page from the results, open it to crawling and add noindex.
What is the maximum size of robots.txt?
Google reads the first 500 KiB of the file and ignores the rest. So long lists of rules are better shortened with wildcard patterns.
How do I check whether a specific URL is blocked?
Enter the site, and in the second field the page address, and choose a bot. We will show whether crawling is allowed or forbidden, which rule applied and on which line of the file.
How much does a robots.txt check cost?
Nothing: the check is free both on this page and in your StayIndexed account. In your account you can save a list of sites, refresh the check in one click and see the history of robots.txt changes.