What robots.txt is and why you need it
The robots.txt file sits in the root of a site (for example, https://example.com/robots.txt) and tells search bots which sections may be crawled and which may not. It is not protection and not a way to remove a page from Google: the file only asks bots not to enter, and a page blocked in robots.txt sometimes still gets into the index (without a description). For a page to disappear from search for certain, you need the noindex meta tag.
What directives the file consists of
- User-agent — which bot the rules apply to:
*means all, andGooglebotonly the Google bot. - Disallow — a path that must not be crawled. An empty value means “everything is allowed”, and
/blocks the whole site. - Allow — an exception to Disallow, for example allowing one file in a blocked folder.
- Sitemap — the full address of the sitemap.xml file. There can be several Sitemap lines.
- Crawl-delay — a pause between requests in seconds. Google ignores it, while Bing and Yandex respect it.
What to block and what not to
Common robots.txt errors
The worst mistake is leaving Disallow: / after the site launch: search engines stop crawling everything. People often mix up case and the slash at the start of a path, block folders with styles, forget the sitemap address or put the file not in the root but in a subfolder. After publishing the file, be sure to check it in Robots.txt Checker: it will show errors and tell you whether a specific address is blocked for the right bot.
How to install the ready file
Save the result as robots.txt (exactly this name, in lower case, in UTF-8 encoding) and upload it to the root folder of the site. In WordPress this file can also be created through an SEO plugin. Make sure it opens at /robots.txt and returns code 200. A list of addresses for the sitemap can be put together with Sitemap Checker.
Limitations of the generator
The generator builds a basic file: groups for all bots, for individual bots and Sitemap lines. Check complex cases with masks, several groups for one bot and directives of specific search engines by hand. Every site is special, so insert only the paths that exist on yours.
What else StayIndexed can do
This page is a free tool from the StayIndexed service. In your account you can save a list of sites and watch whether their robots.txt has changed (this is free too), while the service’s main job is to track how your pages do in Google.
How much it costs
The check is free, both on this page and in your account. Tokens are used to pay for the other services: indexing, speed and backlinks cost 3 tokens per check. 30 tokens are credited at sign-up, no card required.
Frequently asked questions
How do I create robots.txt for a site?
Pick a template or enter the paths by hand, add the sitemap address and copy the result. Save it as robots.txt and upload it to the root of the site.
Where should the robots.txt file be?
Only in the root of the site: https://example.com/robots.txt. A subdomain needs its own file, and bots do not read a file in a subfolder.
Can I hide a page from Google through robots.txt?
It is only a request not to crawl the page. A blocked page can still get into the index without a description. To remove it from search, use the noindex meta tag and do not block it in robots.txt, otherwise the bot will not see it.
How do I block a site from AI bots?
Turn on the “Block AI bots” option: the generator will add separate groups for GPTBot, ClaudeBot, CCBot, Google-Extended and others. Well-behaved bots follow these rules, but this gives no guarantee.
What is the Sitemap line for?
It tells bots where the site map is and helps them find new pages faster. There can be several Sitemap lines, and the address must be full.
How much does the robots.txt generator cost?
Nothing: the generator is free both on this page and in your StayIndexed account, and it runs right in your browser.