Guide · 2 min read ·
What robots.txt does, and how to write one that does not hurt your SEO
A robots.txt file is a short text file that sits at the root of your website and gives crawlers instructions about where they may go. Getting it right takes five minutes. Getting it wrong can quietly keep your whole site out of search. This guide explains the few lines you need to know, what to block, what never to block and the mistakes that trip up even experienced teams.
What robots.txt is
Crawlers are programs that visit web pages to read them. Before a well behaved crawler explores your site, it asks for a file called robots.txt and follows the rules in it. The rules say which crawlers they apply to and which paths those crawlers should stay out of.
The file lives only at one address: yourdomain.com/robots.txt. Put it anywhere else, such as inside a folder, and crawlers will not find it.
The lines you need to know
Four instructions cover almost every case.
- User-agent: names the crawler a group of rules applies to. A star means all crawlers.
- Disallow: a path the crawler should not visit. A single slash means the whole site.
- Allow: an exception inside a blocked area, such as one public folder in a private one.
- Sitemap: the full address of your sitemap, so crawlers can find all your pages.
What is worth blocking
Block areas that are no use in search results and waste a crawler's time: admin screens, checkout and account pages, internal search results, and staging copies of your site. On a typical web app that means a short list such as /admin, /api and /checkout.
The generator has presets for common setups, and you can edit each list by hand. Paths are matched from the start of the address, so blocking /admin also blocks /admin/users.
What never to block
Do not block pages you want people to find, and do not block the CSS and script files that make your pages display. Search engines need those files to see your pages the way a visitor does, and blocking them can hurt how your pages are understood.
The most damaging mistake is leaving Disallow: / in place after moving from a staging site to the live one. It tells every crawler to stay away from the whole site, and pages can drop out of search within days. Always open your live robots.txt after a launch.
Robots.txt is not security
The file is public, and anyone can read it, including people looking for your private areas. Never use it to hide anything sensitive, and use proper login protection instead.
It also does not remove pages from search. A blocked page can still appear in results if other sites link to it, just without a description. To keep a page out of results, add a noindex tag and leave the page open to crawlers so they can see that tag.
Blocking AI crawlers
Several companies run crawlers that collect pages to train AI systems, and each publishes a name for its crawler. The generator lists the common ones, and you can block any of them with a single tick. This is your choice and it has trade offs, so decide based on how you feel about your content being used that way.
Blocking them does not affect your normal search listing, which is built by separate crawlers.
Using the generator
Pick a preset, adjust the paths, add your sitemap address and copy or download the result. The tool warns about risky settings, such as blocking everything or a sitemap without a full address. Upload the file to your site root, then check it loads. For the rest of your technical basics, run your page through the launch page grader.
