robots.txt is a plain text file at the root of your domain that tells crawlers which parts of your site they may fetch. It is a convention, not a security mechanism, but every major search engine and most AI crawlers respect it.
Where it lives
The file must be at https://yourdomain.com/robots.txt. A file in a subfolder is ignored. You can only have one per origin.
The syntax
A robots.txt is made of groups. Each group starts with one or more User-agent lines followed by rules.
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /search
Sitemap: https://yourdomain.com/sitemap.xmlUser-agent: *applies to all crawlersDisallowblocks fetching;Allowcarves out exceptions and wins when it is more specificSitemappoints to your XML sitemap
Disallow is not noindex
This is the most common mistake. Disallow stops a crawler fetching a URL. If a page is already indexed, blocking it in robots.txt will not remove it — and Google cannot see a noindex tag on a page it is not allowed to fetch. To remove a page from the index, either return noindex (and allow crawling so it can be seen), or remove the page and return 404/410.
What to block
- Internal search results (
/search?q=) — they create infinite, low-value pages - Cart, checkout and account pages on e-commerce sites
- Filter and sort parameter combinations that explode into duplicates
- Staging or admin paths
What not to block
- CSS and JavaScript — if you block them, search engines cannot render your page correctly
- Your sitemap, or any page you want indexed
- Pages with a
noindextag you actually want removed
Build one in seconds
Use the free robots.txt generator to assemble allow/disallow rules and a sitemap line, then copy it to your domain root. If you run WordPress, see the WordPress version.