robots.txt is the site-level file that tells crawlers what to access and what to ignore. When written correctly, it helps search engines focus on your important pages and avoid wasting crawl budget on private or irrelevant sections.
User-agent rules explained
Each section starts with one or more User-agent directives. Use "*" to target all crawlers, or a specific bot name to limit rules to a single crawler.
Allow and Disallow
- Disallow blocks a path from being crawled.
- Allow explicitly permits a path under a broader disallow rule.
- Paths should start with / and are case-sensitive on most servers.
- Do not use robots.txt to hide sensitive data; use authentication or noindex instead.
User-agent: * Disallow: /admin Allow: /admin/public Sitemap: https://example.com/sitemap.xml
Add your sitemap
Including a Sitemap directive helps crawlers discover your important URLs. It is especially useful for large sites, new sites, or pages that are not well-linked internally.
Validation matters
A malformed robots.txt can confuse crawlers and cause them to ignore the whole file. Use a generator to ensure the syntax is correct.
ToolOrbit’s Robots.txt Generator lets you build and validate the file with no guesswork.
robots.txt vs meta robots vs auth
| Control | Controls | Removes from results? | Use for |
|---|---|---|---|
| robots.txt Disallow | Crawling (fetching) | No — can still be indexed | Crawl-budget control, heavy assets |
| meta robots noindex | Indexing | Yes | Sheets you do not want listed |
| HTTP auth / login | Access | Yes | Actually private content |
The rule people get backwards: blocking a URL in robots.txt does not keep it out of results — it stops the crawler from reading it, which often makes the snippet *worse*. For pages that must not rank, use noindex or real access control, not a disallow rule.
AI crawlers
Major AI vendors publish dedicated user-agent tokens (GPTBot, ClaudeBot, PerplexityBot, Google-Extended). Many sites add explicit rules for them — often a single Allow for maintainers and a distinct policy for AI crawlers. Decide deliberately what each gets, keep the rules plain and consistent, and remember the same caveat applies: robots.txt shapes crawling, not policy or permissions.
Longest-match wins
Crawlers apply the most specific matching rule, which is how Disallow: /private/ coexists with Allow: /private/public-files/. A more specific path beats a broader one regardless of order, so ordering rules top-to-bottom is organizational, not functional. Watch the wildcard difference: “*” is a legal wildcard, but a trailing “/*” can surprise people used to path globbing — block the exact directories you mean, and re-read your robots.txt after any structural reorganisation.
- Block paths, not whole domains: Disallow: /search? consolidates crawl budget.
- Test a rework with the Robots.txt Generator’s live preview before deploying.
- Add freshness to comments; search engines parse the whitespace, not the prose.