SEO

How to Create robots.txt: Allow, Disallow and Sitemap Rules

ToolOrbit SEO Team 2 min readUpdated
How to Create robots.txt: Allow, Disallow and Sitemap Rules

robots.txt is the site-level file that tells crawlers what to access and what to ignore. When written correctly, it helps search engines focus on your important pages and avoid wasting crawl budget on private or irrelevant sections.

User-agent rules explained

Each section starts with one or more User-agent directives. Use "*" to target all crawlers, or a specific bot name to limit rules to a single crawler.

Allow and Disallow

  • Disallow blocks a path from being crawled.
  • Allow explicitly permits a path under a broader disallow rule.
  • Paths should start with / and are case-sensitive on most servers.
  • Do not use robots.txt to hide sensitive data; use authentication or noindex instead.
User-agent: *
Disallow: /admin
Allow: /admin/public
Sitemap: https://example.com/sitemap.xml
Remember: robots.txt is about crawling, not indexing. A blocked page can still appear in search results if other pages link to it.

Add your sitemap

Including a Sitemap directive helps crawlers discover your important URLs. It is especially useful for large sites, new sites, or pages that are not well-linked internally.

Validation matters

A malformed robots.txt can confuse crawlers and cause them to ignore the whole file. Use a generator to ensure the syntax is correct.

ToolOrbit’s Robots.txt Generator lets you build and validate the file with no guesswork.

robots.txt vs meta robots vs auth

The three controls
ControlControlsRemoves from results?Use for
robots.txt DisallowCrawling (fetching)No — can still be indexedCrawl-budget control, heavy assets
meta robots noindexIndexingYesSheets you do not want listed
HTTP auth / loginAccessYesActually private content

The rule people get backwards: blocking a URL in robots.txt does not keep it out of results — it stops the crawler from reading it, which often makes the snippet *worse*. For pages that must not rank, use noindex or real access control, not a disallow rule.

AI crawlers

Major AI vendors publish dedicated user-agent tokens (GPTBot, ClaudeBot, PerplexityBot, Google-Extended). Many sites add explicit rules for them — often a single Allow for maintainers and a distinct policy for AI crawlers. Decide deliberately what each gets, keep the rules plain and consistent, and remember the same caveat applies: robots.txt shapes crawling, not policy or permissions.

Longest-match wins

Crawlers apply the most specific matching rule, which is how Disallow: /private/ coexists with Allow: /private/public-files/. A more specific path beats a broader one regardless of order, so ordering rules top-to-bottom is organizational, not functional. Watch the wildcard difference: “*” is a legal wildcard, but a trailing “/*” can surprise people used to path globbing — block the exact directories you mean, and re-read your robots.txt after any structural reorganisation.

  • Block paths, not whole domains: Disallow: /search? consolidates crawl budget.
  • Test a rework with the Robots.txt Generator’s live preview before deploying.
  • Add freshness to comments; search engines parse the whitespace, not the prose.

Tools mentioned

More reading

View all guides
How to Write Title Tags and Meta Descriptions That Get Clicks
SEO

How to Write Title Tags and Meta Descriptions That Get Clicks

How Google builds a search snippet, the length limits to respect, and how to write titles and descriptions that earn the click.

2 min readUpdated
What Is an XML Sitemap? Best Practices and Limits
SEO

What Is an XML Sitemap? Best Practices and Limits

What an XML sitemap does for search engines, which URLs belong in it, and when one sitemap file is not enough.

2 min readUpdated
10 Common robots.txt Rules Explained (with Examples)
SEO

10 Common robots.txt Rules Explained (with Examples)

Master robots.txt syntax: understand User-agent, Disallow, Allow, Crawl-delay, and Sitemap directives with 10 real-world examples.

2 min readUpdated
How to Write SEO Meta Tags: Title, Description and Open Graph
SEO

How to Write SEO Meta Tags: Title, Description and Open Graph

Learn how to create effective title tags, meta descriptions, Open Graph and Twitter tags that improve search visibility and click-through rates.

2 min readUpdated
Keyword Density Checklist: Analyse and Improve Your SEO Copy
SEO

Keyword Density Checklist: Analyse and Improve Your SEO Copy

Understand how to measure keyword density, spot over-optimisation, and use phrase frequency to make content more readable and search-friendly.

2 min readUpdated
UTM Parameters Explained: How to Track Campaigns in GA4
SEO

UTM Parameters Explained: How to Track Campaigns in GA4

What utm_source, utm_medium and utm_campaign actually mean, the naming conventions that keep reports clean, and the mistakes that ruin attribution.

2 min readUpdated
Technical SEO Basics Every Site Needs in 2026
SEO

Technical SEO Basics Every Site Needs in 2026

Meta tags, structured data, robots.txt and sitemaps — the technical foundations that help search engines understand your site.

2 min readUpdated
UTM Campaign Builder: Track Links With Confidence
SEO

UTM Campaign Builder: Track Links With Confidence

Build correctly encoded UTM tracking URLs for your campaigns with ToolOrbit's free UTM Campaign Builder, all in your browser.

2 min readUpdated