XML Sitemap Generator
Build sitemap.xml from your URL list, with lastmod, changefreq and priority.
Build a robots.txt file with the rules you actually need.
Blocking a page here stops crawlers requesting it, but does not remove it from search results. To keep a page out of an index, allow the crawl and use a noindex meta tag instead.
Robots.txt is a plain text file at the root of a domain that tells crawlers which parts of the site they should not request. The tool builds one from rules you pick, with a preview of the finished file and a check for the mistakes that make a robots.txt do the opposite of what its author intended.
Those mistakes are specific and common. Disallow with a single slash blocks the entire site, and one stray character is all it takes. Rules are matched by prefix, so disallowing /admin also blocks /administrator-notes. Blank lines inside a group end that group. And the file is only read at the root of the domain, so putting it in a subfolder achieves nothing.
The most important thing to understand is what robots.txt does not do. It stops well-behaved crawlers requesting a page; it does not stop that page appearing in search results if other sites link to it, and it does not keep anything private. A blocked page can still be indexed with no description. To keep a page out of an index, allow the crawl and use a noindex tag; to keep it private, use a login.
The robots.txt generator is used by writers, developers, students, marketers and anyone else who needs the job done once without installing software. Common cases include:
No, and this is the most important misunderstanding about it. It asks crawlers not to request the page. If other sites link to that page, it can still be indexed, usually showing as a bare URL with no description. To keep a page out of an index, let it be crawled and add a noindex meta tag.
Not at all. It is a public file that anyone can read, and it effectively advertises the paths you would rather people did not visit. Anything genuinely private needs authentication.
At the root of the domain, reachable at /robots.txt. It is only honoured there, and rules apply only to the host and protocol it was served from.
The major search engines do. Scrapers and malicious bots generally ignore it, because nothing enforces compliance. Crawl-delay in particular is honoured by some crawlers and ignored by others.
By prefix. Disallow: /admin blocks any path starting with those characters, which includes /administrator. The dollar sign anchors the end of a path and the asterisk is a wildcard, both of which the major crawlers support.
Crawlers assume everything is allowed, which is usually fine for a normal site. A missing robots.txt is not an error.
If the robots.txt generator is not quite what you need, these other free tools solve closely related problems.
Build sitemap.xml from your URL list, with lastmod, changefreq and priority.
Generate canonical link tags, with hreflang alternates and common-mistake checks.
Generate .htaccess rules for redirects, caching, compression and security headers.
Generate redirect rules for Apache, nginx, Netlify, Vercel and Cloudflare from a mapping.
Build security header configuration for Apache, nginx, Netlify or Vercel.