Robots.txt Generator
Updated October 2026Create and validate robots.txt files to control search engine crawlers. Manage your crawl budget and protect sensitive areas. All processing happens locally in your browser.
Advertisement
Advertisement
Bot Selection
The rules below apply to the bots you tick. Untick all of them to write one rule set for every crawler (User-agent: *).
Disallow Paths
Paths that should not be crawled (e.g., /admin/, /private/)
Allow Paths
Exceptions inside a disallowed folder that should still be crawled (e.g., /admin/public/)
Optional: full URL to your XML sitemap, including https://
Time delay between requests (0 = no delay). Bing and Yandex respect it; Googlebot ignores Crawl-delay.
Generated robots.txt
User-agent: Googlebot
User-agent: Bingbot
Disallow:
Validation
- Note: These rules apply only to Googlebot, Bingbot. Other crawlers are not restricted. Untick every bot to write one rule set for all crawlers (User-agent: *).
Upload this file to your website's root directory (e.g., https://example.com/robots.txt).
What is a Robots.txt file and why does your site need one?
A robots.txt file is a simple text file that sits in your website's root directory and communicates with search engine crawlers (also called "robots" or "bots"). Think of it as a polite set of instructions you give to search engines about which parts of your website they should and shouldn't crawl. While it's not a security measure (anyone can still access blocked URLs if they know them), robots.txt is an essential tool for managing how search engines interact with your site.
Every major search engine—Google, Bing, Yandex, DuckDuckGo, and others—respects the robots.txt protocol. When a crawler visits your site, the first thing it does is check for a robots.txt file at your root URL (e.g., https://example.com/robots.txt). If the file exists, the crawler reads it and follows the directives you've specified. If no robots.txt file exists, the crawler assumes it can access everything on your site, which may not be what you want.
Your site needs a robots.txt file for several critical reasons. First, it helps you manage your crawl budget—the limited number of pages search engines will crawl on your site per day. For large websites with thousands or millions of pages, you don't want search engines wasting time crawling duplicate content, admin panels, or private areas. By blocking these areas, you ensure search engines focus their crawling efforts on your important, public-facing content that actually matters for SEO.
Second, robots.txt protects sensitive areas from being indexed. While it doesn't hide pages from the public (anyone with the URL can still access them), it prevents search engines from discovering and indexing these pages through crawling. This is particularly important for admin panels, staging environments, private user areas, or internal documentation that you don't want appearing in search results.
Third, robots.txt can help you avoid duplicate content issues. If you have multiple versions of the same content (like print-friendly pages, mobile versions, or URL parameters), you can use robots.txt to block the duplicates while allowing the canonical version. This prevents search engines from indexing duplicate content, which can dilute your SEO rankings and confuse search algorithms about which version is the "real" one.
Finally, a well-configured robots.txt file demonstrates technical SEO competence to search engines. It shows that you understand how search engines work and that you're actively managing your site's crawlability. While robots.txt itself isn't a ranking factor, the proper management of crawl budget and duplicate content that it enables can indirectly improve your SEO performance.
The "Crawl Budget" – How to stop wasting Google's time on useless pages
Crawl budget is one of the most misunderstood concepts in SEO, yet it's crucial for large websites. Simply put, crawl budget is the number of pages Google will crawl on your site per day. It's not unlimited—Google has finite resources and can't crawl every page on every website infinitely. For most sites, this isn't a problem. But for large e-commerce sites, news sites, or content-heavy platforms with hundreds of thousands of pages, managing crawl budget becomes critical.
Google determines your crawl budget based on several factors: your site's size, how often you update content, your site's crawl rate limit (how fast Google can crawl without overwhelming your server), and your site's overall health and popularity. Sites with frequent updates and high-quality content typically get a larger crawl budget. Sites with crawl errors, slow response times, or low-quality content may see their crawl budget reduced.
The problem is that many websites waste their crawl budget on pages that don't matter for SEO. Think about it: do you want Google spending time crawling your admin panel, your staging environment, your duplicate product pages, or your internal search result pages? Of course not. You want Google focusing on your important product pages, blog posts, category pages, and other content that drives organic traffic and conversions.
This is where robots.txt becomes invaluable. By blocking unnecessary pages, you're essentially telling Google: "Don't waste your time here—focus on the good stuff." For example, if you have an e-commerce site with 100,000 product pages but also 50,000 internal search result pages, duplicate category pages, and admin areas, you're essentially asking Google to crawl 150,000+ pages. But if you block the unnecessary 50,000 pages with robots.txt, Google can focus its entire crawl budget on your actual product pages, leading to better indexing of your important content.
Common pages that waste crawl budget include: internal search result pages (like ?search=query), filter and sort URLs (like ?sort=price&filter=category), duplicate content versions (print-friendly pages, mobile versions), admin and staging areas, user account pages, shopping cart and checkout pages, and pagination URLs beyond the first few pages. By blocking these in robots.txt, you ensure Google spends its limited crawl budget on pages that actually matter for your SEO and business goals.
The impact of proper crawl budget management can be significant. A large e-commerce site that properly blocks unnecessary pages might see 30-50% better indexing of their important product pages. This means more products appearing in search results, better rankings for key pages, and ultimately more organic traffic and revenue. The key is identifying which pages matter for SEO and which are just wasting Google's time—then using robots.txt to guide Google toward the valuable content.
However, it's important to strike a balance. You don't want to be too aggressive with blocking, as you might accidentally block important pages. Always test your robots.txt file using Google Search Console's robots.txt Tester before deploying it live. And remember: robots.txt is just one tool for managing crawl budget. You should also use proper internal linking, XML sitemaps, and canonical tags to help Google understand which pages are most important.
Common Mistakes: Why blocking your CSS/JS files can kill your rankings
One of the most devastating SEO mistakes you can make is blocking CSS (Cascading Style Sheets) and JavaScript files in your robots.txt. This mistake is surprisingly common, often made by well-intentioned webmasters who think they're optimizing their crawl budget by blocking "non-content" files. But the reality is that blocking CSS and JS can completely break how Google sees and understands your website, leading to catastrophic ranking drops.
Here's why: Google doesn't just read your HTML anymore. Modern Google uses a rendering engine (similar to a web browser) to actually render your pages and see them as users would. This process, called "rendering," requires Google to download and execute your CSS and JavaScript files. If these files are blocked in robots.txt, Google can't render your pages properly, which means it can't see your content, understand your page structure, or determine if your site provides a good user experience.
When Google can't access your CSS files, your pages appear as unstyled, broken layouts. Google can't see your visual hierarchy, your navigation structure, or how your content is organized. When Google can't access your JavaScript files, any content loaded dynamically via JavaScript won't be visible to Google. This includes content loaded through AJAX, React components, Vue.js applications, or any other JavaScript-based content rendering. Google might see a blank page or a page with only the initial HTML, missing all the dynamically loaded content.
The consequences are severe. If Google can't see your content because CSS/JS is blocked, your pages may not be indexed at all, or they may be indexed with incomplete or missing content. Your rankings will drop because Google can't determine what your pages are about or whether they're useful to searchers. You might see errors in Google Search Console about "blocked resources" or "render-blocking resources," and your Core Web Vitals scores may suffer because Google can't properly measure page performance.
This is especially critical for modern websites built with JavaScript frameworks like React, Vue, Angular, or Next.js. These sites rely heavily on JavaScript to render content. If you block JavaScript in robots.txt, Google literally cannot see most of your website's content. Your entire site might appear as a blank page or a loading spinner to Google's crawler, even though it looks perfect to human visitors.
The fix is simple: never block CSS or JavaScript files in robots.txt. Always allow these file types. A good robots.txt file should block things like /admin/, /private/, /internal/, or specific duplicate content paths—but it should never block /css/, /js/, /assets/, /static/, or any directory containing stylesheets or scripts. If you're unsure whether a path contains CSS/JS, err on the side of allowing it. It's better to let Google crawl a few extra files than to break your entire site's visibility in search results.
If you've already made this mistake, fix it immediately. Remove any Disallow directives that block CSS or JavaScript files, upload the corrected robots.txt file, and then use Google Search Console to request re-crawling of your important pages. It may take several weeks for Google to re-render and re-index your pages, but the sooner you fix it, the sooner your rankings can recover. Remember: robots.txt is for managing crawl budget on content pages, not for blocking technical resources that Google needs to understand your site.
How to Use
- 1
Select which search engine bots you want to configure (Googlebot, Bingbot, Yandex, DuckDuckBot), or untick them all to write rules for every crawler
- 2
Add Disallow paths for directories you want to block from crawling (e.g., /admin/, /private/)
- 3
Optionally add Allow paths to override Disallow rules for specific directories
- 4
Enter your sitemap URL if you have one
- 5
Set crawl delay if you want to limit how fast bots crawl your site
- 6
Check the validation notes under the output, then copy or download the generated robots.txt file and upload it to your website's root directory
Why This Tool Matters for SEO
A properly configured robots.txt file is essential for managing your website's crawl budget and protecting sensitive areas from search engine indexing. Search engines have a limited crawl budget - the number of pages they will crawl on your site in a given period. Robots.txt doesn't hide pages from the public - it only tells crawlers not to crawl them. For true privacy, you need authentication or noindex meta tags. Common mistakes include blocking CSS/JS files (which can prevent Google from understanding your site) or accidentally blocking important pages, which can significantly hurt your SEO rankings.
Frequently Asked Questions
Can robots.txt hide my pages from the public?
No, robots.txt does NOT hide pages from the public. It only tells search engine crawlers not to index those pages. Anyone who knows the URL can still access the page directly. If you need to hide content from the public, use authentication, password protection, or add a noindex meta tag to prevent indexing. Robots.txt is purely a directive for search engine crawlers, not a security measure.
What happens if I block CSS and JavaScript files?
Blocking CSS and JavaScript files can severely hurt your SEO rankings. Google needs to see these files to properly render and understand your website. If you block them, Google may not be able to see your content correctly, leading to lower rankings or your site appearing broken in search results. Always allow CSS and JS files in your robots.txt.
How do I test my robots.txt file?
You can test your robots.txt file using Google Search Console's robots.txt Tester tool. Simply upload your file to your website's root directory (e.g., https://example.com/robots.txt), then use the tester to see if Google can access specific URLs. This helps you verify that your directives are working correctly.
What is crawl budget and why does it matter?
Crawl budget is the number of pages Google will crawl on your site per day. It's limited, so you want to ensure Google focuses on your important pages. By blocking unnecessary pages (duplicates, admin areas, private content) in robots.txt, you help Google spend its crawl budget on pages that matter for SEO. This is especially important for large websites with thousands of pages.
Should I use robots.txt or meta noindex?
Use robots.txt to block entire directories or sections of your site from being crawled. Use meta noindex tags for individual pages you want to prevent from appearing in search results but still allow crawling. Robots.txt is more efficient for blocking large sections, while noindex gives you page-level control. You can use both together for maximum control.
Where do I upload the robots.txt file?
Upload the robots.txt file to your website's root directory. For example, if your site is https://example.com, the file should be accessible at https://example.com/robots.txt. The file must be in the root directory - it won't work if placed in a subdirectory. Most hosting control panels and FTP clients allow you to upload files to the root directory.
Related Tools You Might Like
Meta Tags Generator
Generate optimized meta tags for your pages
Google SERP Preview
Preview how your page appears in Google search results. Optimize titles and descriptions for CTR
JSON-LD Schema Generator
Generate structured data (JSON-LD) for rich snippets
XML Sitemap Generator
Generate XML sitemaps for search engines with change frequency and priority
Related Guides
More in Technical SEO tools · All free SEO tools