What crawl budget means and why it matters to your site

Crawl budget is the number of pages on your website that a search engine will actually visit and read in a given time period. Search engines like Google don't crawl every page on every site every day — they allocate a limited amount of resources to each domain, and crawl budget is that limit. If your site has 10,000 pages but Google only crawls 2,000 of them in a month, your crawl budget is being spent on those 2,000 pages, and the other 8,000 are being ignored.

This matters because pages that don't get crawled don't get indexed, and pages that don't get indexed won't show up in search results. If important pages on your site aren't being crawled because your crawl budget is exhausted on less important pages, you're losing visibility in search results. The goal is to make sure your crawl budget is spent on the pages that matter most to your business.

Crawl budget is especially important if you run a large site, have a lot of duplicate content, or have pages that change frequently. Small sites with a few hundred pages usually don't need to think about crawl budget at all — search engines will crawl everything. But as your site grows, managing crawl budget becomes a real part of keeping your site visible in search results.

Key Takeaways

  • Crawl budget is the number of pages a search engine will visit on your site in a given period, and it's limited based on your site's size and authority.
  • Pages that aren't crawled won't be indexed, which means they won't appear in search results even if they're well-written and relevant.
  • You can see how many pages Google is crawling by checking the Crawl Stats report in Google Search Console.
  • Blocking low-value pages with robots.txt, removing duplicate content, and fixing broken links all help you use your crawl budget on pages that matter.
  • Crawl budget becomes a real concern when your site has thousands of pages or a lot of pages that change frequently.

How search engines decide how much crawl budget to give your site

Search engines calculate crawl budget based on two main factors: crawl rate limit and demand. The crawl rate limit is how fast the search engine is willing to crawl your pages without overloading your server. The demand is how often the search engine thinks your pages change and how much traffic they get. A site with pages that update every hour and get thousands of visitors will get a higher crawl budget than a site with pages that never change and get no traffic.

Your site's authority also affects crawl budget. A well-established site with many backlinks and a history of quality content gets crawled more often than a new site with few backlinks. This creates a feedback loop: popular sites get crawled more, so their new content gets indexed faster, which keeps them popular. New sites have to earn their crawl budget over time by building authority.

The size of your site matters too. A site with 100 pages will get crawled more completely than a site with 100,000 pages, simply because there's less to crawl. Google has to spread its resources across the entire web, so larger sites have to compete harder for crawl resources.

How to check your crawl budget in Google Search Console

You can see how many pages Google is crawling on your site by opening Google Search Console, going to the left menu, and clicking Settings. Then click Crawl Stats. This report shows you the average number of pages crawled per day, the total kilobytes downloaded, and the time spent crawling your site over the last 90 days.

Look for trends in the data. If the number of pages crawled is dropping month to month, that's a sign your crawl budget is shrinking. If it's staying flat while you're adding new pages, those new pages might not be getting crawled. If you see a spike in pages crawled after you made changes to your site, that's usually a good sign — it means Google noticed something changed and sent more crawlers to check it out.

Keep in mind that the Crawl Stats report only shows you what Google is doing. Other search engines like Bing have their own crawl budgets, but they don't provide the same level of reporting. If you want to see crawling activity from all search engines, you can check your server logs, but that's more technical and usually only necessary for very large sites.

Pages that waste your crawl budget

Some pages on your site consume crawl budget but don't help your search visibility. Duplicate content is the biggest culprit — if you have the same page accessible at multiple URLs, search engines will crawl all of them, wasting budget on copies instead of unique content. This happens often with URL parameters (like ?page=1, ?sort=price), session IDs, or printer-friendly versions of pages.

Broken links and redirect chains also waste crawl budget. When a search engine crawler follows a link and hits a 404 error or a chain of redirects, it's spending resources on a page that doesn't help your site. Similarly, pages that return a 500 server error or time out will be crawled repeatedly as the search engine tries to access them again.

Low-value pages like login pages, thank-you pages, search results pages, and auto-generated pages with little unique content also consume budget without providing search value. If you have thousands of these pages, they can eat up a significant portion of your crawl budget that could be spent on pages that actually drive traffic.

How to protect your crawl budget

The most direct way to protect your crawl budget is to block low-value pages using your robots.txt file. This is a simple text file in your site's root directory that tells search engines which pages to crawl and which to skip. You can block entire directories, specific file types, or individual pages. For example, you can block all pages in your /admin/ folder or all pages with ?sort= parameters.

Remove or consolidate duplicate content. If you have multiple versions of the same page, pick one as the canonical version and use the rel="canonical" tag to tell search engines which one to index. This tells the crawler to spend its budget on the real version instead of wasting time on copies.

Fix broken links and redirect chains. Use a site crawler tool like Screaming Frog or your Google Search Console to find pages that return errors or redirect multiple times. Redirect chains waste crawl budget because the crawler has to follow each redirect before reaching the final page. Ideally, redirects should go directly to the final destination in one hop.

Keep your site structure clean and organized. A logical hierarchy with clear navigation makes it easier for search engines to find and crawl your important pages. Avoid creating pages that are buried deep in your site structure or only accessible through complex navigation — search engines might not find them at all.

When crawl budget becomes a real problem

Most small and medium-sized sites don't need to worry about crawl budget. If your site has fewer than 10,000 pages and you're not adding hundreds of new pages every week, search engines will crawl everything you want them to crawl. Crawl budget only becomes a constraint when you're running a very large site or when you have a lot of pages that change frequently.

E-commerce sites with thousands of product pages, news sites that publish dozens of articles daily, and large content sites with millions of pages all need to think about crawl budget. If you're in one of these categories and you notice that new pages aren't getting indexed quickly or that old pages are disappearing from search results, crawl budget might be the reason.

You can also run into crawl budget problems if your site is slow or if your server is overloaded. Search engines will crawl your site more slowly if your pages take a long time to load, which effectively reduces your crawl budget. Improving your site speed and server performance can increase how much your site gets crawled.

The difference between crawl budget and indexing

Crawl budget and indexing are related but separate. Crawling is when a search engine visits your page and reads the content. Indexing is when the search engine adds that page to its database so it can show up in search results. A page can be crawled but not indexed if it has a noindex tag, if it's blocked by robots.txt, or if the search engine decides the content isn't valuable enough to index.

You can see which pages Google has indexed by searching "site:yoursite.com" in Google. This shows you all the pages in Google's index for your domain. If you see pages in your Google Search Console Crawl Stats that aren't showing up in this search, those pages were crawled but not indexed. This usually means they have a noindex tag or they're duplicate content that Google decided not to keep in its index.

Frequently Asked Questions

Does crawl budget affect my search rankings?

Crawl budget doesn't directly affect your rankings, but it affects whether your pages get indexed at all. If a page isn't crawled, it can't be indexed, and if it's not indexed, it won't rank in search results. So indirectly, a poor crawl budget strategy can hurt your visibility.

How often does Google crawl my site?

This varies based on your site's size, authority, and how often your content changes. A popular news site might get crawled multiple times per day, while a small business site might get crawled once a week or less. You can see the actual crawl frequency for your site in Google Search Console's Crawl Stats report.

Can I increase my crawl budget?

You can't directly request a higher crawl budget from Google, but you can earn one by improving your site's authority, fixing technical issues, and removing low-value pages. A faster, cleaner site with high-quality content and strong backlinks will naturally get crawled more often.

What's the difference between crawl budget and crawl rate?

Crawl rate is how fast Google crawls your pages (pages per second), while crawl budget is the total number of pages Google will crawl in a given time period. You can adjust the crawl rate in Google Search Console if your server is being overloaded, but crawl budget is determined by Google based on your site's characteristics.

Should I use robots.txt to block pages or noindex tags?

Use robots.txt to block pages you don't want crawled at all, like admin pages or duplicate content. Use noindex tags for pages you want crawled but not indexed — this is useful for pages like login pages or search results that you want search engines to see but not include in their index.