Crawl budget is the amount of crawling a search engine allocates to a site. It is a function of how many pages an engine is willing to request (crawl demand) and how much the server can handle without slowing down (crawl rate).
Who needs to care
For a small site of a few hundred pages, crawl budget is rarely a concern; engines crawl it fully with ease. It becomes important for large sites: ecommerce catalogs, news archives, real-estate listings, and any site generating thousands of URLs, including parameterized and faceted ones.
On those sites, wasted crawling on low-value or duplicate URLs means important pages get crawled less often, so new content is discovered slowly and updates take longer to reflect.
How to manage it
- Cut duplicate and low-value URLs. Use canonical tags, and block faceted-navigation traps and infinite parameter combinations.
- Keep a clean sitemap. List only canonical, indexable URLs so engines prioritize them.
- Fix broken links and long redirect chains. Both burn crawl resources.
- Maintain server speed. A fast, stable server raises the crawl rate an engine is willing to use.
Relevance to AI search
AI crawlers face the same efficiency limits. If your important pages are buried behind crawl waste, both traditional indexing and AI citation eligibility suffer, because the engine simply reaches that content less often. A lean, well-structured site gets its best content seen and refreshed sooner.