You publish a page. You wait. Nothing shows up in search. You check the URL in a browser and it loads perfectly, so the page clearly exists. So why doesn't Google know about it?
Nine times out of ten, it's because the page never got crawled, or it got crawled and then dropped before indexing. Those are two completely different failures, and confusing them is why so many site owners waste weeks tweaking the wrong thing. I've done it myself. I once spent eleven days rewriting meta descriptions on a client's blog before realizing the real problem was a stray noindex tag left over from staging that had been sitting in the template for months.
This piece walks through how search engines actually crawl and index pages, where the process breaks, and what you can control. By the end you'll know why some pages get indexed in hours while others sit in limbo for a month, and how to diagnose which stage is failing you.
Key Takeaways
- Crawling and indexing are separate stages: a page can be crawled without being indexed, and never indexed without being crawled first.
- Googlebot discovers URLs through links, XML sitemaps, redirects, and manual submission. No links pointing in means very slow discovery.
- Crawl budget is real, but it only matters on sites with thousands of URLs. On a 40-page site, it's a non-issue.
- Blocking a page in robots.txt does not remove it from the index if it's already there.
- Indexing frequency tracks authority and freshness far more than anything you do on-page.
- Search Console is the only honest source of truth about what Google actually did with your pages.
How search engines crawl and index pages: the two-stage pipeline
Picture a librarian who never sleeps. She walks the shelves looking for new books (crawling), then decides which ones go into the card catalog (indexing). She can pick up a book, flip through it, and put it back without ever cataloging it. That's the whole thing in one image.
The crawler's job is discovery and fetching. It follows links, reads your sitemap, and requests the HTML. Then a separate system parses what came back, extracts text, links, images, and structured data, and decides whether the page deserves a slot in the index.
What crawling actually means
Crawling is a fetch. A bot requests a URL, your server responds, and the bot stores what it received. Nothing is ranked, nothing is searchable yet. The page has simply been seen.
Three things determine whether your URL gets fetched in the first place: whether any link points to it, whether your robots.txt allows it, and whether the crawler has budget left for your host that day. That third point trips people up constantly. If your site has 400,000 product pages and only 200 get crawled daily, most of your catalog is invisible most of the time.
What indexing actually means
Indexing is the storage decision. Once fetched, the page gets rendered (JavaScript executed), deduplicated against other versions, and analyzed for topic, quality, and intent. If it passes, it lands in the index and becomes eligible to appear in results.
Here's the part most guides skip: being in the index doesn't mean you'll rank. It means you're in the pool. Ranking is a third stage, and it happens per query, in milliseconds, across billions of documents.
I've had pages indexed within four hours on a site with solid authority, and I've watched a brand-new domain take 26 days to get a single page indexed. Same content quality, same technical setup. The difference was entirely trust.
Does Google crawl noindex pages?
Yes. And this surprises people every single time.
A noindex tag tells the engine "don't put this in your catalog." It says nothing about fetching. The crawler still visits the URL, reads the HTML, sees the directive, and moves on without indexing it. The page gets crawled and discarded.
Why does that matter? Two reasons. First, a noindexed page still consumes crawl budget. Second, and more dangerously, if that page is blocked in robots.txt and carries a noindex tag, the crawler can't read the noindex at all, because it can't fetch the page. The two directives cancel each other out, and the URL can stay indexed indefinitely.
That exact combination is one of the most common technical SEO mistakes I run into. Someone blocks a staging folder in robots.txt, adds noindex for good measure, and then can't figure out why the pages keep showing up in search months later.
- Want a page crawled but not indexed? Use noindex, keep it crawlable.
- Want it neither crawled nor indexed? Use robots.txt blocking, and accept that removal from the index has to happen separately.
- Want it gone from results right now? Use the URL removal tool, then fix the underlying directive.
- Want it indexed? Do none of the above.
What is the difference between crawling and indexing in search engines?
Crawling is discovery and retrieval. Indexing is storage and eligibility.
Think of it as the difference between a mail carrier and a filing clerk. The carrier brings the letter to the building. The clerk decides which drawer it goes into, or whether it gets shredded. A page can be delivered and shredded. It can never be filed without being delivered.
The practical consequence: when a page isn't performing, you need to identify which stage failed before you touch anything.
| Stage | What happens | Failure symptom | Fix |
|---|---|---|---|
| Discovery | Bot finds the URL via links, sitemap, or submission | URL "unknown to Google" in Search Console | Add internal links, submit sitemap |
| Crawling | Bot requests and fetches the HTML | "Discovered – currently not indexed" | Improve server speed, check robots.txt, reduce URL count |
| Rendering | JavaScript executes, final DOM is built | Content missing in the rendered version | Server-side render key content |
| Indexing | Page is analyzed and stored | "Crawled – currently not indexed" | Improve content depth, remove thin pages, build authority |
| Ranking | Page competes for specific queries | Indexed but no impressions | Relevance and competition problem, not technical |
Those two middle failure states in the third column are the ones I see most, and they mean very different things. "Discovered – currently not indexed" means the crawler knows the URL exists but hasn't bothered to fetch it. Usually a budget or priority issue. "Crawled – currently not indexed" means it fetched the page, looked at it, and decided it wasn't worth storing. That's a quality signal, not a technical one.
How do I stop Google from indexing certain pages?
Match the tool to the situation, because they are not interchangeable.
The meta robots noindex tag
This is the right answer for pages you want reachable but excluded from results: internal search results, thank-you pages, filtered category URLs, tag archives that add no value. Drop it in the <head> of the page and leave the URL crawlable so the directive can actually be read.
robots.txt disallow
Use this when you want to prevent excessive crawling, not when you want to remove something from the index. Blocking a folder stops the fetch. If the URL is already indexed, it stays there, because the crawler can no longer see your noindex instruction. I've watched this trap cost a client nearly three months of confused back-and-forth.
Password protection and removal tools
Genuinely private content shouldn't rely on directives at all. Put it behind authentication. For URLs that are already indexed and need to disappear quickly, the removal request tool gets them out of results within roughly a day, but it's temporary. Fix the root cause in parallel or they come back.
How frequently does Google crawl and index your website?
There's no fixed schedule, and anyone quoting a specific number is guessing.
In practice, frequency tracks a handful of signals. Sites with strong authority and frequently updated content get crawled daily or even multiple times a day. A small site publishing once a month might get visited every few weeks. Pages that get updated after being stale often trigger a recrawl sooner, because the engine learns that this URL changes.
For new sites, expect slower everything. When I launched a niche site from scratch, the homepage took about nine days to get crawled. New posts after that took four to six days each. On an established domain I work with, new articles are typically crawled within two hours and indexed the same day. Same process, wildly different timelines.
You can nudge things: keep your sitemap accurate and current, link new pages from pages that already get crawled often, and resubmit URLs in Search Console after meaningful updates. You cannot force it.
How to diagnose why a page isn't indexed
Start with Search Console's URL inspection tool. It tells you the last crawl date, the rendered HTML, and the exact index status. Then work backward.
If it says discovered but not crawled, your discovery signals are weak or your budget is being eaten elsewhere. Check for redirect chains, crawlable faceted navigation, and soft 404s. If it says crawled but not indexed, stop looking at technical factors. That's almost always a content or authority judgment, and no amount of server tuning will fix it.
One more thing worth internalizing: orphan pages. A URL with zero internal links pointing to it is nearly invisible. I audited a site last year with 340 pages, and 61 of them had no internal links at all. They were in the sitemap, so they'd been discovered, but nothing signaled importance. Adding internal links from relevant pages got most of them indexed within two weeks.
The index isn't a container that holds everything you publish. It's a curated list, and you're competing for a slot. Understanding that reframes the whole problem: you're not trying to make Google see your page, you're trying to make it worth keeping.