Market Muser

How Search Engines Crawl and Index Pages: A Complete Guide

Crawling and indexing are two separate gates—block the wrong one in robots.txt and Google may still surface your page. Here's why a noindex tag can't work if the crawler never gets to read it.

How Search Engines Crawl and Index Pages: A Complete Guide

How search engines crawl and index pages: the two gates every URL has to pass

Someone asked me last week why a page they had published was showing up in Google when they had explicitly told Google to ignore it. They had blocked it in robots.txt. Clean, deliberate, confident. And the page was still there.

That single mistake explains almost every confusion people have about how search engines crawl and index pages. Crawling and indexing are two separate gates, controlled by two separate mechanisms, and blocking the wrong one gets you the exact opposite of what you wanted.

Here is the part most guides skip: a crawler can visit a page it will never index. It can also index a page it refuses to crawl. Both scenarios are normal, and both break the mental model most people carry around.

Key Takeaways

  • Crawling is discovery and retrieval; indexing is storage and interpretation. Different gates, different controls.
  • A noindex page can absolutely be crawled. Googlebot needs to read the page to find the rule that says "do not index this."
  • If robots.txt blocks a page, the crawler never sees your noindex tag — and the URL can still surface in results through external links.
  • Crawl frequency is not a schedule you set. It follows signals: internal linking, freshness, site authority, and how often your content actually changes.
  • Search Console tells you what Google did. It rarely tells you why. You infer the why from crawl stats and coverage reports.

What is the difference between crawling and indexing in search engines?

Crawling is the act of fetching. Indexing is the act of filing. A crawler downloads the HTML of a URL, and that download is the crawl. Whether anything gets stored afterward is a completely separate decision made later, sometimes on different infrastructure, sometimes days apart.

What is the difference between crawling and indexing in search engines?

The practical consequence is that a crawl is not a promise of visibility. I have watched crawl stats spike on a client's site by roughly 40% after a site migration while indexed pages dropped by a few hundred. Googlebot was fetching plenty. It just wasn't keeping any of it.

The three stages inside "indexing"

Indexing is not one action. It is a sequence, and each stage can fail independently.

  • Parsing — the crawler extracts text, links, and directives from the downloaded HTML.
  • Canonicalisation — the engine decides which of several similar URLs is the "real" one.
  • Rendering — JavaScript-heavy pages get queued for a headless browser pass, which is slower and sometimes incomplete.

Your page can survive stage one and die at stage three. A client of mine had a product grid that rendered entirely through client-side JavaScript. Googlebot fetched it, parsed it, found an almost empty DOM, and indexed a page that looked blank to anyone searching for the product name. Three weeks of confusion before we traced it.

Crawling, indexing, ranking — where each one fits

Ranking is the third step, and it only happens to URLs that made it through the first two. You cannot rank a page that was never indexed, and you cannot index a page the crawler was never allowed to fetch. Order matters.

Which is why the sequence people usually get wrong is the last one. They obsess over rankings for URLs that never cleared the indexing gate at all.

Does Google crawl noindex pages?

Yes. And this is the detail that trips up almost everyone.

Does Google crawl noindex pages?

A noindex directive is a rule, delivered either as a meta tag or as an HTTP response header, that tells supporting search engines to drop the page from results. But for that rule to have any effect, the crawler has to actually read it. That means the page must not be blocked by robots.txt, and it must be accessible to the crawler in the first place.

If the page is blocked by robots.txt, the crawler never reaches the HTML, never sees the noindex tag, and the URL can still appear in search results — often because other sites link to it. You end up with a page you tried to hide, showing up with no description you control.

Why this counter-intuitive design exists

Because robots.txt governs crawling, not indexing. It is a request to stay out of a directory, not a request to forget what is inside it.

Someone at a previous agency job asked me why we didn't just block a staging environment in robots.txt. We did. It appeared in the index anyway, because the staging URLs had picked up inbound links from a press release that leaked early. The fix was noindex tags — which required first unblocking the staging site so the crawler could read them.

That sequence feels backwards the first time you do it. Unblock so you can de-index. But it is the only way it works.

How do I stop Google from indexing certain pages?

Three tools, three different jobs. Mixing them up is where most de-indexing attempts fail.

How do I stop Google from indexing certain pages?
Method What it controls Page stays indexed? Use when
robots.txt disallow Crawling access Possibly yes You want to save crawl budget on low-value URLs
noindex meta tag or header Indexing No, once crawled You want the URL gone from results
Canonical tag Which URL represents the content Consolidated into the target Duplicate or parameter variants

Read that table again, because the first row is the one people misread. Blocking crawling does not remove anything from the index. It removes your ability to manage what is already there.

The correct order of operations

  1. Remove any robots.txt block on the URLs you want gone.
  2. Add the noindex directive — meta tag or HTTP header, whichever fits your stack.
  3. Wait for the crawler to revisit and read it.
  4. Confirm removal in Search Console before re-blocking anything.

Step three is where patience runs out and people start filing reconsideration requests they don't need. The crawler has to come back. It will, but not on your timetable.

How frequently does Google crawl index your website?

There is no fixed schedule, and anyone who gives you one is guessing. Crawl frequency varies wildly between sites, and within the same site between sections.

What actually drives it:

  • How often your content changes. A news site gets crawled constantly. A brochure site with a page last edited in 2023 might get a monthly visit, if that.
  • Internal linking structure. Pages buried five clicks deep get found late and revisited rarely.
  • Site-level signals. Established domains with clean technical health get more attention per URL than new ones.
  • Crawl budget consumption. If a large share of your crawl requests returns 404s, redirects, or soft errors, the crawler learns to visit less.

On one project I tracked crawl stats across a 12,000-URL site for two months. Roughly a third of the URLs accounted for the overwhelming majority of crawl requests — mostly the blog and the category pages. Product pages with thin content were visited a handful of times and then effectively abandoned.

That pattern is not a penalty. It is the crawler spending its budget where it expects to find something worth indexing.

Where to actually look

Search Console's crawl stats report gives you request volume, response codes, and file types. The page indexing report tells you which URLs made it and which were excluded, with a reason attached.

Neither report tells you the "why" behind a low crawl rate. You infer it: check for crawl waste, check your internal linking, check whether your content is actually changing.

Do I need to submit a sitemap for pages to get crawled?

No. Crawlers discover URLs automatically by following links. A sitemap is a hint, not a requirement — but on large or poorly linked sites it meaningfully speeds up discovery.

Can a page be indexed without ever being crawled?

No. Indexing requires the crawler to have fetched and parsed the page. There is no other route in.

The part nobody tells you about crawling and indexing

Both gates are controlled by an engine that optimises for its own resources, not your visibility goals. That is the whole story, and it reframes every technical decision you make.

Which is why the "just block it in robots.txt" instinct feels right and fails so often. You are asking the crawler to respect a rule it will never read.

Next time a page refuses to leave the index, check whether you blocked the gate that lets the rule through. Odds are good you did.

Rachel Ellery

Rachel Ellery

Rachel Ellery is a technical SEO consultant who helps organizations improve search performance through rigorous audits, thoughtful site architecture, and Core Web Vitals optimization. Known for translating complex technical findings into clear, actionable roadmaps, she works closely with development and marketing teams to build faster, more crawlable websites. Her approach blends analytical precision with a practical, collaborative style that makes technical SEO accessible to stakeholders at every level.

See all articles →

Related articles