Daily Web Digest

Guides

How Search Engines Find New Pages: Crawling, Indexing and Ranking Explained

What happens after you hit publish? A plain-English guide to crawling, indexing, ranking, sitemaps, IndexNow, robots.txt and Search Console statuses.

By Daily Web Digest Editors
7 October 2026 · 10 min read
In this guide
  1. The three stages in a nutshell.
  2. Stage 1: Exploration and search.
  3. Stage 2: Indexing
  4. Stage 3: Ranking (serving results)
  5. robots.txt vs noindex: two jobs.
  6. Search console "Discovered" and "Crawled - currently not indexed" in Search console.
  7. Easy audit of each new page.
  8. Wrapping up
  9. Sources

You post something, and even send it to several WhatsApp chats and look it up in Google an hour later. Nothing. Nothing a week later. Is something broken?

Usually not. Search engines must locate a page, retrieve it, make decisions of whether the page should be indexed or not, after which only decisions are made on where to be displayed. The real problem and normal waiting are similar except that knowing how each step functions can assist you in identifying which is a practical one.

We will go through one page of the process: a post at example.com/kerala-banana-chips/ published on a food blog about two years old with a number of posts published each month.

The three stages in a nutshell.

Google describes Search as working in three stages, and says "not all pages make it through each stage".

Stage What happens What you can impact.
Crawling Google finds the URL and downloads the page Internal links, sitemaps, robots.txt, server health.
Indexing Google breaks down the page and determines whether to store it or not. Content quality, duplicates, noindex, canonical tags
Serving (ranking) Google selects the indexed pages that it displays in a search. How relevant and useful the page is

Two expectations are set in the same Google page. It is not reliant on payment in order to crawl a site more often, or rank higher, and Google does not in any way promise to crawl, index or serve any of its pages, even those that follow its recommendations. Nobody can offer guaranteed or instant Google indexing, as this is not something that is offered by Google itself.

The way Google gets to know about the existence of a URL.

Google does not have the overall address book of all pages on the web, and hence continues to search. It calls this "URL discovery", and its documentation describes three routes:

  • Pages it has seen before, which it reacts.
  • Links. Google crawl of a well-established page offers a link to a new page, the URL is added to the list. It is a category page on Google that directs to a new blog article.
  • Sitemaps, documents containing the URLs that you want search engines to know about.

In the case of our banana chips post, the most probable route is that Googlebot re-reads the page in the "Snacks" category of the blog, notices the link to the new one, but puts it in the queue.

The links do not assist unless Google can correctly follow it: Google claims to only be able to crawl the link when it is an <a> HTML element that has an href attribute. It also suggests that each of the pages you are interested in ought to be linked to at least one other page within your web.

How a crawl works out.

Discovery doesn't mean an immediate visit. Googlebot relies on an algorithm to determine which websites to crawl, the frequency and the number of pages to retrieve per website. It also attempts not to fill up your server and relinquishes when your site is giving you errors like I (HTTP 500) errors.

When crawling, Google loads the page in Chrome in one of the latest releases and executes its JavaScript. Google may not render contents loaded by your theme using scripts. Common crawl blockers, per Google, are server problems, network issues and robots.txt rules.

IndexNow: Bing (and others) push.

The such manipulation as links and sitemaps are the pull techniques: you provide some URLs and wait. IndexNow, is a push technique, an open protocol and allows a site to inform participating search engines of addition, update, or deletion of a URL. Microsoft Bing, Yandex, Naver, Seznam.cz and Yep are also considered engines listed on indexnow.org and a URL posted to one is shared with the rest.

To demonstrate ownership, you set your key file on your site, and submit modified URLs to an IndexNow endpoint, at a rate of up to 10,000 URLs per request. A lot of platforms will do it automatically; such as Wix, which automatically submits page changes in their Premium plans.

Two caveats. To begin with, submit success is just a confirmation of receipt: according to documentation by the IndexNow status, an HTTP 200 reply indicates to the search engine that you have submitted a URL to it, nothing beyond that. Second, Google doesn't use IndexNow. Google announcement In 2021, Google announced it would test the protocol, but in October 2026 it was not a participant, and Google's official documentation of how to request a recrawl includes only the URL Inspection tool and sitemaps.

IndexNow could then assist Bing to grab our recipe earlier. In the case of Google, internal links, and sitemap matter.

Stage 2: Indexing

After crawling a page, Google determines what is on the page by running the text, the <title> element, images and alt text, videos and other elements.

A large portion of indexing is dealing with duplicates. Google groups use similar pages and select one to be a canonical, the one that most likely represents the group. Google will generally index one of the versions, and consider the rest as alternatives, provided that our recipe can also be accessed at a print-friendly URL, or when tracking parameters are included.

Google points out three possible reasons why pages are not indexed: low-quality content, robots meta rules, and indexing-unfriendly site designs. It is also usual that certain URL can remain out. According to Google indexing pages report, it is not possible to index all of the URLs on your site, rather just the canonical pages.

Stage 3: Ranking (serving results)

As one does a search, Google will go through its index and display the web pages that it thinks are most relevant and of the best quality. Relevance is determined by hundreds of factors, which may be dependent on the location of the searcher, the languages used and the device.

It is so indexed and visible. Our recipe may be categorized but not show on a general search such as banana chips recipe but appears on more specific searches that are closely related to it. By not indexing pages Google observes that they will not be displayed when they are irrelevant to the search terms by people or the quality of such pages is poor.

robots.txt vs noindex: two jobs.

The two are confused all the times.

robots.txt noindex
Controls Crawling: accessibility of a URL by the bots indexing: the possibility of a page to be shown in the results.
Lives in A file in the root of your web address, e.g. example.com/robots.txt. A page X-Robots-Tag in the head of the page, or (X-Robots) HTTP header.
Google's stated purpose Being able to control crawler traffic to prevent overloading your site by bots. Storing an out of search result page.
Faithfully save a page with Google? No Yes, after Google recrawling the page.

Common mistakes

1. Hiding a page with robots.txt. Google declares robots.txt is not a tool of exclusion of a web page by Google. Blocked URLs still may be indexed when they have links to them and in most cases these links have no description. Instead use noindex or password protection.

2. Using a noindex with robots.txt block. The two cancelling, but it is like added insurance. To get noindex working, Google recommends the involvement of a page that is not blocked by robots.txt file. In case the page cannot be loaded by Googlebot, the tag is never looked at.

3. Empirical: Writing noindex in robots.txt. On 1 September 2019, Google ceased to support this. The rule is on the page or on the HTTP header.

4. Leaving launch settings on. Noindex, nofollow robots meta tag Since WordPress 5.3, the box mentioning discourage search engines to index this site (Settings > Reading) adds a noindex, nofollow robots meta tag. Back off a staging environment, malicious if still checked on post-launch. Similarly a Disallow: / rule copied off a test server.

5. Blocking CSS and JavaScript. Google instructs that the resource files should not be blocked when their loss causes a page to be difficult to comprehend by its crawler.

Search console "Discovered" and "Crawled - currently not indexed" in Search console.

These 2 statuses in the Page indexing report in Search Console refer to various steps.

Discovered – currently not indexed. Google has heard of the URL but has not crawled it yet. According to the help page of Google, this normally occurs when the Google wanted to crawl the URL but it anticipated that crawling will burden the site, thus it was re-scheduled. The reason the last crawl date is blank is because of this.

Crawled – indexed not in progress. The page has been fetched by Google though it is not indexed. The help page of Google mentions that it could or could not be indexed in the future, and there is no need to re-post it.

On the first level: According to Google crawl budget manual, crawl demand of Googlebot is determined by the size of a site, its update frequency, the quality and relevance of a page to the rest of the sites, such a site with few links can just wait longer.

On the second one: The URL Inspection assistance of Google says that a crawled page is not yet eligible to meet other requirements, such as it should not duplicate another indexed page and it must be of high quality to be worthy of indexing.

What is realistic on your part.

  1. It is a technical block rule out. Complete a live test in the URL Inspection tool of Search Console. Ensure crawling is enabled, page loads, no extraneous noindexing and you have a canonical where you expect.

  2. Connection with the page on pages visited by Google. Include links in your home page, category pages and related posts that are indexed. In our recipe, that might be a link to an indexed post on Onam snacks, where the anchor text is descriptive such as "Kerala-style banana chips" instead of a text link that reads click here.

  3. Index the page. The helpful content guidelines provided by Google request users to answer whether a page contains original information, reporting, research or analysis. A duplicate of what fifty other sites say accomplishes little to persuade Google to store one also. Include what you alone know to: quantities tested, step photos, what failed the first time.

  4. Merge near-duplicates. Three comparatively thin posts on the same subject are generally more apt to be one efficient post. The crawl budget recommendation by Google also recommends to consolidate on duplicate material.

  5. Use "Request indexing", once. URL Inspection allows you to request that Google be crawled on a single URL, with restrictions:

    • It has a limit of daily requests.
    • "A recrawl request of the same URL is not going to crawl the pages any faster.
    • An indexing is not assured by a request.
    • You can't request indexing for a page the live test finds non-indexable.

    On numerous pages, Google suggests a sitemap, indicating changed pages with the aid of <lastmod>Designations.

  6. Give it time. Google claims that crawling may take few days to a few weeks. After fixing a technical issue flagged in the report, "Validate fix" typically takes up to about two weeks, sometimes longer.

Easy audit of each new page.

  • Linked to no less than one of your site indexed pages, normal <a href.
  • On your sitemap (this is automatically included on most platforms)
  • Not disallowed in robots.txt, and no unintentional noindex.
  • Promotes something that searchers cannot already find on the pages today will rank.
  • Opened one time in URL Inspection, not 20 times.

Wrapping up

Search these engines identify the majority of their pages via links, crawl them according to their schedule, index the pages they consider worth saving and rank them on a scale against all the other pages. It's impossible to jump steps, but you can eliminate obstacles: bring internal connections to light, use an appropriate sitemap, appropriate use of robots.txt and noindex and index-worthy pages. IndexNow is a valuable addition to Bing and its partners. In the case of Google, they are still good links, good pages and time.


Sources