1The three jobs, in order
Google answers a search from a stored copy of the web. A page gets into that copy in three steps: Googlebot fetches it (crawling), Google reads it and files it under one address (indexing), and when somebody searches, Google orders the stored pages and builds a results page (ranking). Search engine optimization (SEO) is work on all three steps. This module covers the first two in full and the shape of the results page; module 03 covers how the ordering is done. By the end you have written down how many results Google shows for the domain, which of our pages shows for three test searches, and what that means.
Module 01 covered what the person typing wants. This module covers what Google does with the typing. When an office manager in Henrico types "managed IT services Richmond VA", Google does not go out and read the web at that moment. It looks in its index, which is a stored and sorted copy of pages it fetched earlier, and picks from that. A Hermetic page that was never fetched cannot appear. A page that was fetched and then filed under a different address cannot appear under its own address either. Everything else in this course sits on top of those two facts.
The three steps happen in a fixed order, and each one has a Hermetic consequence.
Crawling
Googlebot, the program Google uses to fetch pages, requests a URL the way a browser does and gets back a status code and the page's HTML. Our sitemap offers it 392 URLs. Each one is a fetch Googlebot may or may not make.
Indexing
Google renders the fetched page, reads the text, works out what it is about, decides whether it duplicates another page, chooses one address to represent the group, and stores the result. Our three managed IT pages meet this step and lose.
Ranking
A query arrives. Google pulls candidate pages from the index, orders them with its ranking systems, and assembles a page from several indexes at once: web pages, business listings, images, questions. Module 03 covers the ordering.
Four controls act on that pipeline, and all four appear in the Rank Math course as settings. A sitemap and links feed the first step. robots.txt and the server's status code govern the fetch. noindex and the canonical choice govern what is stored and under which address. The third step has no switch of its own; it is what the other two earn.
2The words of a query
A query is the text a person types or speaks into the search box. People compress their language when they search. The office manager who would ask a colleague "do you know which kind of multi-factor authentication (MFA) is the safest" types "most secure mfa" into Google, and both get read as the same request. Google reads a query loosely: it treats singular and plural as the same word, corrects common misspellings, swaps in synonyms, and reads each word in the context of the words around it. Google documents this as the first stage of building results, working out the meaning of the query before matching it to pages.
Loose reading has a consequence for how many pages we write. A single post that answers "which form of MFA is the most secure" can show for "most secure mfa", "safest 2fa method" and "is text message mfa safe", because Google maps those to one meaning. We do not need a page per phrasing. Module 06 of this course groups phrasings so that one page covers a cluster.
Three things about the words still change what comes back, and you can test each one in a private browser window in under a minute.
Word order
"Richmond IT support" and "IT support Richmond" return similar organic results most of the time, and sometimes a different set of features above them. Word order matters more for which features Google adds than for which pages it lists. Test it with both orders and compare the first screen.
Phrases in quotes
Quotation marks force an exact phrase in that order and stop the synonym swapping. "Hermetic Networks" in quotes keeps Google from reading "hermetic" as the adjective for an airtight seal. Google also offers a Verbatim option under Tools on the results page, which does the same for a whole query.
Questions
A query shaped as a question, with "which", "how" or "is" at the front, tends to get a question-shaped answer: a featured snippet or a People also ask block, both covered in section 7. Module 05 collects the questions clients ask us so that posts can answer them in the same words.
Nobody at Hermetic controls how a stranger types, which is why module 05 finds out how they type before anything is written. What this section gives you is a reading habit: when a search fails to show us, first check whether Google read the query the way we assumed.
3Search operators
A search operator is a word with a colon, or a symbol in front of a word, that changes what a search returns rather than what it is about. Google publishes a short list and changes it from time to time, so five of them earn a place here because they have been stable for years and because each one tells you something about our own site in one search. Operators can be combined in one query as long as they do not contradict each other.
| Operator | What it does | Search to run | What it tells you about our site |
|---|---|---|---|
"..." | Exact phrase, in that order, no synonyms | "Hermetic Networks" -site:hermeticnetworks.com | Where the company is named on other sites: directories, the Google Business Profile, partner pages. The raw material for module 03's section on links. |
site: | Limits results to one domain, or one path on it | site:hermeticnetworks.comsite:hermeticnetworks.com managed IT | Roughly how many URLs Google holds for us, and which page it prefers for a topic. Section 8 makes this the first window into the index. |
intitle: | Only pages with the word in their title | site:hermeticnetworks.com intitle:richmond | How many of our pages carry Richmond in the title. Several pages with the same title are competing for the same search. |
inurl: | Only pages with the word in their URL | site:hermeticnetworks.com inurl:oldsite:hermeticnetworks.com inurl:thank | Whether the backup page ending in -old and the two thank-you pages are in Google. None of the three should be. |
-word | Excludes pages containing the word (no space after the minus) | site:hermeticnetworks.com richmond -inurl:guide | Our Richmond pages with the ten henrico-virginia-city-guide pages removed, so the service and city pages can be seen on their own. |
Two cautions. The result count at the top of an operator search is an estimate, and Google says so; it is never a complete listing. Use the count as a rough size and the listed URLs as the evidence. And the cache: operator, which older guides recommend for seeing Google's stored copy of a page, was retired in 2024 along with the cached link on results. The replacement is the URL Inspection tool in Search Console, in section 8.
One search is worth running before you read on. site:hermeticnetworks.com "phising" shows whether the misspelled phishing page is in Google under its misspelled address. If it is, a stored entry exists under the wrong word, and the redirect that fixes the slug in the Rank Math course has a reason beyond tidiness.
4Crawling
Crawling is the step where Googlebot finds out a URL exists and fetches it. The two halves have different failure modes, so take them in turn.
Discovery
Googlebot learns a URL three ways. It follows a link on a page it already has, which is how most of the web is found. It reads a sitemap, which is an XML file listing the URLs a site wants crawled; ours is at https://www.hermeticnetworks.com/sitemap_index.xml and lists 392 URLs. Or a person submits a single URL through Search Console. A page nobody links to and the sitemap omits is an orphan, and Googlebot has no way to reach it. Our tourism pages are the opposite problem: linked, listed, and findable, for a subject no client searches for.
Google says a site of about 500 pages or fewer whose pages link to each other may not need a sitemap for discovery. Ours has 392 URLs, so Googlebot could find them from the menu and the links between posts. The sitemap still matters as the list of pages we tell Google we care about, and today that list includes 84 tag archives, 18 categories, 3 author archives and two thank-you pages. Module 06 in the Rank Math course trims it.
The fetch
Googlebot requests the URL over HTTP and reads the status code first. A 200 means read the page. A 301 means follow the redirect and treat the destination as the page. A 404 or 410 means the page is gone, and after repeated visits Google drops it from the index. A 5xx error or a timeout tells Googlebot the server is struggling, so it slows down and returns later. Googlebot then renders the page, running its JavaScript in a recent version of Chrome, so content added by scripts is read too. Beaver Builder writes most page content into the HTML the server sends, so our pages do not depend on rendering; the View crawled page button in URL Inspection, in section 8, confirms it for any one page.
Crawl budget
Crawl budget is Google's term for how many URLs on a site Googlebot can and wants to fetch in a given period. Google's guidance is that it matters for sites with more than a million pages, or more than ten thousand pages that change daily. A 392-URL site is far below either line, and Googlebot will get through all of it. The cost of our 105 archive URLs, which is 84 tags plus 18 categories plus 3 authors, falls at the index and results steps rather than at the fetch, and sections 5 and 7 explain how.
robots.txt
robots.txt is a plain text file at the root of the site, ours at https://www.hermeticnetworks.com/robots.txt, that tells crawlers which paths they may request. It is a fetch rule and only a fetch rule. A URL that robots.txt disallows can still end up in the index if other pages link to it, because Google knows the address exists; it then appears in results with no description. Keeping a page out of Google is the job of noindex, in the next section, and noindex only works if Googlebot may fetch the page and read the instruction. Those two sentences are the ones most often got backwards.
The file on the site on 2026-09-29 allowed every crawler, including the AI crawlers behind ChatGPT, Claude and Perplexity, to fetch everything except /install; asked for a ten-second pause between fetches with a Crawl-delay: 10 line; and pointed at sitemap_index.xml. Google documents that it ignores Crawl-delay, so that line changes nothing for Googlebot, and nothing in the file blocks a page a client would want. Editing the file is a plugin screen: Rank Math module 04.
What blocks a crawl
- A Disallow line in robots.txt covering the path. Ours covers only
/install. - Server errors and timeouts. Repeated 5xx responses make Googlebot back off. Hosting is module 07 of this course.
- Bot protection. Cloudflare sits in front of the site and blocked our own automated reads on 2026-09-29. Cloudflare keeps a list of verified crawlers that includes Googlebot and lets them through by default; whether this site's rules follow the default is not known yet. The Crawl stats report in Search Console, under Settings, settles it.
- Links that exist only after a click. A link built by a script when a visitor clicks a tab is invisible to a crawler that does not click. Links in the HTML are safe.
- Logins. Googlebot has no password, so a gated page goes uncrawled.
- No link and no sitemap entry. The orphan case above.
5The index
The index is Google's stored, sorted copy of the pages it has crawled and decided to keep. For each page, Google stores the rendered content: text, title, headings, images with their alt text, structured data, outgoing links, the date of the last crawl, and one address under which all of it is filed. Since mobile-first indexing, the stored version is the one fetched as a phone, so what the page shows on a phone is what Google has. Google is explicit that it does not guarantee to index every page it crawls: a page can be left out because it duplicates another, carries a noindex instruction, or was judged too thin to keep.
The canonical URL
When several URLs show the same or nearly the same content, Google groups them and chooses one, the canonical URL, to represent the group in results. The others are recorded as alternates, and the signals that would have gone to them, links in particular, are consolidated onto the canonical. A site can state its preference in two ways: a rel="canonical" link element in the page's head, which both SEO plugins write for every page and point at the page itself by default, and a 301 redirect, which is the strongest statement because the alternate stops answering at all. Google treats the link element as a hint and the redirect as close to a decision.
Run our managed IT pages through it. Three URLs, /managed-services/, /managed-it-support-richmond-va/ and /managed-it-services-in-richmond-expertise-to-elevate-your-business/, describe the same service for the same city. If their text is similar enough for Google to call them duplicates, the index holds one entry for the three, filed under whichever URL Google picked, and today nothing on the site tells it which. If the text differs enough that Google keeps all three, the index holds three entries that compete for every managed IT search, and Google shows one per search, chosen by it. One entry under a URL we chose is the best outcome available, and it takes a redirect, which is why the Rank Math course's page map names a survivor for each duplicate set before anything is configured.
The same grouping covers the hostname. http://hermeticnetworks.com/, https://hermeticnetworks.com/, http://www.hermeticnetworks.com/ and https://www.hermeticnetworks.com/ are four addresses for one homepage. The site's canonical hostname is https://www.hermeticnetworks.com/, and the other three should redirect to it with a 301 so that the index holds one homepage. Module 07 of this course checks that they do.
noindex
noindex is an instruction on a page that tells Google to leave it out of the index even though it was crawled. It travels as a robots meta tag in the page head, <meta name="robots" content="noindex">, or as an X-Robots-Tag HTTP header. Both plugins set it per page and per content type. It is the right tool for the two thank-you pages and the cookie opt-out page, which visitors need and searchers never should. The condition from the last section applies: a page that is disallowed in robots.txt and also marked noindex stays in the index with no description, because Googlebot never fetched it to read the tag.
6The Knowledge Graph and the vertical indexes
The page index is one of several stores Google draws on, and two others matter to us.
The Knowledge Graph is Google's database of entities and the facts and relationships between them, kept separately from the pages. An entity is a specific person, place, organization or thing: Richmond, Virginia is one, Microsoft is one, Hermetic Networks can be one. The graph records what kind of thing an entity is, where it is, what it does, what else it is called, and how it relates to other entities. Google fills it from sources it trusts, from structured data on websites, and from patterns in what people search. It shows up as the knowledge panel on the right of a results page for a well-known name, and as direct answers to factual questions.
For Hermetic the question is whether Google has us as an entity at all, and the test is to search "Hermetic Networks" and see whether a panel appears with the right address and category. Three things feed that panel. The Google Business Profile (GBP), the listing at business.google.com, is the main one for a local company. Organization structured data on the homepage, a machine-readable block stating the company's name, logo, address and phone, is the second, and the homepage has none today: the crawl found WebSite, WebPage, BreadcrumbList, ImageObject and SearchAction blocks and no Organization or LocalBusiness. The third is the company's name appearing consistently on other sites. The word "hermetic" has an older meaning and a philosophical tradition attached to it, so the entity needs those signals more than a company called Richmond IT Partners would.
Vertical indexes are separate indexes Google keeps for one kind of content each, with their own crawling behavior and their own metadata: images, video, news, business listings, shopping, books, scholarly papers. A results page for an ordinary query blends results from several of them, which Google calls universal search. Four verticals and what they mean for us:
Images
Google Images indexes images by their file name, alt text and the text around them. Any Hermetic image with a descriptive alt text can show there. How many of ours have one is not known yet.
Maps
The business listings index. The local pack on a results page, the map with three businesses, is this index showing inside web results. It is fed by the GBP, and whether ours is claimed is not known yet.
News
Google News draws from publishers that meet its content policies. A company blog is normally outside it. We do not plan for it.
Video
Video results come mostly from YouTube. Hermetic has no videos known to this course, so the video block on a results page is one we concede for now.
Two of the four, the web index and the Maps index, carry nearly all of Hermetic's search traffic, which is why the Rank Math course spends a module on the GBP and the website agreeing with each other.
7The results page for a question
The search engine results page (SERP) is what Google assembles for a query, and its parts change with the kind of search. The Rank Math course drew the page for a service-plus-place search, where a local pack of three businesses sits near the top and the GBP does much of the work. This section draws the page for an informational search, one where the person wants an answer rather than a vendor, using the problem search from module 01: "which form of MFA is the most secure". The two pages share little beyond the organic results, and that difference decides which Hermetic asset is doing the work.
Walk down the page. Organic results are the unpaid listings, each a title link, an address and a snippet; Google builds the title and snippet from the page's title element, headings and content, and rewrites them when it judges its own version clearer. A featured snippet is an organic result promoted to a box at the top with a longer excerpt, shown when Google is confident the excerpt answers the question outright. People also ask is a block of related questions, each opening to an excerpt from some page, and it grows as the person clicks. Autocomplete is the list of predicted completions shown while typing, drawn from what people commonly search; the related searches at the foot of the page are the same idea after the fact. Ads carry the Sponsored label and are bought through Google Ads. The local pack is absent because nothing in the query is local, and the knowledge panel is absent because the query names no entity.
| Part of the page | Service plus place: "managed IT services Richmond VA" | Question: "which form of MFA is the most secure" |
|---|---|---|
| Local pack | The GBP. The website confirms it with LocalBusiness data. | Absent |
| Featured snippet | Rare | A blog post with a direct answer paragraph |
| People also ask | FAQ blocks on the service page | Posts and FAQ blocks on neighboring questions |
| Organic results | The one managed IT service page, with its title and description | The one post on the question, with its title and description |
| Knowledge panel | Absent, unless the query is our name | Absent |
| Ads | Google Ads, outside SEO | Google Ads, outside SEO |
| Autocomplete and related searches | Input to module 05 | Input to module 05 |
The asset that matters changes with the column. For the question search, the GBP does nothing and the plugin's title field does a little; the paragraph that answers the question in the first screen of a post does nearly everything. For the service search it is the other way around. Module 01's four kinds of search are, in practice, four different results pages, and knowing which one a query produces tells you which asset to work on.
8Two windows into the index
Nobody can read Google's index directly. Two tools show what it holds for one domain, and they disagree by design, so use both and know what each is good for.
Window 1: the site: search
- Open a private or incognito browser window, so that your own history and account do not shape the results.
- Type
site:hermeticnetworks.comand search. Write down the count Google shows above the results, with the word "about" if it uses it, and the date. - Scroll the first two pages and write down which URLs appear first. The order is Google's judgment of which of our pages matter most when no topic is given.
- Add a topic:
site:hermeticnetworks.com managed IT. The first result is the page Google prefers for that topic among ours. Repeat for backup and for VoIP, since each has duplicates. - Run the two inurl: checks from section 3 for the -old page and the thank-you pages.
The window is good for a rough size, the preferred page per topic, and whether a page that should be out is in. It is bad at the count. Google says the site: count is an estimate and the listing incomplete, and points site owners to Search Console for the real number. A count far below 392 is expected, since duplicates consolidate under one canonical and anything marked noindex is excluded; treat it as a question for the second window.
Window 2: Search Console's page indexing report
Google Search Console is Google's free reporting tool for site owners. It shows which searches show our pages, how often they are clicked, and, in the report that matters here, which URLs are in the index and why the rest are out. Whether a property for hermeticnetworks.com exists, and who owns it, is not known yet. Creating one, verifying ownership and connecting it to the plugin is Rank Math module 06. Once it exists, the steps are these.
- Sign in at
https://search.google.com/search-console/with the Google account that owns or has access to the property, and choose the hermeticnetworks.com property in the top left. - In the left menu, under Indexing, open Pages. The top of the report shows two numbers: how many URLs are indexed and how many are not, with a chart of both over time. Write both down with the date.
- Scroll to the table headed "Why pages aren't indexed". Each row is a reason and a count. Expect reasons such as "Excluded by 'noindex' tag", "Alternate page with proper canonical tag", "Duplicate without user-selected canonical", "Page with redirect" and "Crawled - currently not indexed". Write the list down with the counts.
- Click a reason to see the URLs behind it. "Duplicate without user-selected canonical" is where the managed IT, backup and VoIP duplicates appear if Google has grouped them, and the list tells you which URL it kept.
- For any single URL, paste it into the search bar at the top of Search Console. This is the URL Inspection tool. It says whether the URL is on Google, which canonical Google selected and which the page declared, when it was last crawled, and offers View crawled page, the HTML Googlebot received, and Request indexing, which asks Google to crawl the URL soon.
- Under Settings, open Crawl stats. It shows how many requests Googlebot made, the response codes it got, and whether the host had problems. A run of errors here is a blocked crawl.
Whether a Search Console property exists, who owns it, and what its page indexing report says. Whether Cloudflare's rules let Googlebot through unchallenged. Whether the Google Business Profile is claimed. Until the property is in hand, the site: survey in the next section is the only window open, and every reading from it is an estimate.
9Do this now
- Open a spreadsheet and name a sheet "Site survey". Columns: Date, Query, Count shown, First Hermetic URL, Position on page, Features present, Notes. Every row below goes in here; this sheet is reread in module 04 and compared against Search Console once that exists.
- In a private browser window, search
site:hermeticnetworks.com. Record the count Google shows. Under it, write the comparison in words: the sitemap offers 392 URLs, of which 105 are archives, leaving 287 pages and posts; write whether the count shown is above or below 287 and by roughly how much. Scroll two pages and record the first ten URLs in Notes. - Search
site:hermeticnetworks.com managed IT, then the same withbackupand withvoip. For each, record which of the duplicate URLs comes first. That is Google's current preference, and the Rank Math course's page map should have a reason before it overrules it. - Search
site:hermeticnetworks.com inurl:old,site:hermeticnetworks.com inurl:thankandsite:hermeticnetworks.com "phising". Record yes or no for each: is a page that should be out of Google in it. - Run the three test searches without site:, one row each: "managed IT services Richmond VA", "Hermetic Networks", and "which form of MFA is the most secure". For each, record whether a local pack, a featured snippet, a People also ask block and a knowledge panel appear; the first Hermetic URL if any; its position counted from the top of the organic results, or "not on page one"; and which part of the page it sits in.
- Find out whether a Search Console property exists for the domain. Ask Jeff, then the agency if needed. If it exists and you can get in, run the six steps of window 2 and paste the indexed and not-indexed counts and the reasons table into the sheet. If it does not, write "not known yet" in the sheet and the date you asked.
- Write one paragraph under the table that says what the survey means: how big Google thinks the site is against how big the sitemap says it is, which page Google favors for each duplicated service, whether anything that should be out is in, and which part of each results page our pages appear in or miss. This paragraph is the "Produces" of the module. Keep it; module 04 starts from it.
10Review questions
-
Answer
robots.txt is a fetch rule. Disallowing the path stops Googlebot reading the page, and a page Google knows about but cannot read can stay in the index with no description. The fix is a noindex tag on the page, set in the plugin, with the path left open in robots.txt so Googlebot can fetch the page and read the tag. Adding the Disallow as well makes it worse, because the tag is never seen.
-
Answer
The count is an estimate, and Google says so. Duplicates such as the three managed IT pages consolidate under one canonical, so each set counts once. Pages marked noindex are excluded by design. A fourth benign explanation: Google keeps only what it judges worth keeping, so thin pages may have been crawled and left out. The number to trust is the indexed count in Search Console's page indexing report.
-
Answer
Crawl budget is the wrong reason. Google's guidance puts crawl budget concerns at sites over a million pages, or over ten thousand pages changing daily, and Googlebot will get through 392 URLs. Cut the sitemap anyway, for a different reason: the sitemap is the list of pages we tell Google we care about, and 105 archives and two thank-you pages on that list send the wrong message about what the site is. The cost of those URLs falls at the index and the results page.
-
Answer
Paste each URL into URL Inspection in Search Console; the result shows the Google-selected canonical for each. The two alternates also appear under "Duplicate without user-selected canonical" in the page indexing report. To change it, redirect the two losers to the URL you want with a 301 and leave that page's rel="canonical" pointing to itself; a redirect is the strongest canonical signal Google accepts. The site: search gives an early hint, since the first result for site:hermeticnetworks.com managed IT is the page Google prefers today.
-
Answer
The GBP feeds none of that page. It feeds the Maps index, which shows as the local pack for queries with local intent and as a knowledge panel for a search of our name. A question about MFA has no place in it and names no entity, so neither appears. The parts on that page we can earn, the featured snippet, People also ask and the organic results, are fed by a blog post and its title and description.
-
Answer
Discovery fails: Googlebot does not click, so it never sees the link, and no sitemap entry names the URL. With no discovery there is no fetch, so robots.txt and the status code never come into it. With no fetch there is no index entry, and with no index entry the page cannot rank for anything. The fix is at the first step: a plain HTML link from a page Google already has, or a sitemap entry, or both.
11Glossary
- Query
- The text a person types or speaks into the search box. Google reads it loosely, treating plurals, misspellings and synonyms as the same request.
- Search operator
- A word with a colon or a symbol in front of a word that changes what a search returns rather than what it is about: site:, intitle:, inurl:, quotes, and the minus sign.
- Googlebot
- The program Google uses to fetch pages. It requests a URL, reads the status code and HTML, and renders the page in a recent version of Chrome.
- Crawling
- The step where Googlebot learns a URL exists, through a link, a sitemap or a submission, and fetches it.
- Crawl budget
- How many URLs on a site Googlebot can and wants to fetch in a period. Google says it matters for sites with over a million pages or over ten thousand pages changing daily.
- Sitemap
- An XML file listing the URLs a site wants crawled. Ours is sitemap_index.xml with 392 URLs.
- robots.txt
- A text file at the root of the site telling crawlers which paths they may fetch. A fetch rule only; it does not remove a page from the index.
- Index
- Google's stored, sorted copy of the pages it crawled and decided to keep, each filed under one canonical URL.
- Canonical URL
- The one address Google chooses to represent a group of duplicate or near-duplicate pages. Stated with a rel="canonical" link or, more strongly, a 301 redirect.
- noindex
- A robots meta tag or HTTP header telling Google to leave a page out of the index. Googlebot must be allowed to fetch the page to read it.
- Knowledge Graph
- Google's database of entities, such as people, places and organizations, and the facts and relationships between them. Feeds knowledge panels and direct answers.
- Vertical index
- A separate index for one kind of content: images, video, news, business listings. Universal search blends several into one results page.
- Featured snippet
- An organic result promoted to a box at the top of the results with a longer excerpt, shown when Google is confident the excerpt answers the question.
- People also ask
- A block of related questions on the results page, each opening to an excerpt from some page. Also a source of phrasings for keyword research.
- Page indexing report
- The Search Console report, under Indexing then Pages, that counts indexed and not-indexed URLs and lists the reason for each exclusion.
- URL Inspection
- The Search Console tool, reached by pasting a URL into its top bar, that shows whether a URL is on Google, its selected canonical, its last crawl and the HTML Googlebot received.
12Sources
- Google Search Central, In-depth guide to how Google Search works: the three stages, rendering with a recent Chrome, canonical selection, and that indexing is not guaranteed.
- Google, How results are automatically generated: the meaning of the query, including synonyms and spelling, comes before matching.
- Google Search Help, Refine web searches: quotation marks, the minus sign and site:.
- Google Search Central, Overview of Google Search operators: the operators Google documents, including intitle: and inurl:, and the retirement of cache:.
- Google Search Central, Learn about sitemaps: a site of about 500 pages or fewer with good internal links may not need one.
- Google Search Central, Large site owner's guide to managing your crawl budget: the million-page and ten-thousand-page thresholds.
- Google Search Central, Introduction to robots.txt and How Google interprets the robots.txt specification: a fetch rule, a disallowed URL can still be indexed, and Crawl-delay is ignored.
- Google Search Central, Block Search indexing with noindex: the meta tag and header, and the requirement that the page be crawlable.
- Google Search Central, How to specify a canonical with rel="canonical" and other methods: duplicate grouping, consolidation of signals, redirects as the strongest signal.
- Google Search Central, Mobile-first indexing best practices: the mobile version is the one indexed.
- Google Search Central, Understand JavaScript SEO basics: how Googlebot renders pages and why links need to be in the HTML.
- Google Search Central, Featured snippets and your website, Title links and Snippets: how each part of an organic result is built and may be rewritten.
- Google Search Central, Google Images SEO best practices: file names, alt text and surrounding text.
- Google, A reintroduction to Google's Knowledge Graph and knowledge panels: entities, facts and where panels come from.
- Google Search Help, How Google autocomplete predictions work.
- Search Console Help, Page indexing report and URL Inspection tool: the indexed and not-indexed counts, the reason names, the selected canonical and View crawled page. The report's help also states that site: counts are estimates.
- Google Search Central, Get started with Search Console.
- Stephan Spencer, Eric Enge and Jessie Stricchiola, The Art of SEO, 4th edition (O'Reilly, early release 2021), chapter 2.
- Site facts from a crawl of www.hermeticnetworks.com on 2026-09-29: robots.txt, sitemap_index.xml and its five child sitemaps (392 URLs), homepage source and structured data, and the duplicate service URLs.
Documentation read 2026-10-01. Google changes often; where this page and the documentation disagree, the documentation is current.