Glossary
Indexing terms, defined plainly
The vocabulary this site uses every day - crawling, indexing, IndexNow, the Google Indexing API, Search Console states - in plain language, cross-linked to the guides that go deeper.
A
Anchor text#
The visible words of a link that a reader can click. Search engines use anchor text as a hint about the target page's topic, so a link that says "sitemap guide" tells the engine more than one that says "read more". Internal anchors are under your control; external ones are not, which is why natural variety is expected.
B
BYOK (Bring Your Own Key)#
A SaaS model where the customer connects their own API credentials instead of sharing the vendor's. For indexing tools that means your Google service account or IndexNow key stays in your account, nothing is pooled across tenants, and quota is yours alone.
C
Canonical URL#
The address you declare as the master version of a page that exists under several URLs. Search engines treat it as a strong hint rather than a command, so conflicting signals such as internal links pointing at the duplicates can dilute it.
Cloaking#
Showing search engines content that differs from what people see. It is a spam policy violation in both Google's and Bing's guidelines, and it is unrelated to honest automation that only submits URLs faster.
Coverage report#
The Pages report in Google Search Console. It groups every known URL of a property into indexed, crawled-not-indexed, discovered-not-indexed and error states, which makes it the fastest way to see what indexing actually happened rather than what was submitted.
Crawl budget#
The amount of crawling a search engine is willing to do on a site in a given period. Large sites are affected the most: if the budget goes to parameter URLs or stale pages, fresh content waits longer for its first crawl.
Crawling#
The discovery step where a search engine's bot follows links and fetches pages. Crawling precedes indexing, and a page can be crawled for years without ever being indexed, which is why submission tools target the step after it.
D
Doorway pages#
Pages created mainly to rank for narrow queries and funnel visitors elsewhere, often interlinked in clusters. They are named explicitly in Google's spam policies, and no submission method can make them safe.
Duplicate content#
Substantial blocks of content reachable under more than one URL. It is not a penalty by itself, but it splits signals between the copies and slows indexation until a canonical or a redirect consolidates them.
G
Google Indexing API#
Google's official API for requesting indexing, documented for JobPosting and BroadcastEvent (livestream) pages. Sites outside those types call it at their own discretion: Google states support applies to the documented types, so results for ordinary pages are best treated as opportunistic.
Googlebot#
Google's web crawler. Its schedule is driven by crawl budget and page quality, not by how often a URL is submitted, and it identifies itself in server logs with a verifiable user agent.
I
Index bloat#
An indexed-page set that has grown beyond the pages that earn traffic: tag archives, parameter URLs, stale duplicates. Cleaning it concentrates crawl budget on the pages that matter.
Indexing#
The step where a fetched page is analysed and stored in the search engine's index, making it eligible to rank. Neither submission nor crawling guarantees it; coverage is only confirmed through Search Console or an equivalent.
IndexNow#
The open protocol supported by Bing, Yandex, Naver, Seznam and Yep for notifying participating engines about changed URLs. One key file and one endpoint cover all of them. Google is not a participant, so IndexNow alone does not reach Google Search.
J
JSON Web Token (JWT)#
The signed assertion a Google service account exchanges for an access token when calling the Indexing API. Most 403 errors trace back to a wrong scope, an unpropagated grant or clock skew, and all of them surface at the JWT step.
M
Manual action#
A human reviewer's penalty applied when a site violates a spam policy. It appears in Search Console with an example and a reconsideration path. Automated URL submission is not a violation and cannot trigger one by itself.
Meta robots tag#
A page-level directive, such as noindex or nofollow, inside the HTML head. It must be crawlable to be seen, so blocking a URL in robots.txt while trying to noindex it hides the instruction from the engine.
N
noindex#
A robots directive telling a search engine to keep a page out of its index. Removing it is only the first step of recovery: the page still needs a crawl and a positive indexing decision before it returns.
R
Rate limiting#
The per-time quotas an API enforces, such as the Google Indexing API's per-minute and per-day publish ceilings. Clients that ignore 429 responses make their own queues longer instead of faster.
Rendering#
The step where the engine executes JavaScript and lays out the page before indexing it. Client-rendered SPAs index later than static HTML because rendering is queued separately from crawling.
robots.txt#
The site-level text file that governs crawler access. It controls crawling, not indexing: an excluded URL can still appear in results on external signals alone, just without a fetched snippet.
S
Search Console#
Google's free console for property owners: coverage, crawl stats, URL inspection and manual actions. It is the source of truth for whether indexing happened, and every honest indexing workflow ends there.
Service account#
A Google Cloud identity with its own key file, used to authenticate Indexing API calls after it is added as a property owner in Search Console. The key is a secret: anyone holding it can act with the scopes granted to it.
SERP#
Search engine results page: the ranked set, AI answers and rich modules a query returns. Being indexed is the entry ticket, not the finish line.
Sitemap (XML)#
A machine-readable URL list that helps engines discover pages, including lastmod hints. Sitemaps aid discovery only; they carry no instruction to prioritise or to index.
Sitemap ping#
Notifying engines that a sitemap changed by hitting a ping endpoint. Google deprecated its ping endpoint in 2023 and now discovers sitemaps through robots.txt and Search Console, so automation moved to IndexNow or the Indexing API.
U
URL Inspection Tool#
Search Console's per-URL diagnostic: it shows the indexed version, the live version, and a Request Indexing button with conservative daily limits.
URL submission#
Any act of telling an engine that a URL exists and changed: sitemaps, IndexNow, the Indexing API, or Search Console's request button. A submission is a notification under quota, never a guarantee of indexing.
Want the terms put to work?
Indexer runs the whole loop these words describe: submit, audit, track.