How to Submit URLs to Google with Python and the Indexing API
This guide shows how to use google indexing api python code to submit URLs, check notification status, and handle errors without guesswork. It is for developers, SEO engineers, and site owners who already have a Google Cloud project and want a reliable script they can run from a laptop, a VPS, or a CI job. You will set up authentication with a service account, send URL_UPDATED and URL_DELETED notifications, read metadata, manage quotas, and turn a one off script into scheduled automation.
Key takeaways
- Python needs only two Google libraries plus your service account JSON key to publish notifications and read status.
- The API officially supports JobPosting and BroadcastEvent pages, so set expectations and monitor Search Console for other page types.
- Handle 403 as a permission task, 429 as a queue task, and JWT failures as a key or clock task, with retries only where they help.
- Log every request with timestamp, URL, status, and request ID, and keep daily volume safely below your Cloud quota.
- How google indexing api python submission works end to end
- What you need before you write code
- Installing dependencies and project layout
- Loading the service account and minting tokens
- Submitting a URL with URL_UPDATED
- Checking notification status with getMetadata
- Removing URLs with URL_DELETED
- Handling 403, 429, and JWT errors in Python
- Batching, queuing, and respecting quotas
- Logging, monitoring, and quota tracking
- From script to scheduled automation
- FAQ
<!-- IMAGE-PROMPT cover: 1200x630, DependsIt brand, deep charcoal #121212 or clean white background, vibrant mint #22E3B0 accent glow, thin node-network line art, Clash Display style bold heading space on left, General Sans clean labels, subject: Python code submitting URLs to Google Indexing API with service account key and terminal output, flat vector, high contrast, accessible, no photorealistic faces, no text smaller than 24px, no em dash in rendered text, export PNG then cwebp -q 82 to WEBP -->
How google indexing api python submission works end to end
A Python submission has four stages that repeat for every URL. First, your script loads the service account JSON key from disk or a secret store and builds credentials scoped to the indexing API. Second, it exchanges a signed JWT for a short lived OAuth access token, usually valid for one hour. Third, it POSTs a JSON body with the target URL and a notification type to the publish endpoint. Fourth, it reads the JSON response, logs the notifyTime and any error code, and optionally calls getMetadata later to confirm stored state. Once you see this cycle clearly, debugging gets easier because each failure maps to one stage.
The endpoint and notification types are fixed. You POST to indexing.googleapis.com/v3/urlNotifications:publish for new or updated pages with type URL_UPDATED, and for removed pages with type URL_DELETED. You GET from indexing.googleapis.com/v3/urlNotifications/metadata with a url parameter to read the latest notification for that URL. The API does not crawl the page synchronously or return index status. It records the hint and returns metadata about the notification itself. Actual crawl and indexing happen later through Google normal systems, which is why you must still monitor Search Console coverage and URL Inspection after submission.
Authentication in Python uses the google-auth and google-api-python-client or plain requests with a token. The google-auth library handles JWT signing, token caching, and refresh, so you rarely need to build JWTs by hand. You create a Credentials object from the JSON file, request a token, and attach it as a Bearer header. If you use the API client library, it wraps the HTTP layer for you. If you use requests directly, you manage headers and JSON yourself, which is useful for learning and for minimal dependencies on small hosts. Both paths use the same service account and the same scope, so choose based on how much control you want over retries and logging.
It is important to set scope expectations before you automate. The Indexing API is documented for JobPosting pages and BroadcastEvent pages inside BroadcastEvent or related livestream markup. That scope is enforced in docs and in how Google describes support, even if the endpoint technically accepts other URLs. If you submit job postings or livestream pages, you are inside documented use. If you submit blog posts or product pages, treat the call as a crawl hint with varied outcomes, keep sitemaps and internal links clean, and measure results rather than assuming instant indexing. Our setup prerequisite is covered in the service account setup guide, and the broader context is in the complete setup guide for faster indexing.
For planning, expect about thirty minutes to go from empty folder to first successful publish if your service account already has Search Console Owner access. If that access is missing, add it first, because no Python fix can bypass a 403 permission error. Keep your first test to one or two URLs you control, with valid structured data where applicable, self referencing canonicals, and no robots blocks. Log everything from the start, even for manual tests, so later automation inherits good observability instead of silent failures.
What you need before you write code
You need three things ready before Python can help. A Google Cloud project with the Indexing API enabled. A service account with a JSON key downloaded and stored securely. Search Console Owner delegation for that service account email on each property you will submit for. If any of these is missing, the script will fail with a clear but sometimes confusing error, such as 403 API not enabled or 403 permission denied. Confirm each item in the console UI before you open an editor, and record project ID, service account email, property URLs, and key creation date in a short runbook.
Your workstation needs Python 3.9 or newer, pip, and outbound HTTPS access to oauth2.googleapis.com and indexing.googleapis.com. Most VPS hosts and CI runners already allow this, but corporate proxies sometimes block token endpoints. Test with a simple HTTPS GET to the Google auth endpoint or a curl to the token URI. If you are behind a proxy, set HTTPS_PROXY and verify that Python requests picks it up. Also confirm your system clock is synchronized with NTP. JWTs include issued at and expiry timestamps, and a clock skew of more than a few minutes causes invalid grant errors that look like key problems but are actually time problems.
Prepare two or three test URLs that reflect real use. For documented use, use a job posting with JobPosting markup or a livestream page with BroadcastEvent markup, both passing Rich Results checks. For general testing, use a stable article or listing page you own, with a 200 status, indexable robots directives, and a clean canonical. Avoid faceted search URLs, session IDs, or staging domains for the first run, because those introduce variables that hide whether the Python flow itself works. Note each URL, its property, and its structured data type in your runbook so later log review is straightforward.
Decide where the JSON key will live during development and in production. During development, a file outside the repo with mode 600 is acceptable, referenced by an environment variable such as GOOGLE_APPLICATION_CREDENTIALS or a project specific variable like INDEXING_KEY_PATH. In production, prefer a secret manager or CI secret store, with the file written to a temp path at runtime or loaded into memory. Never commit the JSON to git, even in a private repo, and never paste it into chat or tickets. If a teammate needs access, grant them IAM or Search Console roles rather than sending the file. Key discipline at the start prevents rotation pain later.
Installing dependencies and project layout
Create a small, predictable project layout so later scheduling and logging are simple. A single folder with a virtual environment, one config file, one submit script, one status script, and a logs folder is enough for most sites. Larger teams can split into packages later, but start simple and keep paths explicit. The layout below works on Linux, macOS, and Windows with minimal changes to key paths.
mkdir -p indexer-python/logs indexer-python/urls
python3 -m venv indexer-python/.venv
source indexer-python/.venv/bin/activate
pip install --upgrade pip google-auth google-auth-httplib2 google-api-python-client requests
pip freeze > indexer-python/requirements.txt
The two core libraries are google-auth for credentials and token handling, and requests for direct HTTP calls. The google api python client is optional but useful if you prefer a discovery based client with built-in methods for publish and getMetadata. Installing requests alongside the client gives you flexibility. You can prototype with the client, then drop to python requests indexing api calls for finer retry control without reinstalling. Pin versions in requirements.txt so production runs the same code you tested locally. A typical pin set is stable for months, and you can update on a schedule rather than on every deploy. The same pip install google api step works on a laptop, a VPS, or CI, and most python seo scripts reuse this exact dependency set.
Keep configuration out of code. Use environment variables for key path, property root, and quota limits, with a small config module or .env file that is excluded from git. The example below shows the variables your scripts will read. Adjust paths for your host, and keep the URL list as a plain text file with one URL per line during early testing. Later you can source URLs from a sitemap, database query, or CMS webhook, but a text file makes the first runs easy to reason about.
export INDEXING_KEY_PATH="/etc/secrets/indexing-publisher-key.json"
export INDEXING_PROPERTY="https://example.com/jobs/"
export INDEXING_QUOTA_PER_DAY="150"
export PYTHONPATH="/opt/indexer-python"
Verify the install with a one line import check and a version print. This catches virtual environment mistakes before you debug API errors. Run the check with the same Python interpreter and the same service user that will run submissions, because permission and path issues often hide when you test as root but run as a restricted user. If imports fail, confirm the venv is activated and that pip installed into the expected site packages. If they succeed, you are ready to load credentials and mint a token.
import google.auth, requests, sys
print("python", sys.version.split()[0])
print("google-auth available, requests", requests.__version__)
Loading the service account and minting tokens
Credential loading is the foundation for every later call. Your script reads the JSON key, restricts it to the indexing scope, and requests an access token. The scope string is https://www.googleapis.com/auth/indexing and must match exactly. Using a broader scope or omitting scopes causes token errors or permission mismatches. The snippet below loads from a file path, refreshes the token, and prints the service account email plus a token prefix for verification. It does not call the Indexing API yet, so it isolates key and clock issues from property permission issues.
import os
from google.oauth2 import service_account
from google.auth.transport.requests import Request
KEY_PATH = os.environ.get("INDEXING_KEY_PATH", "/etc/secrets/indexing-publisher-key.json")
SCOPES = ["https://www.googleapis.com/auth/indexing"]
creds = service_account.Credentials.from_service_account_file(KEY_PATH, scopes=SCOPES)
creds.refresh(Request())
print("Service account:", creds.service_account_email)
print("Token prefix:", creds.token[:10], "expiry:", creds.expiry)
Run this on the same host and user that will submit URLs. If it prints an email and token prefix, your key path, file permissions, scope, and clock are correct. If it raises FileNotFoundError, fix INDEXING_KEY_PATH. If it raises PermissionError, fix file ownership to 600 for the runtime user. If it raises invalid grant, invalid JWT signature, or failed to parse JWT, check system time with NTP, re-download the JSON to rule out corruption, and confirm you did not edit the private_key newlines. The private_key field contains literal \n sequences that must stay intact. Opening the JSON in an editor that rewraps lines can break signing.
Token lifetime and reuse deserve attention. Access tokens typically last one hour. Minting a new token for every URL wastes time and can trigger token quota limits. A better pattern is to mint once per worker run and reuse the token until near expiry, then refresh. The google-auth Credentials object handles refresh automatically when used with an authorized session, but if you manage headers manually, check expiry before each batch and refresh when less than five minutes remains. For long running workers that process hundreds of URLs, refresh in a loop and log each refresh at debug level so you can correlate token events with API errors. Correct indexing api jwt python handling depends on intact private key newlines, exact scope, and synced clocks, so verify those three before changing code.
Security handling in code matters as much as file permissions. Never print the full token or private key to logs. Prefixes are enough for correlation. Never include the JSON content in exception messages that go to external error trackers without redaction. Load the key once at startup, keep it in memory, and avoid writing temp copies unless your secret store requires it. If you must write a temp file, use mode 600, delete on exit, and place it outside the web root. These habits keep your logs useful for debugging without turning them into a credential leak.
Submitting a URL with URL_UPDATED
Publishing a URL means telling Google that the page is new, updated, or live. For pages in documented scope, this is a JobPosting creation or a BroadcastEvent going live. For other pages, site owners use the same call as a hint, with varied results. The request is always a POST with JSON containing url and type. The url must be absolute, use https where possible, match the Search Console property you delegated, and return 200 when fetched without session tricks. The type for new or changed pages is URL_UPDATED. Keep the body minimal and the headers explicit.
The requests based snippet below publishes one URL and prints status plus response JSON. It builds an authorized session from your credentials, so token handling stays automatic. Replace the example URL with your own test page. Run it once, inspect the output, and save the full response with request ID in your log. A 200 with urlNotificationMetadata means the notification was accepted. It does not mean the page is indexed. It means the hint is recorded and will be considered by later crawl scheduling.
import os
from google.oauth2 import service_account
from google.auth.transport.requests import AuthorizedSession
KEY_PATH = os.environ.get("INDEXING_KEY_PATH", "/etc/secrets/indexing-publisher-key.json")
SCOPES = ["https://www.googleapis.com/auth/indexing"]
creds = service_account.Credentials.from_service_account_file(KEY_PATH, scopes=SCOPES)
session = AuthorizedSession(creds)
payload = {"url": "https://example.com/jobs/senior-support-specialist", "type": "URL_UPDATED"}
resp = session.post("https://indexing.googleapis.com/v3/urlNotifications:publish", json=payload, timeout=30)
print("HTTP", resp.status_code)
print(resp.text[:2000])
Validate inputs before you send. Strip whitespace, reject relative URLs, reject URLs with fragments that do not change content, and normalize trailing slashes according to your canonical rules. If your CMS generates both /jobs/role and /jobs/role/ as duplicates, submit only the canonical version. To python submit url to google reliably, keep a small indexing api python script that normalizes, validates, then posts one python urlnotification at a time with clear logging. Submitting variants wastes quota and can confuse reporting. Also check robots.txt and meta robots for the URL before submission.
For teams that prefer the discovery client, the equivalent call uses the urlNotifications publish method. It handles URL building and JSON serialization, which reduces typo risk in endpoint strings. The tradeoff is less visible HTTP control for retries. Both approaches are valid. Start with whichever you can read and log clearly, then standardize on one across your codebase so error handling stays consistent. Our broader troubleshooting reference for Indexing API errors including 403, 429, and JWT failures maps each status to the next fix regardless of client choice.
<!-- IMAGE-PROMPT diagram-01: 1600px max, DependsIt brand mint #22E3B0 on charcoal #121212 or white, node-network line art, subject: Python publish flow diagram showing JSON key then token mint then POST URL_UPDATED then 200 metadata response, flat vector, accessible, no em dash, Clash Display style headings, General Sans labels -->
Checking notification status with getMetadata
Reading notification status tells you what Google last recorded for a URL through this channel. Call the metadata endpoint with the URL as a query parameter, using the same authorized session. The response includes the latest update and notifyTime if a notification exists, or a 404 with a clear message if no notification is stored. This is useful for audits, for confirming that a queue worker actually sent an item, and for distinguishing never submitted from submitted but not yet crawled. It is not a ranking or index status check. For index status, use Search Console URL Inspection separately.
from urllib.parse import quote
url = "https://example.com/jobs/senior-support-specialist"
endpoint = "https://indexing.googleapis.com/v3/urlNotifications/metadata?url=" + quote(url, safe="")
r = session.get(endpoint, timeout=30)
print("HTTP", r.status_code)
print(r.text[:2000])
A typical stored notification response echoes the URL, the last type, and timestamps. Save notifyTime alongside your publish log to confirm the round trip. If publish returned 200 but metadata returns 404 shortly after, wait a few minutes and retry, as storage can lag under load. If metadata consistently returns 404 for URLs your worker claims to have sent, check that the worker used the same property scope and the same URL normalization. A trailing slash mismatch or http versus https difference creates a different key in storage, so publish and metadata appear to disagree when they are actually tracking two URL strings.
Use metadata reads sparingly. They consume a separate read quota that is usually generous but not infinite. A good pattern is to check metadata once per URL the day after publish, during a reporting job, rather than in a tight loop after every submit. For bulk audits, sample ten percent of URLs plus all failures, rather than reading every URL every hour. This keeps logs meaningful and leaves headroom for publish volume. Record metadata results in the same table as publish attempts, with columns for publish time, publish status, metadata time, and metadata type, so later analysis can join the two events without guessing.
Distinguish notification metadata from index coverage. A stored URL_UPDATED with a recent notifyTime means Google received your hint. It does not mean the page is indexed, ranked, or even crawled. Coverage in Search Console and server log evidence of Googlebot visits are the signals that show downstream effect. When you report to stakeholders, present both layers. Show submission success rate from your Python logs, then show crawl and index movement from Search Console. That separation prevents the common misunderstanding where a 200 from the API is presented as proof of indexing.
Removing URLs with URL_DELETED
When a job closes or a livestream page is retired, send URL_DELETED so Google can update its records for that URL through the same channel. The call shape matches publish, with type changed to URL_DELETED. Use it only for URLs you control that truly return 404 or 410, or that redirect to a relevant expired notice with noindex where appropriate. Do not send URL_DELETED for pages that moved permanently. For moves, update the new URL with URL_UPDATED and handle the old URL with a proper 301 plus sitemap cleanup, so signals consolidate rather than vanish.
payload = {"url": "https://example.com/jobs/closed-role-123", "type": "URL_DELETED"}
resp = session.post("https://indexing.googleapis.com/v3/urlNotifications:publish", json=payload, timeout=30)
print("HTTP", resp.status_code)
print(resp.text[:2000])
Confirm removal preconditions before you send. Fetch the URL as Googlebot would, check status code, robots, and canonical. A page that still returns 200 with indexable directives should not receive URL_DELETED, even if the job is filled, because the signal contradicts the page state. Update the page content to reflect closed status, add appropriate structured data changes such as validThrough in the past for JobPosting, and let the API notification match the on page reality. Mismatched signals cause confusing coverage states and extra recrawls that waste budget.
Batch deletions carefully. If you close one hundred roles after a hiring freeze, do not blast one hundred URL_DELETED calls in a minute. Pace them like updates, one every 30 to 60 seconds, within daily quota, with logging and backoff. Prioritize pages that still receive search impressions or that show as indexed with stale snippets. Low value expired pages with no traffic can wait for natural sitemap decay. This prioritization keeps quota focused on URLs where freshness visibly matters to users and to hiring managers checking search results.
Handling 403, 429, and JWT errors in Python
Error handling separates a demo script from production automation. The Indexing API uses standard HTTP statuses with JSON error bodies that include code, message, and status. Your Python code should branch on status, log the full body at debug level, and decide between retry, park for review, or fail fast. Retrying everything is a common mistake that turns permission problems into quota problems. The table below gives a practical mapping you can implement directly.
| Status | Meaning in this API | Python action | Retry |
|---|---|---|---|
| 200 | Notification accepted | Log notifyTime, mark done | No |
| 401 | Bad or expired token, wrong scope | Refresh token once, retry once | Once |
| 403 permission denied | Missing Search Console Owner | Park for manual review, alert | No auto retry |
| 403 API not enabled | Wrong project or disabled API | Fix Cloud enablement | No |
| 404 on metadata | No stored notification | Treat as never submitted | N/A |
| 429 | Quota or rate limit | Exponential backoff, pause queue | After delay |
| 5xx | Temporary server issue | Backoff and retry with jitter | Yes, limited |
Implement a small helper that POSTs with timeout, catches network errors, and returns a structured result. The snippet below shows the shape. It refreshes the token once on 401, backs off on 429 and 5xx, and parks 403 for human review. It sleeps between attempts with exponential delays and adds jitter to avoid thundering herds when multiple workers retry at once. Adjust base delays to your quota and volume, starting with 60 seconds for 429 and doubling to a cap.
import time, random, requests
def publish_with_retry(session, url, ntype="URL_UPDATED", max_tries=4):
payload = {"url": url, "type": ntype}
delay = 60
for attempt in range(1, max_tries + 1):
try:
r = session.post("https://indexing.googleapis.com/v3/urlNotifications:publish", json=payload, timeout=30)
except requests.RequestException as e:
print(f"network error attempt {attempt}: {e}")
time.sleep(delay + random.uniform(0, 10))
delay = min(delay * 2, 900)
continue
if r.status_code == 200:
return True, r.json()
if r.status_code == 401 and attempt == 1:
session.credentials.refresh(Request())
continue
if r.status_code in (429, 500, 502, 503):
print(f"retryable {r.status_code} attempt {attempt}, sleep {delay}")
time.sleep(delay + random.uniform(0, 10))
delay = min(delay * 2, 900)
continue
print(f"non retryable {r.status_code}: {r.text[:500]}")
return False, r.text
return False, "max tries reached"
JWT failures deserve special care because they happen before any HTTP to the Indexing API. Failed to parse JWT, invalid signature, and invalid grant usually mean a corrupted key file, edited newlines, wrong service account, or clock skew. Re-download the key, verify client_email matches the Search Console entry exactly, confirm the scope string, and sync time with NTP. Do not add broad IAM roles to fix a JWT error. Roles do not repair signing. Fix the key, the clock, and the scope first, then retest token mint in isolation before resuming publishes.
Batching, queuing, and respecting quotas
Bulk submission without a queue is how teams burn quota in an hour and then wait a day with nothing logged. A queue with pacing, persistence, and backoff turns bulk work into steady progress you can monitor. The simplest reliable queue is a SQLite table or CSV with columns for URL, type, attempts, next_retry, last_status, and last_response. A worker selects due rows, publishes one at a time with a sleep interval, updates the row, and handles 429 by pushing next_retry into the future. Persistence matters because a crash without persistence loses track of what was sent.
Pacing depends on your quota. If your Cloud console shows 200 publishes per day, plan for 150 to leave room for retries and manual tests. At one request every 60 seconds, 150 requests take about two and a half hours, well within a workday window. At one every 30 seconds, the same volume takes about seventy five minutes but leaves less margin for backoff. Start conservative, measure 429 rate, then tune. Never run multiple workers against the same quota without a shared counter or lock, because parallel loops multiply request rate and trigger limits faster than logs can explain. One worker with clear logs beats four workers with interleaved output.
Prioritize URLs so quota goes where it matters. For job boards, new postings and updated closing dates outrank evergreen category pages. For livestream sites, upcoming broadcasts outrank old replays. For general sites using the API as a hint, new or materially updated pages outrank minor typo fixes. Tag each queue row with priority and structured data type, then order selection by priority and age. This ensures that when quota runs out midday, the most time sensitive URLs already went through. Our quota explainer on quota limits and how to stay under them shows dashboards and alerts that complement this queue design.
Handle large backlogs with date windowing. If you have five thousand URLs and a quota of two hundred per day, do not load all five thousand as due today. Load two hundred per day in priority order, with the rest scheduled forward. Each morning, the worker picks up the next window plus any retries from yesterday. This keeps next_retry meaningful and prevents a growing pile of overdue rows that all look urgent. For historical context on pacing and recovery, see the guide to rate limits and recovery when you hit the ceiling.
Logging, monitoring, and quota tracking
Good logs make quota and error behavior visible without guessing. Log one JSON line per attempt with timestamp, URL, type, HTTP status, notifyTime on success, error code and message on failure, request ID if present, and worker hostname. JSON lines are easy to filter with jq, load into BigQuery or SQLite, and graph in a dashboard. Keep human readable console output minimal, one line per URL with status, while the JSON file captures full detail. Rotate logs daily and retain at least thirty days so you can compare week over week patterns.
Track two counters separately. Publish attempts versus publish successes show script health. Quota consumption from Cloud console shows account health. Reconcile them daily. If your log shows one hundred fifty successes but console shows higher consumption, look for other workers, manual tests, or duplicate projects using the same quota. If console shows fewer than your logs, check for requests that never left the host due to network blocks or proxy errors. To automate indexing python safely, run one paced worker with shared counters rather than parallel loops. A simple evening check of log count versus console count catches drift before it becomes a backlog surprise.
Alert on patterns, not on single failures. A single 403 for a new section likely means missing delegation for that prefix, worth a ticket but not a page. Ten 403s in a row likely means a property move or a revoked Owner, worth immediate attention. A single 429 near end of day is normal quota exhaustion. Repeated 429s at start of day suggest a runaway loop or a second worker you forgot about. Set thresholds such as more than five 403s in an hour or any 429 before noon, and route them to the owner of the runbook. Include recent log excerpts with redacted tokens so the responder can act without asking for more data.
Monitor downstream effect in Search Console, not just API responses. Track URL Inspection status for sampled URLs, coverage trends for submitted sections, and server logs for Googlebot fetches after notifyTime. A healthy pipeline shows 200s from the API, followed within days by bot visits and gradual coverage improvement for eligible pages. When you verify google index python submissions, compare notifyTime against later bot visits and coverage movement rather than treating a 200 as proof of indexing. If 200s stay flat with no downstream movement for weeks, revisit page quality, structured data validity, canonicals, and internal linking.
<!-- IMAGE-PROMPT workflow-02: 1600px max, DependsIt brand, subject: Python queue and monitoring workflow showing URL queue then paced worker then logging dashboard then Search Console verification, flat vector, accessible, no em dash, mint #22E3B0 on charcoal #121212, node-network line art, Clash Display style headings, General Sans labels -->
From script to scheduled automation
Once manual runs are stable, automate on a schedule that matches how often your content changes. Job boards with daily postings suit a twice daily cron. News style livestream pages suit an hourly check with a small batch. General blogs suit a post publish hook plus a nightly catch up for missed items. Start with cron on a single host because it is simple, visible, and easy to log. Move to systemd timers, Cloud Scheduler, or CI schedules only when you need retries across hosts or tighter integration with deploys.
A cron entry that runs twice daily, locks against overlap, and logs to a dated file is enough for many teams. Use flock to prevent overlapping runs when a previous batch runs long. Activate the venv explicitly, set required environment variables, and append output to logs. Test the exact command as the runtime user, not as root, to catch permission differences early. Keep the schedule in version control alongside the script so changes are reviewed like code.
0 8,16 * * * /usr/bin/flock -n /tmp/indexer-python.lock /opt/indexer-python/.venv/bin/python /opt/indexer-python/submit_queue.py >> /opt/indexer-python/logs/cron.log 2>&1
Evolve toward event driven triggers when polling becomes wasteful. For example, hook your CMS publish action to append the new URL to the queue file, then let the scheduled worker drain the queue with pacing. This gives you near real time intake with quota safe output. Avoid calling the API directly from the web request path, because a slow token or API response will slow page saves and a traffic spike can burst quota. Queue intake should be fast and local. API output should be paced and logged. That separation keeps authoring snappy and submissions reliable.
Before you scale, revisit facts and limits. Google does not support IndexNow, so this Python flow covers Google only. For Bing, Yandex, Naver, and Seznam, run IndexNow in parallel with its own key file and endpoint. Keep each system quota separate, with separate logs and dashboards. For implementation details in other stacks, share the runbook with teammates who use the Node.js developer workflow or prefer testing with cURL and Postman first. One shared queue schema plus per stack workers is a practical pattern for mixed teams. For scope and auth background, the official references at Indexing API prerequisites and using the Indexing API remain the source of truth.
FAQ
How many URLs can I submit per day from Python?
Your Cloud project quota decides, not your code. Many new projects start near 200 publish requests per day, with separate read limits for metadata. Check Enabled APIs then Quotas for your exact project and plan for about 75 percent of that number to leave room for retries and manual tests. When you automate indexing python workflows, pace with sleeps of 30 to 60 seconds, process in priority order with new jobs first, and window large backlogs across days. Log every attempt with timestamp and status, reconcile log counts against console graphs each evening, and pause the queue when 429 appears rather than retrying in a tight loop.
Why do I get 403 permission denied in Python but token mint works?
Token mint proves your key, scope, and clock are correct, including indexing api jwt python signing. Publish permission is checked separately in Search Console, so the service account email must be an Owner on the exact property that contains the URL. Open Users and permissions for that property, confirm the exact email and Owner role, and watch for prefix versus domain mismatches that cause most of these cases. Also confirm the API is enabled in the same project that owns the key. Park the URL without retry until delegation is fixed, then retry once with a fresh token and the exact canonical URL string.
Should I use google api python client or plain requests?
Both paths power a reliable indexing api python script, so choose based on readability and control. The google api python client reduces endpoint typos and handles discovery for publish and metadata, which helps quick prototypes and small teams. Plain python requests indexing api calls give clearer control over headers, timeouts, and retry logic, which helps production workers that must respect quotas and log uniformly. Start with whichever you read more easily, log timestamp, URL, type, status, and notifyTime either way, and standardize on one across your codebase. Keep most python seo scripts on one pattern so error handling for 403 and 429 stays consistent.
How do I avoid creating duplicate notifications for the same URL?
Normalize URLs before queuing and submit only the canonical variant. Strip whitespace, enforce https, apply your trailing slash rule, and drop fragments that do not change content. Store a hash of normalized URL plus python urlnotification type in your queue and skip inserts that are pending or recently succeeded. This saves quota and keeps reporting clean, especially when CMS hooks and nightly catch ups can enqueue the same link twice. For teams that python submit url to google from several sources, centralize normalization in one helper used by publish, metadata, and delete paths, then audit the queue weekly for near duplicates with different casing or slash forms.
Can I submit normal blog posts and product pages with this google index python script?
The endpoint technically accepts any URL you have rights to, but Google documents support for JobPosting and BroadcastEvent pages. Owners do submit other types as a hint with varied outcomes, so set expectations before you scale. If you run google index python checks for articles or products, monitor Search Console coverage and server logs for bot visits after notifyTime, keep structured data valid where applicable, and keep sitemaps, canonicals, and internal links clean. Do not present a 200 response as proof of indexing. Treat it as an accepted hint, compare submission dates against later crawl movement, and prioritize documented page types when quota is tight.
How should I schedule the indexing api python script in production?
Start with cron twice daily and a lock file to prevent overlap, with the pip install google api dependencies pinned in requirements.txt. Append new URLs to a persistent queue from CMS hooks, then let the worker drain with pacing, exponential backoff, and JSON logging. Move to systemd timers or Cloud Scheduler when you need better observability, dependency control, or cross host retries. Always run as a restricted service user with the key at mode 600, activate the venv explicitly, and keep the schedule in version control. Alert on repeated 403s or early day 429s, and keep cURL spot checks handy for incident reproduction.
Sources
- https://developers.google.com/search/apis/indexing-api/v3/prereqs
- https://developers.google.com/search/apis/indexing-api/v3/quickstart
- https://developers.google.com/search/apis/indexing-api/v3/using-api
- https://support.google.com/webmasters/answer/9008080
- https://www.indexnow.org/documentation