Detect “Discovered – currently not indexed” with the URL Inspection API

This page shows a Python script that takes a list of URLs from a property you own, queries Google’s URL Inspection API for each one, and prints the coverage state — so pages sitting in Discovered - currently not indexed or Crawled - currently not indexed surface as a flat CSV instead of one-by-one clicks in Search Console. Use it when a section of your site, or a batch of pages you point links at, stalls before the crawl stage.

One hard limit first: the API inspects only URLs inside properties verified in your own Search Console account. It does not accept third-party donor pages. For URLs outside your properties, the workable path is a bulk check against the live index — for example, verify the indexing status of a URL list — and that is a different mechanism: parsing what Google actually serves, not asking the API what it thinks.

Requirements

The script

Full file: scripts/inspect_url_status.py

# see scripts/inspect_url_status.py in this repository

The script reads urls.txt (one URL per line), calls urlInspection.index.inspect for each, and writes inspection_report.csv with three columns: URL, coverageState, lastCrawlTime.

How to run

  1. Export the key path: export GOOGLE_APPLICATION_CREDENTIALS=~/keys/gsc-sa.json. The shell prints nothing — silence is success.
  2. Put your URLs into urls.txt and run python inspect_url_status.py sc-domain:example.com. Expected console output per URL: https://example.com/page -> Discovered - currently not indexed | last crawl: none.
  3. Open inspection_report.csv. Rows where coverageState is Discovered - currently not indexed and lastCrawlTime is empty are pages Google knows about but has never fetched.
  4. Mind the quota: Google documents a limit of 2,000 inspection calls per property per day and 600 per minute. The script sleeps 150 ms between calls; a 10,000-URL list needs to be split across days or properties.

Limitations

Further