Detect “Discovered – currently not indexed” with the URL Inspection API
This page shows a Python script that takes a list of URLs from a property you own, queries Google’s URL Inspection API for each one, and prints the coverage state — so pages sitting in Discovered - currently not indexed or Crawled - currently not indexed surface as a flat CSV instead of one-by-one clicks in Search Console. Use it when a section of your site, or a batch of pages you point links at, stalls before the crawl stage.
One hard limit first: the API inspects only URLs inside properties verified in your own Search Console account. It does not accept third-party donor pages. For URLs outside your properties, the workable path is a bulk check against the live index — for example, verify the indexing status of a URL list — and that is a different mechanism: parsing what Google actually serves, not asking the API what it thinks.
Requirements
- A Search Console property (domain or URL-prefix) where the target URLs live.
- A Google Cloud service account added to that property as a user, with the Search Console API enabled.
- Python 3.10+, packages
google-api-python-clientandgoogle-auth. - The service-account JSON key, exposed through an environment variable — never hard-coded.
The script
Full file: scripts/inspect_url_status.py
# see scripts/inspect_url_status.py in this repository
The script reads urls.txt (one URL per line), calls urlInspection.index.inspect for each, and writes inspection_report.csv with three columns: URL, coverageState, lastCrawlTime.
How to run
- Export the key path:
export GOOGLE_APPLICATION_CREDENTIALS=~/keys/gsc-sa.json. The shell prints nothing — silence is success. - Put your URLs into
urls.txtand runpython inspect_url_status.py sc-domain:example.com. Expected console output per URL:https://example.com/page -> Discovered - currently not indexed | last crawl: none. - Open
inspection_report.csv. Rows wherecoverageStateisDiscovered - currently not indexedandlastCrawlTimeis empty are pages Google knows about but has never fetched. - Mind the quota: Google documents a limit of 2,000 inspection calls per property per day and 600 per minute. The script sleeps 150 ms between calls; a 10,000-URL list needs to be split across days or properties.
Limitations
- Only your verified properties. Donor pages on other people’s sites are out of scope by design.
- The verdict is Google’s internal state, which can lag the live index by days. A page can serve in search while the API still reports an older state, and the reverse.
Discovered - currently not indexedmeans the URL was never fetched;Crawled - currently not indexedmeans it was fetched and declined. The crawled but not indexed status needs content-side fixes, while a discovered-only page usually needs crawl paths: internal links, sitemap freshness, server response time.- The script does not request indexing. The Indexing API is a separate endpoint restricted to job postings and broadcast events, and this page does not pretend otherwise.
Further
- Official reference: Method: index.inspect — request body, response fields, quotas.
- Previous build in this series: Google index checker API after the Custom Search shutdown.