Documentation

Methodology · own logs

How we compute bot crawl-waste

Last updated 2026-08-09

Crawl-waste rows come from parsing the access logs you upload. Each log line is streaming-parsed (Apache combined, Nginx, Cloudflare Logpush JSON, or CloudFront TSV), classified against a bot-UA table, and rolled up to per-URL × bot-class × day counts in ClickHouse.

The `is_orphan` flag is computed at read time: a URL that Googlebot has hit but our crawler has never seen is flagged, typically indicating a sitemap gap or a stranded page with no internal links.

Bot classes

We classify googlebot, bingbot, duckduckbot, yandexbot, applebot, and every published AI crawler (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Meta-ExternalAgent, CCBot). SEO crawlers (Ahrefs, Semrush, MJ12) are reported separately.

With `log parser evidence policy=true`, each hit's bot claim is forward-confirmed via reverse DNS + hostname suffix allowlist. Unverifiable hits are downgraded to `spoofed_<class>` so reports can filter them out.