Methodology · own logs
How we compute bot crawl-waste
Last updated 2026-08-09
Crawl-waste rows come from parsing the access logs you upload. Each log line is streaming-parsed (Apache combined, Nginx, Cloudflare Logpush JSON, or CloudFront TSV), classified against a bot-UA table, and rolled up to per-URL × bot-class × day counts in ClickHouse.
The `is_orphan` flag is computed at read time: a URL that Googlebot has hit but our crawler has never seen is flagged, typically indicating a sitemap gap or a stranded page with no internal links.
Bot classes
We classify googlebot, bingbot, duckduckbot, yandexbot, applebot, and every published AI crawler (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Meta-ExternalAgent, CCBot). SEO crawlers (Ahrefs, Semrush, MJ12) are reported separately.
With `log parser evidence policy=true`, each hit's bot claim is forward-confirmed via reverse DNS + hostname suffix allowlist. Unverifiable hits are downgraded to `spoofed_<class>` so reports can filter them out.