Expand description
Content-Type shift detection (T104) — rolling baseline tracker.
Catches the publisher-cloaking pattern: HTML replaced with a Markdown stub for AI-bot User-Agents, same URL, same 200, parse succeeds, but a different edition of the page. Content-Type shift detection (T104) — rolling baseline tracker.
The 2026 scraping guide
(docs/dev/project/scraping-guide-2026-llm-context.md §“POST-EXTRACTION”,
L2536) calls out that publishers are starting to serve a deliberately
different document to AI-bot User-Agents (HTML replaced with a
Markdown stub, the same URL, the same 200, the same parse
succeeding). Selector-based validators that only count fields don’t
notice because the Markdown stub still has plenty of fields — they’re
just different fields.
This module catches the publisher-cloaking pattern by tracking
(Content-Type, byte_length) per identity and emitting a
ContentTypeShiftReport when either dimension drifts past the
configured threshold.
Structs§
- Content
Type Observation - One observation per scrape.
- Content
Type Shift Report - Aggregated shift report for one identity.
- Rolling
Baseline Detector - Default rolling-baseline adapter.
Enums§
- Content
Type Drift - Detected drift between the most recent observation and the historical baseline.
- Content
Type Error - Errors raised by
ContentTypeShiftDetector::record. - Mime
Class - Coarse MIME classification. Two responses count as “same class” only
if they share the same variant — the
Stringform (e.g.text/html; charset=utf-8) is normalised to the variant.
Traits§
- Content
Type Shift Detector - Port trait for content-type shift detection.
Functions§
- current_
unix_ secs - Unix-seconds timestamp helper for callers that need to record the observation’s wall-clock time alongside the drift event.
- drift_
summary - Helper for the snapshot builder: format a drift report as a short human-readable string suitable for inclusion in a diagnostic summary line.