Skip to main content

Module content_type_shift

Module content_type_shift 

Source
Expand description

Content-Type shift detection (T104) — rolling baseline tracker.

Catches the publisher-cloaking pattern: HTML replaced with a Markdown stub for AI-bot User-Agents, same URL, same 200, parse succeeds, but a different edition of the page. Content-Type shift detection (T104) — rolling baseline tracker.

The 2026 scraping guide (docs/dev/project/scraping-guide-2026-llm-context.md §“POST-EXTRACTION”, L2536) calls out that publishers are starting to serve a deliberately different document to AI-bot User-Agents (HTML replaced with a Markdown stub, the same URL, the same 200, the same parse succeeding). Selector-based validators that only count fields don’t notice because the Markdown stub still has plenty of fields — they’re just different fields.

This module catches the publisher-cloaking pattern by tracking (Content-Type, byte_length) per identity and emitting a ContentTypeShiftReport when either dimension drifts past the configured threshold.

Structs§

ContentTypeObservation
One observation per scrape.
ContentTypeShiftReport
Aggregated shift report for one identity.
RollingBaselineDetector
Default rolling-baseline adapter.

Enums§

ContentTypeDrift
Detected drift between the most recent observation and the historical baseline.
ContentTypeError
Errors raised by ContentTypeShiftDetector::record.
MimeClass
Coarse MIME classification. Two responses count as “same class” only if they share the same variant — the String form (e.g. text/html; charset=utf-8) is normalised to the variant.

Traits§

ContentTypeShiftDetector
Port trait for content-type shift detection.

Functions§

current_unix_secs
Unix-seconds timestamp helper for callers that need to record the observation’s wall-clock time alongside the drift event.
drift_summary
Helper for the snapshot builder: format a drift report as a short human-readable string suitable for inclusion in a diagnostic summary line.