What Did Compression Destroy?
Measure coverage, factuality, temporal order and caveat survival across video compression levels.
A ladder with contracts is still promises until audited. Chapter 17 built the rungs; this chapter decides whether they hold. No new representation mechanism enters - the chapter’s own restraint is the point. What it earns instead is enforcement: a measured stopping rule answering, per item and per use, how far this particular information may be compressed before the next reduction destroys something the current use cannot afford to lose.
The formal backbone is information monotonicity. Chapter 17 derives upper rungs by omission from one unit set; this chapter states the commitment law:
GLANCE ⊆ BRIEF ⊆ SUMMARY ⊆ DETAIL ⊆ SOURCE
— not as literal strings but as informational commitments. A lower rung may omit a claim, omit evidence, omit background. It may not silently strengthen a claim, reverse a relation, strip a qualifier while presenting the remainder as unconditional, change a number, or convert uncertainty into certainty. The compression law:
Reduced resolution may reduce commitments; it must not mutate the commitments it retains.
Omit, do not invent — now with a subset relation enforcing it.
The profile, unchanged, at every rung
No new fidelity ontology. The Chapter 9 profile runs identically at DETAIL, SUMMARY, BRIEF, and GLANCE — eight dimensions each — producing a degradation curve rather than a compression ratio:
DETAIL all critical dimensions pass
SUMMARY minor evidence loss
BRIEF qualifier survival fails
GLANCE relation survives; qualification and uncertainty fail
The operating point follows mechanically: SUMMARY, though GLANCE is cheaper. Different dimensions fail at different rates — qualifiers and uncertainty typically first, topic and entities last — which is why the curve matters more than any single score and why a scalar would hide exactly the failure that decides safety.
Losses classify against the Chapter-17 contracts as acceptable loss (detail the rung never promised), contract debt (something the rung promised, absent), or hard violation (retained information turned misleading or unsupported — the subgroup qualifier dropped from “18% improvement, only in subgroup Z”). The compression stop is then mechanical: stop at the lowest rung whose required preservation contract passes — never “looks good enough.” And the book’s loss asymmetry recurs a fourth time (false skips, missed streams, false-known suppression, now destructive compression): over-retention costs attention; under-retention costs decisions. The two are never combined into one efficiency score.
The evaluator is itself an instrument that can fail
Before trusting any automatic judgement in the autopsy, the chapter validates the validator — because the 2026 evidence says it must. Mujahid, Wright and Augenstein (ACL 2026 long paper, pp. 31914–31933, verified via ACL Anthology) stress-tested six widely used reference-free factuality metrics on long-document summarisation across science-fiction, legal, and scientific datasets: seven meaning-preserving perturbations (paraphrase, simplification, synonym replacement, equivalent negations, vocabulary reduction, compression, source-text insertion), plus retrieval-context and claim-density analysis. Findings: inconsistent scores for semantically equivalent summaries, declining reliability on information-dense claims resembling many source parts, and no metric consistently maintaining factual alignment in long context. The chapter’s rule, mirroring RELATE’s philosophy:
Measurement earns permission — but only after the measurement itself has been validated for the distinction.
EXP-18 therefore opens with controlled pairs: meaning-preserving paraphrase, simpler wording, reordered content (must not trigger failure) versus qualifier removal, number change, relation reversal, uncertainty deletion, unsupported insertion (must trigger). An automatic metric that cannot separate those groups authorises nothing. And the non-monotonic-evaluator trap is named explicitly: SUMMARY 0.82, BRIEF 0.76, GLANCE 0.84 does not mean GLANCE regained faithfulness — it means the metric is blind to what disappeared. Metric monotonicity is not preservation monotonicity, and any stop rule driven by unvalidated scores is decorative rigour performing safety without providing it.
QEVA serves as auxiliary corroboration under its established fence: reference-free coverage/factuality/chronology against source video on appropriate video cases (800/200 MLVU(VS)-Eval base), never replacing the profile; where the two disagree, the disagreement is reported as instrument localisation. Ou & Lapata return as the generation-side warning — regenerative compression (each stage rewriting the last) versus source-linked progressive omission (each stage traceable to the unit set) — explaining why EXP-18 tests that distinction rather than assuming levels are levels. Two final fences: the profile is validated for these constructs (claim/relation/qualifier/numeric/uncertainty/provenance/addition/emphasis under progressive compression) — validated-for-this-construct never implies universally valid evaluator; and the lowest-safe-rung verdicts govern decision tasks until a learning-goal condition with delayed-transfer measures exists, because semantic preservation is not learning preservation.
The autopsy design
EXP-18 runs the exact EXP-17 ladders on the same frozen corpus: SOURCE→DETAIL→SUMMARY→BRIEF→GLANCE per candidate, recording at each transition what disappeared, what changed, what was added, which contracts failed, which profile dimensions moved, and whether the user’s decision would change. Conditions: A independent summaries; B progressive omission ladder; C ladder plus preservation contracts; D contracts plus fail-closed stop. The headline result is lowest safe rung, broken down by source and task type — the wonderfully non-uniform outcome the book wants (GLANCE safe here, SOURCE required there). Failure criteria: no differential degradation across rungs (instrument blind); independent summaries tie progressive ladders on contract passage (derivation unnecessary); the stop rule never triggers (thresholds decorative); automatic scores non-monotonic without attributable cause (evaluator invalid for authorisation). Artifacts: per-transition autopsy records, degradation curves, lowest-safe-rung tables, evaluator-validation results. What a positive result would not justify: delivery timing — with qualification, resolution, and stops now measured, the remaining question is when anything should appear, which is Chapter 19’s entire and only job.
We now know what qualifies, how much can safely survive, and where compression must stop. But we have still assumed that eligible information should appear immediately.
References
- Mujahid, Z.M., Wright, D. & Augenstein, I. (2026). Stress Testing Factual Consistency Metrics for Long-Document Summarization. Proc. ACL 2026 (long), pp. 31914–31933. DOI 10.18653/v1/2026.acl-long.1472. Verified via Anthology. Used: six metrics × seven perturbations × three long-form datasets; inconsistency under equivalence; info-density decline; no consistent long-context alignment. Licensed as the evaluator-validation requirement.
- Jung & Kim QEVA (auxiliary corroboration under Ch-13 fence); Ou & Lapata (generation-side warning); Ch 09 profile (reused unchanged); Ch 17 contracts + ladders (audited, not rebuilt).
Proposed experiment EXP-18: ladder autopsy with evaluator validation
Status: PROPOSED. Two phases: (1) evaluator validation on controlled meaning-preserving vs meaning-changing pairs (authorisation gate for phase 2); (2) per-transition autopsy across A–D with contract classification (acceptable/contract-debt/hard-violation), over/under-retention separation, lowest-safe-rung headline, non-monotonicity trap checks.