← Language From First Principles

What Survived the Transformation?

Build a preservation profile for every generated representation.

Part I has spent five chapters transforming information. Visual previews compress pages into triage cards. The router adds diagrams, tables, and timelines — or declines. The semantic record stabilises structure. The sidecar brings in surrounding material. Relation readout classifies how that material bears on the source. Every one of these operations can fail silently: the output looks fluent, looks professional, looks right, while the content has shifted underneath. A diagram drops the exception. A summary upgrades a correlation to a cause. A sidecar panel restates the source’s speculative conclusion as its headline finding. Nothing in the rendering announces the damage.

This chapter builds the instrument that announces it. Not a new transformation — the last mechanism of Part I is a measuring device — and its question covers everything the Part has done:

What did we lose, distort, add, or overstate while doing all of that?

The chapter earns two things: a multidimensional preservation object that scores every transformation type in the book, and the empirical machinery beneath the sentence that has hung over the manuscript since Chapter 1: every transformation creates an obligation to measure what survived.

A vector, not a score

The preservation object is a profile, never a scalar:

PreservationProfile {
    claim_survival
    relation_survival
    qualifier_survival
    numerical_survival
    uncertainty_survival
    provenance_survival
    unsupported_additions
    emphasis_shift
}

The fields may evolve — Chapter 18 will stress-test them against video, Part IV against proxy reports — but the structural point is permanent: different transformations fail differently, and one number hides the difference. A diagram preserves topology while losing caveats. A summary preserves claims while weakening uncertainty. A visual preview preserves topic while exaggerating emphasis. A relation readout preserves entities while reversing direction. fidelity = 0.84 describes none of these; it averages them into reassurance. The profile exists so that a specific failure — qualifier survival at 0.3 while claim survival sits at 0.9 — is visible, attributable, and fixable. Any evaluation in this book that reports a single fidelity number without the profile behind it is, from here on, methodologically out of bounds.

Lost, added, changed

Every profile entry decomposes into three failure classes, and the triad recurs everywhere the book compresses — video levels, radar digests, proxy consequence summaries:

lost / added / changed

SOURCE FACT PRESENT
→ missing from derived view
= omission (lost)

SOURCE DOES NOT SUPPORT CLAIM
→ claim appears in derived view
= addition (added)

SOURCE CLAIM PRESENT
→ derived representation makes it stronger, weaker,
  more certain, differently scoped
= transformation (changed)

Omission and invention are different engineering problems — one is fixed by coverage, the other by constraint — and conflating them produces fixes for the wrong disease. The third class matters most for AI-mediated rendering because it is the most characteristic: the claim survives as words while its force changes. A “may” becomes an “is.” A “in subgroup Z” detaches from its sentence and generalises. An association acquires an arrow. Nothing was lost and nothing was fabricated from nothing; the content was edited by transformation, which is precisely what fluent generation does best and what scalar metrics punish least.

Relations are first-class

This is Chapter 8’s payoff arriving on schedule. A representation can preserve every noun while destroying the assertion:

SOURCE:  A causes B only when C is present.
VIEW:    A causes B.

Topic preserved. Entities preserved. Relation partially preserved. Qualifier destroyed. The statement’s practical content — reversed, in any decision context where C is absent.

SOURCE:  A acquired B.
VIEW:    B acquired A.

Nearly all tokens survive. The relationship does not — and with it goes everything a reader would act on.

Relation survival is therefore a central measure, not a subfield of claim fidelity. The profile scores relation direction, relation type, and attached qualifiers as their own entries, using the adversarial constructions of Chapter 6 (conditioned causation, rejected causation, scoped exception) as permanent regression cases. This is also where DiagramEval’s lesson hardens into method: node alignment without path alignment is entity preservation without relation preservation, and any diagram evaluation that reports only the former is measuring the boxes while the arrows rot. The profile refuses that bargain structurally.

Distortion by emphasis

A subtler failure needs its own field, because it passes every fact-check while changing the communication. Suppose a paper is, by weight of content:

70% method
20% limitation
10% speculative conclusion

and the generated visual devotes 80% of its area, colour, and motion to the speculative conclusion. No statement need be false. The reader nevertheless leaves believing the speculation is the paper. Nothing was lost, added, or individually changed — prominence was redistributed.

A representation can distort by changing prominence even when every individual statement remains technically defensible.

Emphasis shift is operationalised comparatively: the distribution of attention-weight across the source’s components (by tokens, sections, or evidential role) against the distribution across the derived view (by area, position, ordering, repetition). Large divergences flag editorialising-by-layout — the spin study of Chapter 4 (a third of infographics spinning non-significant primaries, concentrated in results presentation) is emphasis distortion with prevalence data. The concept earns its keep far beyond Part I: personalised interfaces, news sidecars, recommendations, filtering, and the Personal AI all redistribute prominence as their core operation, and each will be scored against this field. A system that never states anything false but consistently promotes the engaging over the important is not faithful. From here, the book can say exactly why.

Traceability is separate from preservation

A view can be imperfect yet safe, or fluent yet dangerous, depending on one independent property: whether its elements point back to source evidence.

CONTENT PRESERVATION       TRACEABILITY
did the content survive?   can I walk it back?
claim/relation/qualifier   generated claim → record entry
  survival scores             → source span → document → evidence
generated claim → source span → original document → evidence reference

A trace does not make the claim correct — a perfectly sourced falsehood is still false — but it makes correction and inspection possible, which is the book’s agency philosophy in measurement form. Every profile therefore reports provenance survival alongside content scores: what fraction of derived statements resolve to record entries with source spans, and what fraction of those spans actually support the statement. An untraceable view fails closed regardless of its apparent quality; a traced view with content gaps is repairable. The escape hatch of Chapter 3, the spans of Chapter 6, and the contestability requirements of Part III all meet in this field.

Fidelity is not truth

One separation guards the whole instrument against its most dangerous misreading. Suppose the source states something false, and the generated view reproduces it faithfully. Preservation is high. That does not make the claim true. Three distinct judgements, three owners:

SOURCE FIDELITY        What did the transformation preserve?
                       Owned by THIS chapter.

EVIDENTIAL VALIDITY    Was the source claim well supported?
                       Assisted by Chapter 7's evidential sidecar.

TRUTH                  What is actually the case?
                       Owned by no pipeline in this book.

Preservation must never quietly become verification. A high-profile view that faithfully transmits a badly supported claim is working correctly as a transformation and failing as an information environment — and the book needs both sentences simultaneously, which only separate instruments allow. The Hallucination-book inheritance (evidence, truth, containment, acceptance as distinct stages) arrives here as method: this chapter measures containment of source content; acceptance of claims as believable belongs elsewhere entirely.

The climax experiment: one instrument, every transformation

EXP-09 is what makes this chapter the Part’s climax rather than another evaluation proposal. A single controlled source set — ordinary positive claims, negations, directed relations, temporal statements, exact numbers, qualifiers, exceptions, uncertainty markers, causal/non-causal distinctions, evidential citations — is run through every Part I transformation: visual compression, semantic enhancement, record-mediated rendering, sidecar enrichment, relation readout. Each output is scored with the same profile. The result is not five evaluations but one reusable instrument demonstrated across transformation classes — and the interesting findings will be comparative: the AI diagram preserves relationships better than prose summarisation but loses qualifications more often; the preview preserves topic while shifting emphasis; the relation readout preserves entities while straining on direction. Differential failure signatures, not leaderboard numbers.

Two controls keep the comparisons honest. A copy / minimally reformatted source condition sets the ceiling — what perfect preservation scores — so every transformation’s cost is visible against doing nothing. A human-written short summary sets the reference for competent compression, so AI systems are beaten by, or beat, genuine craft rather than straw. Without these, the experiment measures deviation from an ideal nobody attains; with them, it measures the price of each transformation in the currency the profile defines.

What this chapter earned

AI-mediated representations are to be evaluated against explicit preservation dimensions — never fluency, visual quality, semantic similarity, or preference alone. The profile (claims, relations, qualifiers, numbers, uncertainty, provenance, additions, emphasis) with its lost/added/changed decomposition, first-class relation survival, emphasis-shift detection, and independent traceability is the book’s standing obligation made operational: every transformation creates an obligation to measure what survived — now with machinery underneath the sentence. Fidelity stays fenced from validity and truth. And the instrument is reusable by design: Chapter 18 inherits it for video rather than reinventing it.

Part I can now compose representation, related information, relation-aware discovery and fidelity into a semantic browser. One standing rule travels with the instrument: the preservation profile is an experimental instrument, not ground truth — it authorises measurement only through the validation pattern of remove-and-restore and controlled perturbation (instrument → validation → authorised measurement, never metric → truth). Chapter 18 demonstrates the full gate; later evaluators inherit the obligation.

References

  • The author’s Embeddings From First Principles (preservation-profile methodology) and Hallucination From First Principles (evidence/truth/containment/acceptance separation): conceptual inheritance; mechanism-level verification against sibling sources scheduled at the Part I coherence pass.
  • Visual-abstract spin and RIVA-C evidence (Ch 04 base): reused as emphasis-distortion prevalence data; no new claims taken.
  • DiagramEval node/path alignment (Ch 05/06 base): reused as the relation-survival instance check for diagrams.
  • Long-form summarisation factuality and qualification-loss literature: leads for the Part I coherence pass to deepen the uncertainty-survival and qualifier-survival fields; no specific results cited as established here.

Proposed experiment EXP-09: preservation across all Part I transformations

Status: PROPOSED. Corpus: controlled sources seeded with the ten construction types (positive claim, negation, directed relation, temporal statement, exact number, qualifier, exception, uncertainty, causal/non-causal pair, evidential citation) plus adversarial Chapter-6 cases. Conditions: visual compression, semantic enhancement (router incl. UNCHANGED), record-mediated rendering (pipeline B), sidecar enrichment items, relation readout labels — each scored with the full profile, plus COPY and human-summary controls. Human adjudication per field; emphasis shift via source-vs-view attention-weight distributions; traceability via span-resolution rates. Remove-and-restore tests: delete each decisive element in turn and confirm the profile registers the specific loss (validates the instrument discriminates rather than merely scores). Failure criteria: profile fails to distinguish the seeded distortions (instrument invalid — the chapter’s claim falls); all transformations score identically (no differential signatures); COPY control itself scores imperfectly on traceability (scoring procedure broken). Artifacts expected: seeded corpus, ledger keys, per-transformation profiles, differential-failure matrix. What a positive result would not justify: evidential validity or truth of any source claim — fidelity only, per the fenced separation.