29: How to Fail Correctly
By the time we reached Chapter 28, the Digital Crystal had accumulated an uncomfortable history.
Several experiments had looked convincing.
Some had even produced large effects.
Then a stronger control arrived.
Or an implementation defect was discovered.
Or the original question turned out to contain two different questions.
Or a result remained directionally clear but failed the predeclared magnitude requirement.
The easy response would have been to describe all of these as failures.
That would have been wrong.
A failed implementation is not the same thing as evidence against a hypothesis.
An unresolved result is not the same thing as a negative result.
A precise negative result is not the same thing as nothing happening.
A stronger control that narrows an interpretation does not erase the lower-level phenomenon that survived.
Those distinctions had become important enough that they could no longer remain informal.
So Chapter 29 stopped asking another question about the crystal.
Instead, it asked a question about the investigation itself:
When an experiment appears to fail, can we determine exactly what failed without silently changing the question, erasing surviving evidence, or converting an invalid run into a negative result?
This chapter is about epistemic bookkeeping.
Not as administration.
As science.
Failure Is Not One Thing
The word failure compresses several very different situations.
Suppose an experiment does not support the claim we wanted.
At least four possibilities immediately exist.
The implementation may have been invalid.
The experiment may have been valid but too imprecise.
The experiment may have been valid and precise enough to bound the effect below a meaningful threshold.
Or the lower-level result may still be supported while a stronger interpretation fails under a better control.
Those are not stylistic differences.
They imply different things about what we are allowed to conclude.
The working taxonomy became:
INVALID
UNRESOLVED
BOUNDED NEGATIVE
SUPPORTED BUT NARROWED
Later we also needed:
DESCRIPTIVE CLOSEOUT
because a post hoc mechanism can explain a result without changing its confirmatory status.
The core rule was simple:
FAILURE OF A CLAIM โ FAILURE OF AN EXPERIMENT โ ABSENCE OF A PHENOMENON
Chapter 29 tested whether our own recent experiment history actually respected that distinction.
flowchart LR
subgraph FailureClasses
A[INVALID]
B[UNRESOLVED]
C[BOUNDED NEGATIVE]
D[SUPPORTED BUT NARROWED]
E[DESCRIPTIVE CLOSEOUT]
end
A -.->|not evidence against| F[Claim]
B -.->|insufficient precision| F
C -.->|effect bounded below SEI| F
D -.->|stronger control narrows| F
E -.->|explains only| F
Building the Failure Ledger
The audit registered ten cases from Chapters 26 through 28.
Each case recorded:
- the claim being tested,
- whether the run itself was valid,
- the inferential status,
- the kind of transition that occurred,
- the role of the evidence,
- the relevant threshold where one existed,
- the expected mean, interval and MDE where available,
- and, critically, what evidence survived.
The registered validity categories were:
VALID
INVALID_IMPLEMENTATION
INVALID_CONSTRUCT
INVALID_REFERENCE
UNKNOWN
The inferential statuses were:
SUPPORTED
BOUNDED_BELOW_SEI
BOUNDED_NEAR_ZERO
UNRESOLVED
DIRECTION_SUPPORTED
DESCRIPTIVE_ONLY
INVALID
UNTESTED
And the transitions included:
IMPLEMENTATION_INVALIDATION
CONSTRUCT_NARROWING
PRECISION_RESOLUTION
PRECISION_LIMIT
MECHANISTIC_DECOMPOSITION
DESCRIPTIVE_CLOSEOUT
REPLICATION
CONTROL_STRENGTHENING
The point was not to create vocabulary for its own sake.
The point was to make illegal moves visible.
The Forbidden Moves
The Chapter 29 protocol explicitly banned six common forms of epistemic slippage.
We could not:
count INVALID as evidence against the hypothesis
We could not:
count CI-crossing-zero as FAILED
We could not:
call a result bounded negative without an explicit meaningful threshold
We could not:
erase valid sub-results when a larger inference became invalid
We could not:
promote a descriptive closeout into confirmatory rescue
And we could not:
silently change the estimand or control after seeing the result
This is less glamorous than generating a strange new Digital Crystal behavior.
It is also more important.
Because without these rules, every later claim can be rewritten after the fact until the project appears cleaner than it really was.
flowchart LR
subgraph Forbidden
A[INVALID as negative evidence]
B[CI crosses zero as FAILED]
C[Bounded negative without SEI]
D[Erase valid sub-results]
E[Descriptive closeout as confirmatory rescue]
F[Post-hoc change of estimand/control]
end
A --> X[Epistemic error]
B --> X
C --> X
D --> X
E --> X
F --> X
Chapter 26: When the Reference Is Wrong
Chapter 26 began with a plausible question.
If finite computation forces the system to subsample its frontier, does that amplify the downstream causal consequence of a local perturbation?
The first design tried to match background construction rate between bounded and effectively exhaustive evaluation.
But the matching was incomplete.
The calibration was fixed from an early condition rather than dynamically matched across the process.
Worse, the supposedly exhaustive f = 1 reference was not actually the same thing as true unbounded evaluation.
The comparison was therefore not the comparison the experiment claimed to make.
That made the V1 primary:
INVALID_REFERENCE
The correct conclusion was not:
no causal amplification exists.
The correct conclusion was:
this experiment does not establish the answer.
That distinction matters.
An invalid reference arm provides no clean negative evidence against the hypothesis.
The failure belongs to the experiment design.
Not to the phenomenon.
The lesson became:
A controlled initial condition is not a controlled process.
Correcting Chapter 26
V2 rebuilt the comparison properly.
The PREVENT branch was dynamically matched at every lag.
The exhaustive comparison became genuinely unbounded.
The frozen horizon remained twelve steps.
The smallest meaningful difference was:
So the result was not merely nonsignificant.
It was precision-resolved.
Strong candidate subsampling did not produce a mean twelve-step causal consequence differing meaningfully from exhaustive evaluation under the matched-rate protocol.
That became:
BOUNDED_NEAR_ZERO
But something else survived.
A Negative Aggregate Can Hide a Changed Mechanism
The Chapter 26 mechanism audit decomposed the immediate causal effect.
Under strong subsampling, much of the local effect came from:
force-only opportunities
Under exhaustive evaluation, more of it came from:
shared probability shifts
The aggregate consequence stayed matched.
The pathway changed.
That produced another construct separation:
CAUSAL ROUTING โ CAUSAL AMPLIFICATION
This is exactly the kind of evidence that gets lost if we summarize Chapter 26 as:
the experiment failed.
It did not fail.
The first design was invalid.
The corrected primary bounded amplification near zero.
And the mechanism audit revealed a genuine change in causal routing.
Three different epistemic outcomes existed inside one chapter.
The ledger had to preserve all three.
Chapter 27: When the Intervention Itself Is Wrong
Chapter 27 made the distinction even sharper.
V1 tested whether hidden material state could redirect future causal response.
The design compared:
FORCE
against:
PREVENT
But the PREVENT branch contained a defect.
At lag 1, the supposedly prevented site was not explicitly excluded from ordinary attachment.
That meant the intervention could partially undo itself.
The twelve-step primary effect was therefore:
INVALID_IMPLEMENTATION
Again, the invalid run was not negative evidence.
But this time something stronger happened.
Not every part of the experiment depended on the defective downstream intervention semantics.
The immediate expected causal effect remained interpretable.
And that immediate result was strong.
Locally accessible decaying material reduced immediate causal sensitivity.
So Chapter 27 V1 simultaneously contained:
INVALID PRIMARY
and:
SUPPORTED SUB-RESULT
This forced a more precise rule:
Invalidate only the inference touched by the defect. Preserve independently valid sub-results.
That is harder than throwing the run away.
It is also more honest.
Correcting Chapter 27
V2 fixed the intervention.
PREVENT explicitly blocked the target at lag 1.
FORCE exposed the target for exactly one growth step.
Both branches then removed it and resumed ordinary dynamics.
The corrected primary used a Rao-Blackwellized downstream causal consequence.
The mean accessible-versus-remote difference was approximately:
So the direction was supported.
But the achieved MDE was approximately:
The correct classification was therefore split:
DIRECTION_SUPPORTED
and:
MINIMUM_MAGNITUDE_UNRESOLVED
This gave us another identity:
DIRECTIONAL EFFECT โ ESTABLISHED MINIMUM MAGNITUDE
A conventional summary might have called the experiment successful because the interval excluded zero.
Our protocol would not allow that shortcut.
The scientific question had included a magnitude requirement.
So the magnitude question remained unresolved.
The Closeout Did Not Rescue the Claim
The Chapter 27 trajectory audit found something fascinating.
A substantial share of the downstream causal difference continued to accumulate after the material trace had fallen below half of its initial mass.
The result was consistent with:
early material-state modulation
โ early construction difference
โ redirected later trajectory
rather than a purely contemporaneous bias.
That was useful.
But it was descriptive.
It did not change the frozen confirmatory magnitude status.
The ledger therefore recorded:
DESCRIPTIVE_TRAJECTORY_PERSISTENCE
while leaving:
MINIMUM_MAGNITUDE_UNRESOLVED
untouched.
This protects against a subtle rescue pattern.
When the primary result weakens, it is tempting to run another analysis, find something interesting, and let the interesting mechanism retroactively improve the original claim.
Chapter 29 forbids that.
A descriptive explanation can illuminate a result without rescuing its confirmatory status.
Chapter 28: When a Strong Result Is Still the Wrong Construct
Chapter 28 was the most dangerous case because the first result was not weak.
It was extremely strong.
The raw causal modularity score at radius 4 was approximately:
It was supported.
Internal perturbations expressed much more causal mass inside the selected region than external perturbations penetrated into it.
Nothing about the raw measurement was wrong.
The problem was construct validity.
The module score increased as the disk radius increased.
That suggested a simple alternative explanation:
local causal propagation
+
observer-drawn containment
could generate the same effect.
The issue was no longer:
is the module score real?
It was:
does the module score indicate a privileged causal individual?
Those are different claims.
The Stronger Control
Chapter 28 V2 kept the same radius, same intervention, same causal-mass definition and same unbounded dynamics.
It added same-checkpoint geometry-matched controls.
Each selected region was compared against other radius-4 regions matched on:
occupancy fraction
radial position
occupied count
internal frontier count
external frontier count
probe depth
boundary occupancy
The observed regions still showed a large module score:
The selected regions were not meaningfully more modular than matched arbitrary regions.
So the stronger interpretation failed.
But the V1 phenomenon remained.
This is the key distinction:
RAW CAUSAL CONTAINMENT
SUPPORTED
while:
PRIVILEGED CAUSAL REGION
NOT ESTABLISHED
The resulting identity was:
CAUSAL RETENTION โ CAUSAL INDIVIDUATION
A stronger control narrowed the claim.
It did not erase the earlier measurement.
The Audit
Chapter 29 then ran the ledger against the actual project artifacts.
All ten registered cases had source artifacts.
No case required invented values.
No case depended on manual evidence alone.
The audit checked:
10 registered cases
and:
10 / 10 passed their consistency rules
Five cross-case checks were also predeclared.
All five passed.
The checks verified that:
- Chapter 27’s invalid V1 primary preserved the independently valid immediate effect.
- Chapter 27 separated directional support from minimum-magnitude support.
- Chapter 28’s stronger control narrowed the claim without erasing raw containment.
- Chapter 26’s bounded aggregate result preserved mechanistic rerouting.
- Chapter 27’s descriptive closeout did not rescue the unresolved confirmatory magnitude claim.
The final status was:
FAILURE_LEDGER_CONSISTENT
That status does not mean the project never made mistakes.
It means the mistakes and weakened claims can be represented without turning one evidence class into another.
That is a much more modest result.
And a much more useful one.
Construct Separation
The failure ledger also made visible a pattern that had been accumulating throughout the book.
Again and again, two things that initially seemed interchangeable turned out not to be.
Earlier chapters had already established distinctions such as:
MARGINAL EFFECT
โ
AVERAGE PAIRED EFFECT
AVERAGE PAIRED EFFECT
โ
PATHWISE DIVERGENCE
CAUSAL DIFFERENCE
โ
SYSTEMATIC SIGNATURE
PERSISTENT TRACE
โ
READABLE TRACE
ATTACHMENT
โ
NEW CONSTRUCTION
REOCCUPATION
โ
REPAIR
The recent chapters added:
CAUSAL ROUTING
โ
CAUSAL AMPLIFICATION
CAUSAL RETENTION
โ
CAUSAL INDIVIDUATION
These are not failures in the ordinary sense.
They are discoveries about where our concepts split.
A useful experimental system does not merely tell us whether a phenomenon exists.
It tells us which concepts we had incorrectly bundled together.
flowchart TD
subgraph EarlierDistinctions
A[Marginal effect / average paired effect]
B[Average paired effect / pathwise divergence]
C[Causal difference / systematic signature]
D[Persistent trace / readable trace]
E[Attachment / new construction]
F[Reoccupation / repair]
end
subgraph RecentDistinctions
G[Causal routing / causal amplification]
H[Causal retention / causal individuation]
end
A --> I[Conceptual split discovered]
B --> I
C --> I
D --> I
E --> I
F --> I
G --> I
H --> I
Invalid Is Not Negative
This deserves its own rule.
Suppose an experiment contains a broken intervention.
The result is invalid.
It does not matter whether the measured effect is:
positive
negative
zero
large
small
significant
nonsignificant
The intended estimand was not correctly implemented.
Therefore the experiment cannot serve as evidence for or against that intended claim.
This seems obvious when stated abstractly.
In practice, it is easy to violate.
Especially when the invalid result is convenient.
If a broken experiment produces zero, people are tempted to say:
we found no effect.
If it produces a large effect, they are tempted to salvage it.
Both are mistakes.
The correct status is:
INVALID
and then the question becomes:
which parts, if any, were not touched by the defect?
That is exactly what happened in Chapter 27.
Unresolved Is Not Failed
Now consider a valid experiment.
Suppose the interval crosses zero.
That alone does not tell us the result is negative.
The experiment may simply lack precision.
The correct question is:
What effect sizes are still compatible with the data?
If scientifically meaningful positive and negative effects both remain plausible, the result is:
UNRESOLVED
Not:
FAILED
Chapter 27’s magnitude question is a good example.
The direction was supported.
But the MDE remained too wide for the frozen ยฑ0.15 magnitude claim.
So the magnitude remained unresolved.
The fact that another inferential layer was supported did not make the unresolved layer disappear.
A Bounded Negative Is a Real Result
The reverse mistake is equally common.
A result with an interval near zero is sometimes summarized as:
nothing happened.
That discards information.
If the experiment is precise enough to exclude effects larger than a predeclared meaningful threshold, then the negative result is substantive.
Chapter 26 V2 did this for amplification.
Chapter 28 V2 did it for excess modularity.
In both cases, the claim was not merely unsupported.
A meaningful positive effect was bounded away.
That is stronger.
The correct language is:
BOUNDED BELOW SEI
or:
BOUNDED NEAR ZERO
depending on the estimand.
The threshold matters because zero is rarely the only scientifically interesting boundary.
Preserve the Surviving Phenomenon
Perhaps the most important rule emerged from Chapters 27 and 28.
When a higher-level interpretation fails, do not automatically delete the lower-level observation.
Chapter 27 V1’s broken downstream intervention did not erase the valid immediate material-state effect.
Chapter 28 V2’s matched null did not erase raw spatial causal containment.
This gives us a hierarchy:
measurement
โ
mechanism
โ
construct
โ
interpretation
A failure higher in that hierarchy need not propagate all the way down.
If the measurement remains valid, keep it.
If the mechanism remains supported, keep it.
Fail only the layer the evidence actually defeats.
That leads to the central rule of the chapter.
Fail the Smallest Claim
The final Chapter 29 principle is:
Fail the smallest claim justified by the evidence. Preserve everything that still survives.
This prevents two opposite errors.
The first is over-rescue:
the main claim weakened
โ reinterpret until something sounds successful
The second is over-destruction:
the main claim weakened
โ throw away the entire experiment
Both distort the evidence.
The correct response is surgical.
If the implementation is invalid, invalidate the affected estimand.
If precision is insufficient, leave the claim unresolved.
If the confidence interval and MDE bound the effect below the SEI, record a bounded negative.
If a stronger control removes a higher-level interpretation, preserve the lower-level phenomenon that still survives.
If a descriptive audit explains the mechanism, label it descriptive.
Do not use it as retroactive rescue.
flowchart TD
A[Claim weakens] --> B{What failed?}
B -- Implementation invalid --> C[Invalidate affected estimand only]
B -- Precision insufficient --> D[Mark unresolved]
B -- Effect bounded below SEI --> E[Record bounded negative]
B -- Stronger control narrows --> F[Preserve lower-level phenomenon]
B -- Descriptive mechanism --> G[Label descriptive only]
C --> H[Smallest claim fails]
D --> H
E --> H
F --> H
G --> H
H --> I[Preserve all surviving evidence]
Why This Matters for Digital Life
The project began with a risk.
We were deliberately searching for properties associated with life.
That makes us vulnerable to seeing them everywhere.
Growth looks like reproduction.
Refill looks like repair.
Persistence looks like memory.
Coherent motion looks like flocking.
Causal containment looks like individuality.
The danger is not merely false positives.
The deeper danger is conceptual compression.
We see one measurable phenomenon and immediately promote it into a richer biological category.
The failure ledger is a defense against that.
It forces every promotion to survive:
construct definition
implementation
measurement
control
precision
alternative explanation
And when a promotion fails, it tells us what remains underneath.
That is why the negative results have become some of the most useful results in the book.
They tell us where one concept stops and another begins.
What the Audit Does Not Prove
The Chapter 29 audit passed.
But its scope is deliberately bounded.
It does not establish that:
the project never made mistakes
It does not establish that:
the taxonomy is universally complete
It does not establish that:
every future experiment will classify cleanly
And it does not establish a general philosophy of science.
It establishes something narrower:
Across the registered Chapter 26โ28 experiment chains, invalid implementations, unresolved questions, precision-bounded negative results, descriptive closeouts and stronger-control claim narrowings can be represented without erasing surviving evidence or converting one evidence class into another.
That is enough.
Because Chapter 30 now has a foundation.
We can finally ask what survived the entire investigation.
Not what looked alive.
Not what we hoped would be alive.
Not what received a biological name.
What survived.
After every invalidation.
Every null.
Every stronger control.
Every distinction.
Every failed interpretation.
What remains?
That is the question for the final chapter.