← Jev From First Principles

Confidence Is Not Probability

Measure what a provider's scores mean before your program branches on them.

You have a program that can approve a request or send it to a person. The provider returns a valid distribution, the selected option has a large probability, and your next line is if p > 0.9. What does that comparison buy you?

The Decision Contract checks that an answer covers the requested options and that inability is explicit. It cannot check how often a prediction is correct. You need observations for that second question. This chapter measures the available providers on CPU. It fits no calibration method; Calibration owns that work.

PROPOSED: separate four jobs. The specification names the decision and its answer set. The provider produces scores. The runtime checks their shape. An evidence record compares those scores with outcomes. Your application still has to choose an action and the acceptable consequences of a mistake.

The expectation must be allowed to fail

DOCUMENTED: Guo and colleagues define top-label calibration through the frequency of correct predictions conditional on the probability of the selected class. Their ECE approximates that relationship by grouping predictions into bins. Their experiments also show that accuracy and calibration can move in different directions; their positive temperature transformation preserves the winning class. These are definitions and findings from their experiments, not an assertion about every provider in this book. Guo et al., §§2–5.

DOCUMENTED, challenging our starting expectation: Desai and Durrett found that pretrained transformers could be more calibrated than their smaller comparators. Their in-domain and out-of-domain results also differed. You cannot diagnose calibration from architecture or sophistication alone. Desai and Durrett, §§4.3–4.4.

DOCUMENTED: Ovadia and colleagues evaluated uncertainty under shift across several modalities. Calibration on a validation distribution did not generally transfer to changed distributions. They also reported mostly consistent method ordering across their experiments: a ranking reversal is possible, not required. Ovadia et al., §§3–5.

We committed the numeric predictions before measuring. They include ECE bands, a prediction that LR beats fastText on calibration, larger errors under safety shift, a binning reversal, and changes with sample size and embedding scale. The predictions can fail without making the chapter fail.

Make the measures disagree in the open

PROPOSED definitions, implemented and checked here: top-label ECE averages the absolute gap between a bin’s mean maximum probability and its observed correctness, weighted by bin count. Classwise ECE instead checks each class’s probability against whether that class occurred, then takes an unweighted mean across classes. We include all class probabilities. A low top-label ECE does not establish calibration of the other options; classwise macro averaging can dilute the importance of a rare but costly class.

We use equal-width and equal-mass bins, at the two preregistered bin counts. Width bins are right-closed, with zero in the first bin. Mass bins use quantiles; duplicate edges collapse so identical scores stay together. The requested and effective counts are recorded. Empty bins have null means and contribute no weight. These choices are part of the estimator, not formatting preferences.

Brier here is the mean sum of squared errors across the full distribution and one-hot truth. It uses the same scale for all providers within a task. Its binary value is twice the conventional scalar binary Brier score. Ovadia’s paper uses a class-count-normalized convention, so you cannot compare those numbers without rescaling. Brier combines aspects of calibration and discrimination; it is not an isolated calibration error.

Log loss scores the probability assigned to the actual class. We clip at the preregistered epsilon, then renormalize and use natural logarithms. The clipping rule changes the penalty for impossible-but-observed events. Accuracy measures which label wins; macro one-versus-rest AUROC measures ranking and gives tied scores their average rank. Neither establishes calibration. AUROC is undefined when a required class has no positive or negative observations.

Each ECE carries a percentile bootstrap interval. Head-to-head differences use the same resampled item indices for both providers. Test-versus-shift differences resample independently because these are different items. The fit stays fixed: these intervals quantify evaluation-sample uncertainty, not all uncertainty in training, dataset selection or deployment.

DOCUMENTED limitation: Ciosek and colleagues distinguish heuristic binned estimates from bounds on population calibration under stated assumptions. Their experiments show ECE behaving competitively on some synthetic calibration functions and failing on another. Their guarantees concern binary classifiers; the multiclass extension is future work. Our bootstrap does not implement those guarantees. A narrow interval around an estimator does not remove its binning bias. Ciosek et al., §§3,7,9–11.

OBSERVED — JEV-09-03: reference and hand-fixture agreement had maximum absolute error 2.22e-16 across 17 checks. The suite passed 72 tests, including earlier contract tests. The planted bin-edge defect was caught through bin membership/counts; a scalar ECE alone can miss a boundary defect through cancellation. The earlier recorded-score replay still preserved 1540 decisions. Chapter 5 compatibility used a fake encoder. Chapter 6’s real NLI smoke was not run; its provider is covered by the fake contract suite.

The measures, running

The definitions above are easier to trust once you have watched them disagree. Everything below is examples/ch09-confidence-is-not-probability/walkthrough_ch09.py, which you can run as it stands. It uses the chapter’s own estimators from benchmarks/harness/calibration.py. Parts 1 to 3 use hand-written numbers small enough to check by hand. Part 4 is a seeded simulation, labelled ILLUSTRATIVE. None of it is a measurement of any provider.

import numpy as np

ROOT = Path(__file__).resolve().parents[2]
sys.path[:0] = [str(ROOT), str(ROOT / "src")]

from benchmarks.harness.calibration import brier_score, clipped_log_loss, macro_auroc, top_label_ece


def softmax(logits, temperature):
    z = np.asarray(logits, dtype=float) / temperature
    z -= z.max(axis=1, keepdims=True)
    e = np.exp(z)
    return e / e.sum(axis=1, keepdims=True)


def report(name, y, p, bins=10):
    accuracy = float(np.mean(p.argmax(1) == y))
    auroc = macro_auroc(y, p)  # undefined (None) when every label is the same class
    shown = "n/a" if auroc is None else f"{auroc:.3f}"
    print(f"   {name:<22} accuracy {accuracy:.3f}  ECE {top_label_ece(y, p, bins=bins):.3f}  "
          f"Brier {brier_score(y, p):.3f}  log loss {clipped_log_loss(y, p):.3f}  AUROC {shown}")


def main() -> None:
    # 1. A predictor that always says 60/40 is perfectly calibrated, and useless.
    y = np.array([1] * 6 + [0] * 4)  # class 1 is right 60% of the time
    always_prior = np.tile([0.4, 0.6], (10, 1))
    print("1. a perfectly calibrated predictor that knows nothing")
    report("always 60/40", y, always_prior)
    print("   it is right 6 times in 10 and says 0.6 each time, so ECE is exactly 0;")
    print("   AUROC 0.5 says it cannot tell one case from another")

    # 2. A predictor that claims far more than it earns.
    y2 = np.ones(10, dtype=int)  # the right answer is class 1 every time
    says_class_1 = np.array([1, 1, 1, 1, 1, 1, 1, 0, 0, 0])  # but it is wrong on the last three
    p2 = np.where(says_class_1[:, None] == 1, [0.05, 0.95], [0.95, 0.05])
    print("2. a confident predictor that is right 7 times in 10")
    report("always 95%", y2, p2)
    print("   it claims 0.95 on every item and earns 0.70: the gap is the ECE, 0.250")

    # 3. A scale changes ECE and leaves every decision alone.
    logits = np.array([[2.0, 0.0, -1.0], [1.5, 0.5, 0.0], [0.2, 0.1, 0.0], [3.0, 0.0, 0.0],
                       [0.5, 0.4, 0.3], [1.0, 2.0, 0.0], [0.0, 0.1, 0.9], [2.5, 2.0, 0.0]])
    y3 = np.array([0, 1, 0, 0, 2, 1, 2, 0])
    print("3. temperature rescales the probabilities and never changes the winner")
    for t in (0.5, 1.0, 2.0, 4.0):
        report(f"temperature {t}", y3, softmax(logits, t), bins=5)

    # 4. ECE depends on how many items you have, even for a perfectly calibrated source.
    print("4. a perfectly calibrated source, scored on samples of different sizes (ILLUSTRATIVE)")
    rng = np.random.default_rng(0)
    print("   n      mean ECE over 200 draws")
    for n in (10, 30, 100, 1000, 10000):
        eces = []
        for _ in range(200):
            conf = rng.uniform(0.5, 1.0, n)
            hit = rng.random(n) < conf  # correct with exactly the stated probability
            y_s = np.where(hit, 0, 1)
            p_s = np.column_stack([conf, 1 - conf])
            eces.append(top_label_ece(y_s, p_s, bins=10))
        print(f"   {n:>5}  {np.mean(eces):.3f}")
    print("   the source is calibrated by construction, yet small samples report a gap")
1. a perfectly calibrated predictor that knows nothing
   always 60/40           accuracy 0.600  ECE 0.000  Brier 0.480  log loss 0.673  AUROC 0.500
   it is right 6 times in 10 and says 0.6 each time, so ECE is exactly 0;
   AUROC 0.5 says it cannot tell one case from another
2. a confident predictor that is right 7 times in 10
   always 95%             accuracy 0.700  ECE 0.250  Brier 0.545  log loss 0.935  AUROC n/a
   it claims 0.95 on every item and earns 0.70: the gap is the ECE, 0.250
3. temperature rescales the probabilities and never changes the winner
   temperature 0.5        accuracy 0.750  ECE 0.178  Brier 0.391  log loss 0.649  AUROC 0.823
   temperature 1.0        accuracy 0.750  ECE 0.209  Brier 0.399  log loss 0.685  AUROC 0.823
   temperature 2.0        accuracy 0.750  ECE 0.259  Brier 0.469  log loss 0.813  AUROC 0.872
   temperature 4.0        accuracy 0.750  ECE 0.338  Brier 0.551  log loss 0.933  AUROC 0.872
4. a perfectly calibrated source, scored on samples of different sizes (ILLUSTRATIVE)
   n      mean ECE over 200 draws
      10  0.213
      30  0.133
     100  0.067
    1000  0.023
   10000  0.007
   the source is calibrated by construction, yet small samples report a gap

Read it part by part.

  1. ECE can be zero for a predictor that knows nothing. Saying 60/40 every time against a 60% base rate is perfectly calibrated, so its ECE is exactly 0, while its AUROC of 0.5 says it cannot tell one case from another. This is why ECE is never reported alone: it measures whether the stated confidence is honest, not whether the predictor is useful.
  2. The gap is the ECE. A predictor that claims 0.95 on every item and earns 0.70 has an ECE of 0.250, exactly the difference. (AUROC is undefined here because every label is the same class, and the helper says so instead of inventing a number.)
  3. A scale changes the probabilities and leaves every decision alone. Accuracy is 0.750 at every temperature. ECE moves (0.178, 0.209, 0.259, 0.338) and so do Brier and log loss. It happens to rise with temperature in this toy, but there is no universal direction. AUROC also shifts, from 0.823 to 0.872, because it ranks per-class probabilities across items and a rescaling is not a monotone map across items with different logit gaps. Accuracy cannot move; every other measure can. Chapter 10 fits the temperature on purpose. The warning here is that a provider whose scale was chosen arbitrarily, as the embedding provider’s was, can report almost any ECE without changing a single answer.
  4. ECE depends on how much data you scored it on. The simulated source is calibrated by construction, because each item is correct with exactly its stated probability. Scored on 10 items it still reports a mean ECE of 0.213, falling to 0.067 at 100 and 0.007 at 10,000. A small sample reports a gap that a large one does not, which is why the safety splits (n of 90 and 116) carry wide intervals in this chapter.

What ran, and what did not

OBSERVED method: we refitted the existing LR and fastText-style providers on the frozen Chapter 4 training splits. LR uses the inherited threshold-selected configuration. FastText uses its inherited configuration and all declared seeds. No hyperparameter or calibration method was selected in this chapter. The embedding provider uses the exact cached small-encoder revision on CPU. A local subclass supplies that snapshot explicitly because the old loader did not pass its declared revision.

LR and fastText reuse their historically selected configurations. Embedding has no calibration grid and retains its availability choice and fixed earlier scale. This is an explicit difference in tuning treatment. These comparisons do not isolate architecture or tuning effort as a cause.

The intent comparison gives every provider the same complete answer set. Zero-Shot Decisions used a different seen/unseen comparison for LR. This is a matched-set refit, not a claim to reproduce that chapter’s headline. Embedding descriptions use the existing author-written paraphrase fixture. We keep every wording result; lexical overlap and one author’s choices remain limitations. Import the geometry from Embeddings From First Principles, and the distinction between an assertion and evidence from Hallucination From First Principles.

We acquired each task’s test bundle once for this chapter, under the shared harness gate, and persisted the acquisition ledger. Historical touches retain their earlier chapter keys. The declared configurations, seeds, wordings and headline measures belong to that one logical evaluation. Resumption reads frozen predictions. Binning searches, sample-size demonstrations, scale variation and clipping sensitivity use calibrate or threshold, never test.

The safety shift is the pinned jackhhao dataset with the same binary labels. It changes source, content and class mix together. NOT_OBSERVED: an intent shift. You cannot turn a binary jailbreak dataset into a banking-intent shift by changing a column name.

NOT_OBSERVED: hosted Jev calibration, probability saturation, and the commenters’ claims about its confidence errors. The provider and response cache were absent. No API calls were made. Chapter 6 has no committed per-item output appropriate for this analysis, so its NLI comparison is also absent. We ran no NLI or option-scored LM inference. Deferred Chapter 7 model arms stay deferred.

The evidence beside the number

OBSERVED — JEV-09-01:

Dataset/split Provider/seed/wording ECE10 width [95% bootstrap] Classwise ECE10 Brier sum Log loss Accuracy AUROC n
intent/test prior/s0/w0 0.0064 [0.0028, 0.0103] 0.0026 0.9878 4.3820 0.0130 0.5000 3080
intent/test tfidf-lr/s0/w0 0.0519 [0.0430, 0.0616] 0.0023 0.1835 0.5147 0.8779 0.9971 3080
intent/test fasttext-style/s0/w0 0.0408 [0.0343, 0.0540] 0.0026 0.2251 0.6051 0.8526 0.9962 3080
intent/test embed-sim/s0/w0 0.5630 [0.5469, 0.5784] 0.0100 0.8226 2.4848 0.6799 0.9804 3080
safety/test prior/s0/w0 0.1484 [0.0622, 0.2432] 0.1484 0.5434 0.7380 0.4828 0.5000 116
safety/test tfidf-lr/s0/w0 0.0593 [0.0339, 0.1279] 0.0859 0.1710 0.3392 0.8966 0.9661 116
safety/test fasttext-style/s0/w0 0.0642 [0.0378, 0.1212] 0.0743 0.1580 0.2471 0.9052 0.9658 116
safety/test embed-sim/s0/w0 0.0574 [0.0181, 0.1554] 0.0685 0.4827 0.6753 0.5345 0.6423 116
safety/shift prior/s0/w0 0.1617 [0.1044, 0.2266] 0.1617 0.5504 0.7452 0.4695 0.5000 262
safety/shift tfidf-lr/s0/w0 0.4375 [0.3754, 0.4937] 0.4396 0.8646 3.4633 0.5496 0.8644 262
safety/shift fasttext-style/s0/w0 0.4361 [0.3731, 0.4934] 0.4378 0.8633 2.1766 0.5382 0.5115 262
safety/shift embed-sim/s0/w0 0.1861 [0.1209, 0.2420] 0.1861 0.5652 0.7600 0.3740 0.3085 262

The table uses the declared primary seed and wording. It reports ECE next to accuracy, full-distribution scores, AUROC and sample size. Full tables retain every seed, wording, split, binning choice and paired interval in the evidence report. Reliability data, including bin counts, are recorded in the JSONL file.

On these in-distribution samples, the classical providers’ top scores are much closer to correctness frequencies than banking embedding scores. This supports an approximate frequency reading under the declared estimator and distribution, with the reported uncertainty. The prior control is also approximately calibrated on banking while offering no useful discrimination. No provider earns a universal probability guarantee, and the safety shift changes the answer substantially.

OBSERVED: the signed mean-confidence-minus-accuracy gap makes the direction visible. Banking embedding is underconfident by 0.5630; banking LR and fastText are overconfident by 0.0514 and 0.0376. On safety shift, LR and fastText gaps are 0.4375 and 0.4330. These are aggregate directions, not explanations of their mechanisms or claims that every bin behaves alike.

Reliability and bin counts for the fixed primary fits

The diagonal asks whether mean top probability matches correctness within a bin. The lower panels show how many observations support each point. An empty region of this diagram supports no claim about requests that land there in production.

OBSERVED, paired intervals at the fixed primary seed/wording:

For intent, LR minus fastText ECE is 0.0111 [-0.0029, 0.0184]; Brier difference is -0.0416 [-0.0551, -0.0279] and accuracy difference 0.0253 [0.0143, 0.0364]. Negative ECE/Brier differences favor LR; positive accuracy differences favor LR. The predicted LR calibration advantage is refuted.

For safety, LR minus fastText ECE is -0.0049 [-0.0397, 0.0382]; Brier difference is 0.0130 [-0.0453, 0.0694] and accuracy difference -0.0086 [-0.0603, 0.0347]. Negative ECE/Brier differences favor LR; positive accuracy differences favor LR. The predicted LR calibration advantage is refuted.

On safety shift, the ECE changes (shift minus test, independently resampled) are tfidf-lr: 0.3782 [0.2877, 0.4411]; fasttext-style: 0.3719 [0.2880, 0.4336]; embed-sim: 0.1287 [0.0175, 0.1966]. These compare different datasets; they do not isolate the cause of the change. The numeric degradation prediction has provider-specific verdicts in metadata.

Across the declared seeds, intent: fastText ECE range 0.0389–0.0423, accuracy 0.8516–0.8526; safety: fastText ECE range 0.0439–0.0675, accuracy 0.8793–0.9052. All seed results and their LR comparisons are retained; the primary table is not seed selection.

Wording variation is intent/test: embedding accuracy 0.5000–0.6799, ECE 0.4209–0.5630; safety/test: embedding accuracy 0.4741–0.6983, ECE 0.0574–0.2205; safety/shift: embedding accuracy 0.3740–0.6641, ECE 0.0399–0.1861. No wording was chosen on test. These ranges do not establish generalization beyond the frozen descriptions.

A probability-shaped number can be useless

OBSERVED: the banking train-prior control has ECE 0.0064 [0.0028, 0.0103], accuracy 0.0130, Brier 0.9878, log loss 4.3820 and AUROC 0.5000, with n=3080. It ranks no state above another. In the exact synthetic class-prior fixture, ECE is 0.0000 while accuracy is 0.6000, AUROC 0.5000, Brier 0.4800 and log loss 0.6730, with n=10. The latter’s finite-sample bootstrap interval is in the results record.

Wrong: the lowest ECE identifies the best decision provider.

Correct: ECE asks one frequency question under one estimator. Read it beside discrimination, proper scoring rules, accuracy and the action’s error costs.

The train-prior control is deliberately boring. It gives every state the same distribution. Its low ECE is not a software defect or proof of useful decisions. It is the case the suite must accept while making its lack of discrimination visible. A provider’s ability to rank requests and its ability to report useful frequencies are distinct requirements.

Bins and sample sizes can move the conclusion

OBSERVED: the preregistered calibrate/threshold search found 3 provider-pair/split cases with a point-estimate ordering reversal. For safety/calibrate, tfidf-lr versus fasttext-style, the ECE differences were width10 0.0070, width15 -0.0070, mass10 -0.0095. The two bin-count/scheme comparisons are reported separately.

Safety threshold provider n Median ECE10 95% subsampling spread Full-sample ECE10
tfidf-lr 30 0.0837 [0.0226, 0.1584] 0.0685
tfidf-lr 60 0.0722 [0.0421, 0.1023] 0.0685
tfidf-lr 90 0.0685 [0.0685, 0.0685] 0.0685
fasttext-style 30 0.1126 [0.0457, 0.2013] 0.1051
fasttext-style 60 0.1090 [0.0732, 0.1502] 0.1051
fasttext-style 90 0.1051 [0.1051, 0.1051] 0.1051
embed-sim 30 0.0954 [0.0269, 0.2210] 0.0796
embed-sim 60 0.0756 [0.0401, 0.1375] 0.0796
embed-sim 90 0.0796 [0.0796, 0.0796] 0.0796

The spread comes from repeated subsets without replacement, not an iid confidence interval. Each row also records a declared first-subset example with ECE bootstrap interval, accuracy, Brier, log loss and AUROC. We do not call its median ECE a population error bound.

These comparisons use the declared search space. We did not hunt for a reversal on test or keep changing bins until one appeared. A point-estimate rank flip does not by itself establish a statistically resolved performance reversal. At small sample sizes, bins may describe only a handful of observations. Bootstrap uncertainty is part of the result, and small-n resampling spreads have a different interpretation from confidence intervals.

Normalization is not a calibration method

The existing embedding provider computes cosine similarities and multiplies them by a positive scalar before softmax. Its parameter is called temperature in the earlier code, but operationally it is an inverse temperature. Here we record the multiplier as k. We preserve the earlier operation rather than silently switching to division.

HYPOTHESIS tested here: changing that positive multiplier keeps the winning label and can change ECE. We did not assume arbitrary scores are miscalibrated by logical necessity. They may happen to be calibrated. What normalization fails to supply is evidence for a frequency interpretation.

OBSERVED, threshold only:

Task Multiplier k ECE10 width [95% bootstrap] Accuracy Brier Log loss AUROC n
intent 1 0.6506 [0.6196, 0.6805] 0.6670 0.9808 4.1304 undefined 1000
intent 10 0.5573 [0.5273, 0.5852] 0.6670 0.8339 2.5267 undefined 1000
intent 30 0.1058 [0.0841, 0.1317] 0.6670 0.4617 1.2567 undefined 1000
safety 1 0.1478 [0.0483, 0.2377] 0.6556 0.4915 0.6846 0.7356 90
safety 10 0.0796 [0.0350, 0.1717] 0.6556 0.4331 0.6243 0.7356 90
safety 30 0.1182 [0.0641, 0.2225] 0.6556 0.4005 0.5806 0.7356 90

Log-loss clipping sensitivity on threshold is retained separately. At the declared epsilon values, intent/tfidf-lr: range 0.0258; intent/fasttext-style: range 0.0212; intent/embed-sim: range 0.0000; safety/tfidf-lr: range 0.0003; safety/fasttext-style: range 0.0429. No threshold, scale or clipping rule was selected from these results.

No scale is selected from this demonstration. Choosing the best one would start Chapter 10’s calibration work and would require its own fitting protocol.

The intent threshold split omits a required class, so its macro AUROC is undefined under the declared rule. The complete intent test split supports the AUROC in the main table. We do not silently drop the missing class.

Jev’s concentration statistic needs its own question

DOCUMENTED — vendor: TypeSafe explicitly distinguishes confidence from the option probabilities. For a Choice it derives confidence from the selected probability and answer-set size:

\[ c = \frac{p_{\max}-1/K}{1-1/K}. \]
It returns no separate confidence field for Noul; that answer is a probability of yes. Concentration and correctness frequency are different claims. Applying correctness ECE to the derived statistic would diagnose whether it happens to match correctness, not test its documented arithmetic definition. [TypeSafe confidence documentation](https://docs.typesafe.ai/confidence).

C, untested: the Hacker News discussion contains the claimed confidence errors that motivated this chapter. We cannot confirm them from our absent Jev cache, and a percentage in that discussion is not automatically this chapter’s estimator. Discussion.

3P, not our observation: AnyJev’s historical receipt reports its raw, L0 and L1 readouts on a smaller banking answer set. L1 includes fitted temperature; our chapter fits none. The kickoff warns that its accuracy may reflect teacher agreement. The inspected aggregate receipt does not establish item-label provenance, so we keep that caveat rather than assert independent ground truth. Different answer sets, models and protocols also prevent a direct league table. AnyJev receipt.

The refreshed Red Hat benchmark reports classification comparisons, not a calibration evaluation. Its accuracy findings cannot fill the missing Jev calibration arm.

Put provenance beside the score

PROPOSED: the versioned contract view adds calibration_state without breaking the earlier constructors. The default is UNCALIBRATED. A CALIBRATED record requires the dataset or split, sample count, date, method and evidence reference. The shim upgrades an old result or decision into this view.

The validator checks that the required fields exist and have usable types. It does not verify that a method worked, that an evidence reference is true, or that another distribution will behave similarly. Every actual provider in this chapter remains UNCALIBRATED; measurement alone is not a fitted calibration method. A provenance field does not make your action safe by declaration.

Run the small demonstration from the repository root:

python examples/ch09-confidence-is-not-probability/demo_ch09.py

The implementation imports the shared harness. Its reference checks, independent hand fixtures and compatibility tests run with:

python -m pytest tests/calibration tests/contract -q

The staged reproduction and evidence verification commands are in the example README. Stages keep incomplete outputs under results/partial/; the final results and compressed item distributions appear only after completion.

What you can carry forward

OBSERVED: the predictions have mixed outcomes. ECE bands and shift-degradation thresholds are checked provider by provider, rather than collapsed into a universal claim. The LR-versus-fastText calibration prediction does not authorize selecting the lowest-ECE provider for the application. Binning and sample-size sensitivity are findings under the declared estimators.

PROPOSED conclusion: the contract should carry calibration provenance while keeping its default uncalibrated. A score may support a frequency interpretation on a particular evaluation distribution. These results do not certify that interpretation for an individual request or a future distribution.

The required papers were read partly from full text, with inspected sections recorded. Appendix proofs were not audited. We measured existing providers on fixed public datasets, not general calibration across deployments. Wording bias, finite-sample bias, conditional bootstrap intervals, missing intent shift and the absent hosted/NLI arms bound the conclusion. Untested explanations for a provider’s calibration are hypotheses, not mechanisms established here.

The numeric band prediction held for LR and fastText on both tasks, but failed for embedding in opposite directions: banking ECE was above its predicted band, and safety ECE below it. The predicted LR advantage failed on both tasks; the paired ECE intervals include zero, so these data do not resolve an ECE winner. Shift degradation, the declared binning reversal, scale sensitivity and the small-sample median prediction were observed. The embedding shift interval includes changes smaller than its predicted minimum, even though its point estimate passes that minimum.

Fit wall times are recorded for replay planning. CPU-seconds, peak memory and comparable serving latency were not measured; these results support no cost ranking.

Chapter 10 inherits checked estimators, unfitted calibration splits, fixed per-item outputs, and an explicit place for calibration provenance. What would you have to know before you let a program branch on this number?