Cascades
Does uncertainty-driven escalation beat always using the strongest provider?
Cascades
This chapter is a simulation over assumed stage profiles. No cascade is measured. Every number is a consequence of assumed accuracy curves, costs, confidence noise, shift and human capacity, all swept and stated next to the results. In particular: the naive product of stage accuracies does not equal the measured end-to-end accuracy, so any “cascades save X%” headline below is a statement about the swept assumptions, not about real providers.
The problem
Chapter 29’s router chose one provider per request. A cascade is the unrolled version: run a cheap stage, escalate the uncertain ones, and (Chapter 17’s machinery) hold the end to a risk target. This chapter asks: does uncertainty-driven escalation beat always using the strongest provider?
What we expect and why
Our starting hypothesis is that the saving is large on skewed traffic and collapses under shift. Three sources frame it:
-
Chen, Zaharia and Zou (2023) — FrugalGPT cascades LLMs and reports matching GPT-4 at up to 98% lower cost. (Read at abstract level.)
-
Gupta et al. (2024) — token-level uncertainty cascades show simple confidence deferral is not enough. (Read at abstract level.)
-
Viola and Jones (2001) — the original boosted cascade rejects cheaply and escalates only the hard; the architecture is old and it worked in vision. (Verified CVPR 2001, IEEE 10.1109/CVPR.2001.990517; classic, not on arXiv.)
The build
src/arbiter/cascade.py runs stages rule → linear → judge → human. A stage answers when its assumed confidence clears a threshold learned on a calibrate stream, else it escalates. The human stage is capacity-limited: once its queue is full, overflow items fail. run_ch30.py sweeps the answer fraction, the shift (a drift factor on stage-1 confidence) and the human capacity, writing results/ch30.jsonl (mode: simulation).
from arbiter.cascade import Stage, p_ok
# 1. A request escalates until a stage's assumed confidence clears a threshold.
print("1. escalation (ASSUMED stage accuracies)")
stages = [
Stage("rule", 0.95, 0.45, 1.0, 0.2),
Stage("judge", 0.99, 0.92, 20.0, 0.05),
]
d = 0.7
for s in stages:
print(f" {s.name:<8} P(correct)={p_ok(s, d):.2f} cost={s.cost}")
# 2. The naive product does not equal the chain.
print("2. the chain is not a product")
p_rule, p_judge = p_ok(stages[0], d), p_ok(stages[1], d)
print(f" P(rule ok)={p_rule:.2f} P(judge ok)={p_judge:.2f} product={p_rule*p_judge:.2f}")
print(" the chain's joint is P(rule|escalated) style, not the product")
1. escalation (ASSUMED stage accuracies)
rule P(correct)=0.60 cost=1.0
judge P(correct)=0.94 cost=20.0
2. the chain is not a product
P(rule ok)=0.60 P(judge ok)=0.94 product=0.56
the chain's joint is P(rule|escalated) style, not the product
The sweep
run_ch30.py, n=4,000, assumed stage curves in every row. Navigation:
| config | accuracy | cost/req | naive product | judge traffic |
|---|---|---|---|---|
| always-judge | 0.952 | 20.00 | — | 4,000 |
| answer 0.5, drift 1.0 | 0.830 | 11.80 | 0.577 | 433 |
| answer 0.5, drift 0.5 | 0.918 | 20.08 | 0.805 | 874 |
| answer 0.5, drift 1.5 | 0.746 | 4.96 | 0.543 | 121 |
| answer 0.7, drift 1.0 | 0.778 | 5.02 | 0.548 | 226 |
| human capacity 50 | 0.673-0.752 | 3.6-4.2 | — | — |
What it says
-
At matched accuracy the cascade does not save here. The cheapest way to reach the judge’s 0.952 is to send almost everything to the judge (drift 0.5, cost 20.08). The cascade saves 41% (11.80) only by accepting 12 points less accuracy. The author’s warning is the finding: a headline “cascades save 41%” would be true here without the accuracy match, and false with it. P1 holds only at an accuracy cost.
-
The guarantee does not compose. The naive product (0.577) is far below the measured end-to-end accuracy (0.830) at answer 0.5. Multiplying per-stage risks assumes independence, and Chapter 24 already showed the product is exact only there; escalation is correlated with difficulty, so the end-to-end risk must be measured on a held-out split. P2 holds.
-
Shift breaks it in both directions. Drift 1.5 (stage 1 overconfident) sends only 121 items to the judge and accuracy collapses to 0.746: hard items are answered by the weak rule. Drift 0.5 (under-confident) swamps the judge (874 items, cost 20.08). Thresholds calibrated in-distribution go stale; P3 holds.
-
The human stage is a real, capped cost. At capacity 50 the cascade’s accuracy falls to 0.673-0.752 because overflow items fail. A capacity-free human stage is a fantasy; P4 holds.
Wrong / Correct. Wrong: “Cascades save most of the cost at the same quality.” Correct: “Under these assumed profiles the cascade’s saving is bought with accuracy unless the confidence signal is near-perfect; the guarantee does not compose; shift and human capacity break it. Anyone claiming a cascade saving must state the accuracy match and the assumptions.”
What to carry forward
Chapter 31 learns the router from the cascade’s outcomes. This chapter hands it a warning: the cascade’s thresholds were calibrated in-distribution and broke under shift, and the router will be worse — a router that only sees the outcomes of providers it selects learns the wrong thing. That is Chapter 31’s topic.
Close by
Which model should handle each decision, and who checks the cascade? Under this evidence: no single model; a controller whose thresholds are re-validated, because the guarantee does not compose and shift breaks it.
Limitations
- SIMULATION over assumed stage profiles;
results/ch30.jsonlismode: simulationwith the assumptions in every row. FrugalGPT and Gupta et al. are ABSTRACT_ONLY; none is measured. - The naive-product comparison is computed from the measured per-stage accuracies; the correlation it exposes is a property of the sampled stream.
- “Saving” figures are for the specific answer fractions and curves swept; the direction of the findings (accuracy match needed, no composition, shift breaks it) is the durable part.
- Viola-Jones cited at summary level (IEEE 10.1109/CVPR.2001.990517).