The Decision Router
Can a router choose the cheapest competent provider for each request?
The Decision Router
This chapter is a simulation over assumed provider profiles. The provider accuracy-by-difficulty curves, costs and signal correlation are assumed and swept; every number is a consequence of them. No real router (RouteLLM, Hybrid LLM) is measured. The assumptions sit next to the results.
The problem
Chapter 28 ended with a table mostly full of NOT_OBSERVED. A router makes that honest gap useful: it chooses, per request, the cheapest provider that is competent for that request, so quality comes from where it is measured and cost is saved elsewhere. This chapter asks: can a router choose the cheapest competent provider for each request?
What we expect and why
Our starting hypothesis is that explicit rules capture most of the saving; learned routers add a small increment and a new failure mode.
Three papers frame it:
-
Ong et al. (2024) — RouteLLM learns routers between strong and weak LLMs and halves cost without quality loss. (Read at abstract level.)
-
Ding et al. (2024) — Hybrid LLM routes with a difficulty predictor that asks whether the small model suffices. (Read at abstract level.)
-
Hu et al. (2024) — RouterBench benchmarks routing with cost-quality curves and oracle routers — the evaluation shape this chapter uses. (Read at abstract level.)
The build
src/arbiter/router.py is a deterministic simulation. Providers have assumed linear accuracy-by-difficulty curves. The routers are always_best, always_cheapest, random, oracle (cheapest provider that meets the target at true difficulty), explicit_rule (routes via a metadata signal correlated with difficulty) and difficulty_pred. The router’s own latency is a per-request overhead added to the routed cost.
from arbiter.router import Provider, p_correct
# 1. Provider accuracy as a function of difficulty (ASSUMED curves).
print("1. provider curves (ASSUMED)")
cheap = Provider("cheap", 0.95, 0.40, cost=1.0)
strong = Provider("strong", 0.99, 0.90, cost=20.0)
for d in (0.0, 0.5, 1.0):
print(f" d={d}: cheap {p_correct(cheap, d):.2f} strong {p_correct(strong, d):.2f}")
# 2. The routing decision on one request.
print("2. one request, target 0.85")
d = 0.1
for p in (cheap, strong):
print(f" {p.name:<8} P(correct)={p_correct(p, d):.2f} cost={p.cost}")
print(" -> route cheap if it meets the target, else strong")
1. provider curves (ASSUMED)
d=0.0: cheap 0.95 strong 0.99
d=0.5: cheap 0.68 strong 0.95
d=1.0: cheap 0.40 strong 0.90
2. one request, target 0.85
cheap P(correct)=0.90 cost=1.0
strong P(correct)=0.98 cost=20.0
-> route cheap if it meets the target, else strong
The sweep
run_ch29.py sweeps the strong-provider cost ratio (5, 20), the signal correlation rho (0.0, 0.5, 0.95) and the router overhead, writing results/ch29.jsonl (mode: simulation, assumptions in every row). At cost ratio 20, quality target 0.85:
| router | rho | accuracy | cost/request |
|---|---|---|---|
| oracle | any | 0.912 | 11.06 |
| explicit_rule | 0.95 | 0.911 | 11.01 |
| explicit_rule | 0.0 | 0.869 | 11.21 |
| always_best | any | 0.945 | 20.00 |
| always_cheapest | any | 0.679 | 1.00 |
What it says
-
The oracle saves at most about 45% under these curves (20.00 → 11.06) and the explicit rule reaches it only when the signal is good: at rho 0.95 the rule is 11.01 vs the oracle’s 11.06, at rho 0.0 its accuracy falls to 0.869. The rule inherits the signal’s noise (P2), and “close to the oracle” is exactly a statement about how good the difficulty signal is (P1).
-
The difficulty predictor and the explicit rule coincide here by construction — both route on the same signal. That is the finding: a difficulty predictor is an explicit rule over a learnt signal; the only new failure mode it can have is a signal that is worse than the metadata one. P3 is a restatement, not a discovery.
-
Routing is a decision and its cost counts. With an overhead of 1.5 per request the explicit-rule total is about 12.5 vs 20 for always-best; routing still wins. With overhead near the per-request saving (11 vs 20, so ~9), always-cheapest-adequate takes over. P4 holds: the router stops paying for itself when its own cost approaches the saving.
-
always-cheapest is not competent (0.679 at target 0.85). Routing is not about always picking cheap; it is about picking the cheapest that is competent on this request, and the competence judgement is where the signal matters.
Wrong / Correct. Wrong: “Routing saves cost because small models are cheap.” Correct: “Routing saves the cost of competence: the cheap model on the easy requests, escalated elsewhere, with the saving bounded by how well the router can tell easy from hard — a signal-quality statement, not a model statement.”
What to carry forward
Chapter 30 turns the router’s escalation into a cascade: instead of choosing one provider per request, it runs a cheap stage and escalates the uncertain ones, under a target end-to-end risk. The router chapter’s lesson carries over: the value is bounded by the quality of the difficulty/uncertainty signal, which this chapter swept and did not measure.
Close by
When does routing cost more than it saves? When its own decision overhead approaches the per-request saving, and when its signal is too noisy to tell competence apart — both swept here, neither measured.
Limitations
- SIMULATION over assumed profiles;
results/ch29.jsonlismode: simulationwith the assumptions in every row. RouteLLM, Hybrid LLM and RouterBench are ABSTRACT_ONLY; none is measured. - The
explicit_ruleanddifficulty_predrouters coincide because both route on the same signal; a genuinely learnt predictor with separate noise would differ. - The provider curves are linear, which is an assumption; a different curve shape changes the numbers (the direction of the finding — signal quality bounds the saving — does not).
- The real router overhead and provider costs come from Chapter 28’s unmeasured cells; this chapter uses assumed costs instead and says so.