← Jev From First Principles

The Decision Router

Can a router choose the cheapest competent provider for each request?

The Decision Router

This chapter is a simulation over assumed provider profiles. The provider accuracy-by-difficulty curves, costs and signal correlation are assumed and swept; every number is a consequence of them. No real router (RouteLLM, Hybrid LLM) is measured. The assumptions sit next to the results.

The problem

Chapter 28 ended with a table mostly full of NOT_OBSERVED. A router makes that honest gap useful: it chooses, per request, the cheapest provider that is competent for that request, so quality comes from where it is measured and cost is saved elsewhere. This chapter asks: can a router choose the cheapest competent provider for each request?

What we expect and why

Our starting hypothesis is that explicit rules capture most of the saving; learned routers add a small increment and a new failure mode.

Three papers frame it:

  1. Ong et al. (2024) — RouteLLM learns routers between strong and weak LLMs and halves cost without quality loss. (Read at abstract level.)

  2. Ding et al. (2024) — Hybrid LLM routes with a difficulty predictor that asks whether the small model suffices. (Read at abstract level.)

  3. Hu et al. (2024) — RouterBench benchmarks routing with cost-quality curves and oracle routers — the evaluation shape this chapter uses. (Read at abstract level.)

The build

src/arbiter/router.py is a deterministic simulation. Providers have assumed linear accuracy-by-difficulty curves. The routers are always_best, always_cheapest, random, oracle (cheapest provider that meets the target at true difficulty), explicit_rule (routes via a metadata signal correlated with difficulty) and difficulty_pred. The router’s own latency is a per-request overhead added to the routed cost.

from arbiter.router import Provider, p_correct
    # 1. Provider accuracy as a function of difficulty (ASSUMED curves).
    print("1. provider curves (ASSUMED)")
    cheap = Provider("cheap", 0.95, 0.40, cost=1.0)
    strong = Provider("strong", 0.99, 0.90, cost=20.0)
    for d in (0.0, 0.5, 1.0):
        print(f"   d={d}: cheap {p_correct(cheap, d):.2f}  strong {p_correct(strong, d):.2f}")
    # 2. The routing decision on one request.
    print("2. one request, target 0.85")
    d = 0.1
    for p in (cheap, strong):
        print(f"   {p.name:<8} P(correct)={p_correct(p, d):.2f} cost={p.cost}")
    print("   -> route cheap if it meets the target, else strong")
1. provider curves (ASSUMED)
   d=0.0: cheap 0.95  strong 0.99
   d=0.5: cheap 0.68  strong 0.95
   d=1.0: cheap 0.40  strong 0.90
2. one request, target 0.85
   cheap    P(correct)=0.90 cost=1.0
   strong   P(correct)=0.98 cost=20.0
   -> route cheap if it meets the target, else strong

The sweep

run_ch29.py sweeps the strong-provider cost ratio (5, 20), the signal correlation rho (0.0, 0.5, 0.95) and the router overhead, writing results/ch29.jsonl (mode: simulation, assumptions in every row). At cost ratio 20, quality target 0.85:

router rho accuracy cost/request
oracle any 0.912 11.06
explicit_rule 0.95 0.911 11.01
explicit_rule 0.0 0.869 11.21
always_best any 0.945 20.00
always_cheapest any 0.679 1.00

What it says

  1. The oracle saves at most about 45% under these curves (20.00 → 11.06) and the explicit rule reaches it only when the signal is good: at rho 0.95 the rule is 11.01 vs the oracle’s 11.06, at rho 0.0 its accuracy falls to 0.869. The rule inherits the signal’s noise (P2), and “close to the oracle” is exactly a statement about how good the difficulty signal is (P1).

  2. The difficulty predictor and the explicit rule coincide here by construction — both route on the same signal. That is the finding: a difficulty predictor is an explicit rule over a learnt signal; the only new failure mode it can have is a signal that is worse than the metadata one. P3 is a restatement, not a discovery.

  3. Routing is a decision and its cost counts. With an overhead of 1.5 per request the explicit-rule total is about 12.5 vs 20 for always-best; routing still wins. With overhead near the per-request saving (11 vs 20, so ~9), always-cheapest-adequate takes over. P4 holds: the router stops paying for itself when its own cost approaches the saving.

  4. always-cheapest is not competent (0.679 at target 0.85). Routing is not about always picking cheap; it is about picking the cheapest that is competent on this request, and the competence judgement is where the signal matters.

Wrong / Correct. Wrong: “Routing saves cost because small models are cheap.” Correct: “Routing saves the cost of competence: the cheap model on the easy requests, escalated elsewhere, with the saving bounded by how well the router can tell easy from hard — a signal-quality statement, not a model statement.”

What to carry forward

Chapter 30 turns the router’s escalation into a cascade: instead of choosing one provider per request, it runs a cheap stage and escalates the uncertain ones, under a target end-to-end risk. The router chapter’s lesson carries over: the value is bounded by the quality of the difficulty/uncertainty signal, which this chapter swept and did not measure.

Close by

When does routing cost more than it saves? When its own decision overhead approaches the per-request saving, and when its signal is too noisy to tell competence apart — both swept here, neither measured.

Limitations

  • SIMULATION over assumed profiles; results/ch29.jsonl is mode: simulation with the assumptions in every row. RouteLLM, Hybrid LLM and RouterBench are ABSTRACT_ONLY; none is measured.
  • The explicit_rule and difficulty_pred routers coincide because both route on the same signal; a genuinely learnt predictor with separate noise would differ.
  • The provider curves are linear, which is an assumption; a different curve shape changes the numbers (the direction of the finding — signal quality bounds the saving — does not).
  • The real router overhead and provider costs come from Chapter 28’s unmeasured cells; this chapter uses assumed costs instead and says so.