Transfer Across Decisions
Design the decisive test of general decision capability: diverse tasks versus relabeled volume, with H3 held at insufficient evidence until it runs.
Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.
Everything so far has been about one decision at a time. Zero-Shot Decisions showed that runtime-defined labels buy unseen-label capability at a price. Tuning an Arbiter designed, but did not run, the comparison that decides whether tuning earns its cost. This chapter asks the question those chapters were always pointing at: when you train on some decisions, do unseen decisions get better, or did you just learn twenty separate classifications?
That is H3, general decision capability, and this chapter is its decisive experiment. Decisive does not mean triumphant. The literature contains both the strongest evidence for transfer and a direct measurement of transfer going backward. The design below is built to survive either outcome, and to tell the difference between task diversity and mere data volume, which is the confound that would otherwise swallow the result whole.
Because the run needs a GPU and a task catalog that does not exist yet, this chapter does something different from the earlier design drafts. It builds and tests the design itself: the registry, the k-subsets, the relabel control and a verifier that refuses a plan that confounds diversity with volume. The confounds are exactly the part that can be checked now, and one of them turned up while building the checks.
What the instruction-tuning literature actually says
The three required papers agree that instructions help. They disagree, usefully, about what “help” costs and where it stops.
Super-NaturalInstructions: breadth substitutes for scale
DOCUMENTED: Wang and colleagues assemble Super-NaturalInstructions: 1,616 tasks across 76 types, each with an expert-written instruction. Their Tk-Instruct model, meta-trained on instructions, outperforms InstructGPT by over 9% on their benchmark despite being an order of magnitude smaller. Generalisation grows log-linearly with the number of training tasks, while instances per task saturate at just 64. Task definition plus two positive examples is the effective encoding; explanations can hurt smaller models. [Wang et al., arXiv:2204.07705, Abstract, §§4–7.1]
The result most relevant to us is the substitution of diversity for scale: a T5-large trained on 757 tasks matches a T5-3B trained on 128 (48.0 against 48.4 ROUGE-L).
Flan: a scaling curve with a knee
DOCUMENTED: Chung and colleagues scale the same idea to 1.8K tasks and 540B parameters. Flan-PaLM beats PaLM by 9.4% normalised average across held-out benchmarks, and Flan-T5 11B beats PaLM 62B on BBH-direct (43.7% against 37.5%). [Chung et al., arXiv:2210.11416, Abstract, §§1, 3, 7]
But the scaling curve has a knee. Most task-count gains arrive by about 282 tasks, with diminishing though still positive returns after. Instruction tuning costs 0.2% of pretraining compute. And non-CoT instruction tuning degrades reasoning tasks unless CoT data is mixed in: joint non-CoT plus CoT finetuning restores reasoning while keeping the non-CoT gains. That is why our k-sweep keeps reasoning-adjacent tasks in the mix instead of testing transfer on a diet of pure classification. Breadth helps, but the help plateaus, and the wrong breadth actively harms.
Crowdsourced instructions: quality beats breadth alone
DOCUMENTED: Mishra and colleagues run the cleanest version of the argument with 61 crowdsourced tasks and 193K instances. Instructions buy 19% cross-task generalisation; without instructions, more seen tasks buy nothing. Leave-one-category-out still transfers. [Mishra et al., arXiv:2104.08773, Abstract, §§1, 5.1, 6.1–6.4, 7]
The ceiling is sobering: task-specific models average 66% against 32% for the best generaliser. Instruction elements are task-dependent (definitions help classification, examples help generation), and verification tasks barely move. Negative examples, surprisingly, can hurt: models do better without them. Breadth plus quality beats breadth alone.
What those results mean for this book
PROPOSED: Wang’s scale substitution touches the book’s compression thesis. If diversity can substitute for parameters for decision tasks too, then the compromise at the heart of this book, small models deciding well, gets cheaper through task breadth rather than parameter count. But substitution is not transfer. Matching a bigger model’s score is consistent with learning twenty classifications well, which is exactly what the relabel control below exists to expose. The book will not let a scale-substitution number do transfer’s evidential work.
PROPOSED: The negative-examples finding constrains our task registry more than it first appears. If “things to avoid” and negative demonstrations can degrade transfer, then instruction quality is a treatment arm in its own right, and an uncontrolled one in most transfer studies. JEV-15-01 holds instruction format fixed (definition plus two positive examples, the Tk-Instruct-effective encoding) so that k measures task diversity rather than formatting luck. A future arm may vary instruction elements; this one must not, or P3’s diversity claim dissolves into a wording confound of the kind Chapter 5 already measured.
Two newer papers supply the guardrails
DOCUMENTED: Jung and Jung test 51 task mixtures across three 7 to 8B model families. With small data, 1-2 task mixtures win; with larger budgets, 3-4 task mixtures win. Programming and math synergise; programming and creative writing interfere. Effects stabilise near 750 examples, and the winners are model-specific. Their conclusion cuts against any naive “more tasks always help” reading: increasing diversity has limited benefit in some settings, and the wrong mixture is worse than a narrow one. [Jung & Jung, Findings of ACL 2026, Abstract, §§4.3–4.6, 8/Limitations]
DOCUMENTED, and it contradicts the optimistic half outright: Mueller and colleagues instruction-tune T5-Large on 619 Super-NaturalInstructions tasks and find that transfer to 93 held-out tasks gets worse than the pretrained base: catastrophic forgetting. Transfer ability and in-context generalisation correlate strongly (Spearman-R 0.705), and proportional sampling plus continued unsupervised-data training mitigate the damage. [Mueller et al., Findings of ACL 2024, Abstract, §§1–2.4, Limitations]
Their sampling analysis is the methodological point our design steals: uniform sampling overfits low-resource tasks, reverse-annealing underfits them, and only proportional sampling balances both. Multi-task training can therefore destroy exactly what it is supposed to build, and how you sample the mix decides which. Any transfer experiment without an untuned-base baseline cannot distinguish “transfer failed” from “tuning hurt.”
Wrong: “Train on more decisions and unseen decisions improve; the only question is how many.”
Correct: “Diversity, volume, instruction quality, mixture composition and forgetting all move the needle, and forgetting can move it backward. The experiment must separate them.”
flowchart TD
A[base model] --> B[train on k diverse tasks]
A --> C[train on 20 relabelings of 1 task]
A --> D[no tuning]
B --> E[held-out related tasks]
B --> F[held-out unrelated tasks]
C --> E
D --> E
D --> F
E --> G{diversity beats volume?}
What the earlier chapters already fixed
OBSERVED: Chapter 5 froze the materials this chapter reuses: seen-60 / unseen-17 BANKING77 (2,400 and 680 test items), with LR refusing all 680 unseen-label requests and the size-matched control putting seen ahead. Those rows are why the unseen-17 pool is a legitimate held-out intent task with zero training rows, and why candidate-set size must be controlled whenever label counts differ.
OBSERVED: Chapter 4’s earn rule governs the cost half: a tuned provider must beat the boring baseline by more than about three points with an interval, or match it more cheaply.
PROPOSED: Chapter 14 designed the tuning comparison but did not run it. JEV-15-01 therefore fixes its own training method now (LoRA r=8, learning rate 2e-4, 3 epochs) rather than waiting for Chapter 14’s unresolved winner. Depending on a deferred result would let this chapter’s design inherit another chapter’s delay. Fixing the method now has a second virtue: if Chapter 14 later crowns a different winner, the transfer result still stands as a clean statement about this method, and rerunning the sweep under the new winner becomes a dated amendment rather than a rescue.
PROPOSED: The task registry is a schema plus a plan, not data pretending to be data. Only three datasets are pinned locally, so a 20-task catalog cannot honestly be listed yet. The preregistration freezes what can be frozen now: the schema, the held-out tasks that already exist, the relatedness rubric, the relabel construction, the budgets and the seeds. The catalog (task IDs, questions, answer sets, pinned revisions) is frozen in a separate step before the run, and recorded per results row. A transfer experiment whose training tasks are chosen after seeing results would be worthless; the freeze step is what makes the future run checkable, and the loader refuses unpinned sources.
The design
PROPOSED: JEV-15-01 trains LoRA-tuned Arbiter-2 on k tasks for k in {2, 5, 10, 20} and evaluates on frozen held-out tasks:
- unseen-17 intent (17-way, related, new labels);
- a 5-way seen-label subset (related, new label-set size);
- safety shift plus unrelated-family tasks (unrelated).
The hard control trains on 20 isomorphic relabelings of one 77-way intent task: seeded label-ID permutations over identical inputs and identical total volume. Diversity and volume then differ while everything else is held constant.
Relatedness is predeclared structurally, not judged after the fact. Related means sharing an answer-type or skill family with at least one training task; unrelated means a different family. The rubric is applied to the frozen registry before evaluation. Per-task accuracy and top-label ECE are reported and never pooled across different k, with macro averages over the related and unrelated groups and diversity-minus-relabel differences with bootstrap intervals. Seeds 0 to 2 throughout; test pools are evaluated once.
Three mechanics each guard a different confound:
- The relabel control permutes label IDs, not label semantics: the same inputs, the same class frequencies, the same total volume, and only the mapping from input to answer changes. If diverse-20 still wins, the win cannot be volume, frequency or input coverage; it must be cross-task structure.
- ECE is reported per task because a 77-way and a binary task have different chance floors. Pooling them would let the easy task launder the hard one’s miscalibration.
- The Llama confirmatory arm runs only k=20 plus the control: enough to check that the verdict is not a Qwen artefact, cheap enough not to triple the GPU bill. A confirmatory arm that reran the whole sweep would be rigor theatre at someone else’s compute expense.
The design, made executable
The design is src/arbiter/transfer_design.py, tested by 12 tests in tests/transfer/. Everything below is examples/ch15-transfer-across-decisions/walkthrough_ch15.py, which you can run as it stands. The registry is illustrative toy metadata: 25 invented task specs with families, label sets and pinned revisions. It is not the real catalog. What matters is the checks, which apply unchanged to the real registry when it exists.
The setup
from arbiter.transfer_design import (
Registry,
RegistryError,
TaskSpec,
apply_relabeling,
build_diverse_plan,
build_relabeled_plan,
has_unrelated_heldout,
heldout_relatedness,
nested_subsets,
relabelings,
relatedness,
verify_plan,
volume_matched,
)
FAMILIES = ["intent", "intent", "safety", "sentiment", "topic", "intent"]
def toy_registry() -> Registry:
tasks = [
TaskSpec(
task_id=f"toy-{i:02d}",
family=FAMILIES[i % len(FAMILIES)],
question=f"toy question {i}?",
labels=tuple(f"c{j}" for j in range(3 + i % 4)),
revision=f"rev-{i:02d}",
n_examples=500,
)
for i in range(24)
]
tasks.append(TaskSpec("toy-intent77", "intent", "which intent?",
tuple(f"i{j}" for j in range(77)), "rev-77", 500))
return Registry(tuple(tasks))
def main() -> None:
registry = toy_registry()
heldout = ["toy-00", "toy-02"] # an intent task and a safety task
common = dict(per_task_examples=100, base_model_id="toy/model", base_model_rev="r0")
1. The registry refuses unpinned tasks
A task without a pinned data revision cannot be registered at all. This is the loader’s rule moved up to where a plan is made, so an unfreezable task never enters a plan.
# 1. A task without a pinned revision cannot be registered at all.
print("1. the registry refuses unpinned tasks")
try:
TaskSpec("toy-x", "intent", "q?", ("a", "b"), revision="", n_examples=10)
except RegistryError as exc:
print(f" {exc}")
2. Nested k-subsets
Each point on the sweep adds tasks to the one before it, so a change between points is attributable to the added tasks. Held-out tasks are removed from the pool first.
# 2. Nested k-subsets: each point on the sweep adds tasks to the one before.
print("2. nested k-subsets (held-out tasks excluded)")
subsets = nested_subsets(registry, heldout, [2, 5, 10, 20], seed=0)
for k, ids in subsets.items():
print(f" k={k:<2} -> {len(ids)} tasks, first three {list(ids[:3])}")
print(f" k=2 inside k=5 inside k=20: {set(subsets[2]) <= set(subsets[5]) <= set(subsets[20])}")
print(f" any held-out task in training: {bool(set(heldout) & set(subsets[20]))}")
3. A trap in the relatedness rubric
The rubric is structural, so it can be computed from the registry alone, before any result exists. Computing it exposed a flaw that is easy to miss on paper. With 20 training tasks drawn from a pool of 22, every family appears in training, and so every held-out task comes out related. P2, which says unrelated tasks should stay put, then has nothing to test. The fix is a design requirement: reserve at least one whole family for held-out tasks, so that it never appears in training.
# 3. The relatedness rubric is structural, which exposes a design trap.
print("3. relatedness, from the registry alone")
plain = build_diverse_plan(registry, heldout, k=20, seed=0, **common)
for tid, label in heldout_relatedness(plain, registry).items():
print(f" {tid} ({registry.get(tid).family}) -> {label} [no family reserved]")
print(f" anything unrelated to test P2 with: {has_unrelated_heldout(plain, registry)}")
reserved = build_diverse_plan(registry, heldout, k=20, seed=0, reserved_families=["safety"], **common)
for tid, label in heldout_relatedness(reserved, registry).items():
print(f" {tid} ({registry.get(tid).family}) -> {label} [safety reserved]")
print(f" anything unrelated to test P2 with: {has_unrelated_heldout(reserved, registry)}")
4. The relabel control
A relabeling is a permutation of label IDs: the same inputs, the same class frequencies, a different mapping from input to answer. The check shows class frequencies unchanged and the mapping changed, and that 20 relabelings of 77 labels are distinct and none is the identity.
# 4. The relabel control: same inputs, same class frequencies, new mapping.
print("4. relabelings keep inputs and class frequencies")
labels = [0, 0, 1, 2, 2, 2, 1, 0]
perm = relabelings(3, 1, seed=3)[0]
out = apply_relabeling(labels, perm)
freq = lambda xs: sorted(xs.count(c) for c in set(xs))
print(f" permutation {perm}: {labels} -> {out}")
print(f" class frequencies before {freq(labels)}, after {freq(out)}, mapping changed {out != labels}")
twenty = relabelings(77, 20, seed=0)
print(f" 77 labels, 20 relabelings: {len(set(twenty))} distinct, identity among them: {tuple(range(77)) in twenty}")
5. The decisive comparison at k=20
The diverse arm and the relabeled arm are both verified and have the same total volume: 20 tasks of 100 examples each, against one task relabeled 20 times at 100 examples each. Only k=20 is the controlled comparison; at k=10 the volumes differ, so a difference there would confound diversity with volume.
# 5. The two arms of the decisive comparison, verified and volume-matched.
print("5. the decisive comparison at k=20")
diverse = build_diverse_plan(registry, heldout, k=20, seed=0, reserved_families=["safety"], **common)
relabeled = build_relabeled_plan(registry, heldout, "toy-intent77", 20, seed=0,
reserved_families=["safety"], **common)
print(f" diverse : {len(set(diverse.train_task_ids))} tasks x {diverse.per_task_examples} = {diverse.total_examples} examples, violations {verify_plan(diverse, registry)}")
print(f" relabeled : 1 task x 20 relabelings x {relabeled.per_task_examples} = {relabeled.total_examples} examples, violations {verify_plan(relabeled, registry)}")
print(f" volume matched: {volume_matched(diverse, relabeled)}")
small = build_diverse_plan(registry, heldout, k=10, seed=0, reserved_families=["safety"], **common)
print(f" k=10 vs relabeled-20 volume matched: {volume_matched(small, relabeled)} (only k=20 is the controlled comparison)")
6. The verifier earns its place
A verifier that has never failed proves nothing. Each of these mistakes is deliberate, and each is caught by name.
# 6. The verifier earns its place by catching the confounds.
print("6. the verifier catches deliberate mistakes")
leaky = dataclasses.replace(diverse, train_task_ids=diverse.train_task_ids[:-1] + ("toy-00",))
print(f" held-out task leaked into training -> {verify_plan(leaky, registry)}")
greedy = build_diverse_plan(registry, heldout, k=5, seed=0, per_task_examples=900,
base_model_id="toy/model", base_model_rev="r0")
stray = dataclasses.replace(reserved, train_task_ids=reserved.train_task_ids[:-1] + ("toy-08",))
print(f" asking for more examples than exist -> {verify_plan(greedy, registry)[0]}")
print(f" a reserved-family task in training -> {verify_plan(stray, registry)}")
dup = dataclasses.replace(relabeled, relabel_perms=relabeled.relabel_perms[:-1] + (relabeled.relabel_perms[0],))
print(f" a duplicated relabeling -> {verify_plan(dup, registry)}")
What it prints
1. the registry refuses unpinned tasks
toy-x: unpinned revision; refuse to register
2. nested k-subsets (held-out tasks excluded)
k=2 -> 2 tasks, first three ['toy-22', 'toy-23']
k=5 -> 5 tasks, first three ['toy-22', 'toy-23', 'toy-13']
k=10 -> 10 tasks, first three ['toy-22', 'toy-23', 'toy-13']
k=20 -> 20 tasks, first three ['toy-22', 'toy-23', 'toy-13']
k=2 inside k=5 inside k=20: True
any held-out task in training: False
3. relatedness, from the registry alone
toy-00 (intent) -> related [no family reserved]
toy-02 (safety) -> related [no family reserved]
anything unrelated to test P2 with: False
toy-00 (intent) -> related [safety reserved]
toy-02 (safety) -> unrelated [safety reserved]
anything unrelated to test P2 with: True
4. relabelings keep inputs and class frequencies
permutation (1, 2, 0): [0, 0, 1, 2, 2, 2, 1, 0] -> [1, 1, 2, 0, 0, 0, 2, 1]
class frequencies before [2, 3, 3], after [2, 3, 3], mapping changed True
77 labels, 20 relabelings: 20 distinct, identity among them: False
5. the decisive comparison at k=20
diverse : 20 tasks x 100 = 2000 examples, violations []
relabeled : 1 task x 20 relabelings x 100 = 2000 examples, violations []
volume matched: True
k=10 vs relabeled-20 volume matched: False (only k=20 is the controlled comparison)
6. the verifier catches deliberate mistakes
held-out task leaked into training -> ["held-out task(s) in training: ['toy-00']"]
asking for more examples than exist -> toy-12: only 500 examples, plan needs 900
a reserved-family task in training -> ["toy-08: family 'safety' is reserved for held-out tasks"]
a duplicated relabeling -> ['duplicate relabelings']
Read it in order. Section 3 is the one to dwell on: without a reserved family there is nothing unrelated to test against, and with safety reserved the safety held-out task becomes unrelated while the intent one stays related. Section 5 shows the two arms at 2,000 examples each, which is what lets a difference between them be attributed to diversity. Section 6 shows leakage, over-asking, a reserved-family leak and a duplicated relabeling each reported as a specific violation.
The experiment, designed but not run
JEV-15-01 is preregistered with status: NOT_RUN; the full block is in metadata/15-chapter.yaml. The reserved-family requirement from section 3 sharpens the design before any run and is recorded in the ledger.
Predictions (HYPOTHESIS):
- P1. At k=20, related held-out tasks gain at least 5 accuracy points over the untuned base (related mean, 3 seeds).
- P2. Unrelated tasks stay within ±2 points of base at every k.
- P3. Diverse-20 beats relabeled-20 by at least 3 points on the related mean; a smaller gap means volume, not diversity, did the work.
- P4. Gains diminish: the k=10 to 20 related gain is smaller than the k=2 to 10 gain.
PROPOSED: k stops at 20 for two reasons, one statistical and one honest. Statistically, the literature’s knees live far above our range (282 tasks for Flan, hundreds for Tk-Instruct), so a 2-to-20 sweep cannot observe the literature’s plateau directly. P4 tests diminishing gains within our range instead. Honestly, twenty decision tasks with pinned revisions is already the largest mix this book can assemble without inventing data. A k=50 arm built on unpinned downloads would trade catalog integrity for curve smoothness, which is exactly the trade the harness loader was written to refuse.
Refutation: P1 fails if related tasks do not improve. P2 fails if unrelated tasks move. P3 fails if relabelings match diversity. P4 fails if late gains match early ones. Sustained negative transfer against base is reported as evidence against the H3 direction, per Mueller, and not as a failed run.
PENDING_RUN: result for JEV-15-01, P1–P4 Command:
python examples/ch15-transfer-across-decisions/run_ch15.py --k 2,5,10,20 --control relabeled-20 --seeds 0,1,2Fills:results/ch15.jsonl,static/figures/ch15-heldout-vs-k-*.pngStage plan (each ≤15 min, GPU-bound, resumable; adapters outside git): S1 freeze registry, k-subsets, and 20 relabelings (now checkable withverify_plan); S2 tune k-sweep + control shards; S3 verify freeze and provenance; S4 evaluate each held-out task once with intervals; S5 write rows and figures.
What would change your mind
On the evidence available now, the literature plus the measured rows of Chapters 1 to 9, H3 is INSUFFICIENT_EVIDENCE, and this chapter leaves it there. That is the honest answer to the prompt’s closing question, and planning/hypotheses.md is untouched accordingly.
What would change it:
- P1 and P3 holding with P2 intact would support H3 directionally (partially, pending replication).
- Relabelings matching diversity would refute the diversity mechanism while leaving volume effects open.
- Outright negative transfer would count against the H3 direction.
A fourth outcome is worth naming because Jung makes it plausible: transfer that is real but spotty, with some held-out tasks up, some down and the mean flat. That is neither support nor refutation; it is interference, and the per-task table is what lets the book say so instead of averaging it away.
No pivot is proposed: a design supplies no new measured evidence.
Limitations
JEV-15-01 is NOT_RUN; results/ch15.jsonl does not exist and the 20-task catalog is not assembled. The toy registry in the walkthrough is illustrative and says nothing about any real task.
Required papers are PARTIAL by section; appendices, code and full ablations were not reproduced. The added sources carry their own bounds: Jung and Jung is 7 to 8B direct-answer tuning, and Mueller is T5-Large on one dataset. The training method is fixed to LoRA rather than Chapter 14’s unresolved winner, which keeps the transfer question clean at the cost of generality. Sparse low-k budgets will miss classes and tasks; missing coverage is reported, not smoothed. ECE across different k is reported per task and never pooled.
Three further bounds stay visible. First, the relabel control is one family of volume controls (permuted IDs), and a skeptic may demand a second, such as repeated resampling of one task; the design records the family it uses rather than claiming exhaustiveness. Second, only Qwen3-1.7B runs the full k-sweep, so a Qwen-specific verdict with Llama confirmation is the most this run can license. Third, held-out tasks number in the single digits per group, so intervals will be wide and the chapter must report them wide.
What the next chapters inherit
Chapter 16 inherits the transfer-tested models only conditionally. If JEV-15-01 ever shows general decision capability, the decision expression builds on tuned models with known transfer; until then it builds on the cheaper providers.
Chapter 28 inherits the diversity-versus-volume control as the deployment-side complement: Chapter 15 tests whether training transfers, and Chapter 28 tests whether it matters where providers win.
The Part IV arc closes honestly whatever happens. Probes asked what is readable, tuning asked what is worth paying for, and transfer asks what generalises. Until the GPU runs happen, all three answers stay open, and the open state is written down, not wished away. The book also inherits a sharper H3: not “do models transfer?” but “which relatedness, at what k, against which volume control, on which base?” That is a question an experiment can answer, and this one is designed to.