The Safe but Useless Model
Chapter 9 changed the unit of evaluation.
Instead of asking whether one answer looks good, we began studying a family of related executions.
That immediately reveals a failure that static evaluation can miss almost completely.
Consider two organizations.
ORGANIZATION A
3 months of runway
falling demand
negative cash flow
credit line nearly exhausted
ORGANIZATION B
5 years of runway
rapidly growing demand
strong margins
large cash reserve
Ask both:
What should management prioritize over the next twelve months?
Now imagine receiving essentially the same answer in both cases:
Invest in innovation.
Improve operational efficiency.
Stay close to customers.
Use data to guide decisions.
Build organizational capabilities.
Balance short-term execution with long-term growth.
There may be no fabricated fact. There may be no contradiction. The advice may be professionally worded and broadly sensible.
And yet something is badly wrong.
The answer barely depends on the problem.
That is the subject of this chapter.
A response can look safe, fluent, grounded, and still be useless because the decisive context never materially entered the decision.
The word safe in the title is deliberately informal. Here it means safe-looking under the static checks developed so far: no obvious fabrication, contradiction, or policy violation. It does not mean that an acceptance policy has certified the recommendation as safe for action.
This failure is broader than hallucination. It is a failure of context utilization.
Where we are
By this point the book can detect several distinct failures:
unsupported semantic extension
structural contradiction
role or polarity failure
instability under irrelevant transformations
insensitivity to declared counterfactual changes
Chapter 10 narrows the last category into an especially important production problem:
What happens when a model repeatedly converges to a polished default even when the task requires different decisions?
The objective is not to reward novelty. It is to test whether decision-relevant context has the leverage the task contract says it should have.
1. The failure is not falsehood
The easiest AI failures to notice are explicit errors:
wrong date
invented citation
reversed relation
unsupported claim
false tool state
The safe-but-useless failure is harder because every individual sentence can survive ordinary inspection.
A static evaluator may report:
Grammatically clear? PASS
Obvious factual fabrication? PASS
Internal contradiction? PASS
Obvious policy violation? PASS
Commonly accepted advice? PASS
The missing question is relational:
Would the decision have been substantially the same if the important facts had been different?
That question cannot be answered from one completion. We need the response-surface machinery from Chapter 9.
2. Context-insensitive convergence and the default response basin
Call the broader failure:
context-insensitive convergence
A system exhibits context-insensitive convergence when scenarios that the task contract declares decision-distinguishing repeatedly map to the same or nearly the same substantive decision.
The pattern may be:
the same recommendation
the same prioritization
the same ranking
the same trade-off
the same refusal
the same hybrid non-choice
Different inputs do not automatically require different outputs. Two patients may correctly receive the same treatment. Two companies may correctly choose the same liquidity intervention. Two bugs may need the same fix.
So low output diversity is not itself a failure.
Low sensitivity is evidence of genericity only when the oracle says the changed variable should alter the decision.
We will use default response basin as a descriptive behavioral term for a region of a declared scenario family that maps to the same structured decision pattern:
graph TD
X1[scenario x1] --> DRB[DEFAULT RESPONSE BASIN]
X2[scenario x2] --> DRB
X3[scenario x3] --> DRB
X4[scenario x4] --> DRB
X5[scenario x5] --> DRB
This is not a claim that the model contains a literal dynamical attractor.
Let:
If:
The text need not be similar.
prioritize customer-centric innovation
invest in differentiated customer value
build innovation capabilities around customer needs
may be three different phrasings of one decision.
Genericity is not textual similarity.
3. Three kinds of collapse
| Collapse | What remains the same? | Is it necessarily a reasoning failure? |
|---|---|---|
| Template collapse | rhetorical scaffold | no |
| Decision collapse | substantive recommendation | yes, when oracle requires divergence |
| Choice avoidance | non-committal hybrid | yes, when task requires exclusion |
A stable template can be useful.
Decision collapse is more serious. Different scenarios receive the same action even though the expected action should change.
Choice avoidance is subtler. The model converts a real trade-off into a balanced-sounding combination:
centralize strategically while decentralizing operationally
pursue radical innovation while maintaining incremental improvement
optimize the short term while investing for the long term
Sometimes a hybrid is genuinely correct. Sometimes combination is a way to evade a required choice.
The task contract must therefore declare the admissible output states:
A
B
HYBRID
ABSTAIN
and whether HYBRID is actually feasible under the stated constraints.
When the task requires exclusion, combination can be evasion.
4. Trendslop is one observed instance
In March 2026, Angelo Romasanta, Llewellyn D. W. Thomas and Natalia Levina reported a closely related pattern in strategic-advice experiments with leading LLMs.[1]
They tested seven recurring business tensions:
exploration vs exploitation
centralization vs decentralization
short-term vs long-term performance
competition vs collaboration
radical vs incremental innovation
differentiation vs commoditization
automation vs augmentation
Across thousands of simulations, the models repeatedly favored fashionable strategic positions rather than context-specific strategic logic. The authors called the pattern trendslop.[1][2]
The important part for this book is the experiment shape:
change scenario context
change framing
repeat across models
โ
observe persistent decision tendencies
That is a response-sensitivity experiment.
The broader technical class is context-insensitive convergence toward a recurring default recommendation basin.
Trendslop is one domain-specific manifestation.
5. Static evaluation can select for genericity
Generic answers can perform well under static evaluation because genericity reduces falsifiable commitment.
Compare:
Cut discretionary R&D immediately and preserve twelve months of payroll.
with:
Balance near-term efficiency with long-term innovation.
The second answer is harder to falsify. It is also harder to act on.
If an evaluator rewards:
fluency
non-toxicity
plausibility
broad helpfulness
absence of explicit factual error
then an evaluation regime can create selection pressure toward answers that minimize contestable commitment.
That does not imply the model consciously optimizes for genericity.
It is a systems-level observation about what kinds of outputs survive the objective.
A response can become safer to score by becoming less useful to decide with.
6. Mechanism hypotheses are not diagnoses
Why might a default basin exist?
Possible mechanisms include:
training-distribution frequency
preference optimization
prompt ambiguity
weak context utilization
evaluator pressure
uncertainty avoidance
These are hypotheses, not conclusions from the behavioral pattern.
It is tempting to blame RLHF or preference tuning directly for the Hybrid Trap. The evidence is not strong enough for that universal claim.
There is, however, relevant evidence that human-preference optimization can create undesirable conditioning behavior. Sharma and colleagues found that human preference data often favored responses matching a user’s stated views, and that optimizing against preference models could sometimes sacrifice truthfulness for sycophancy.[4] Earlier model-written evaluations also found sycophancy and other inverse-scaling behaviors associated with RLHF in some settings.[5]
That supports a narrower conclusion:
Preference objectives can shape which kinds of answers are rewarded, including answers that are agreeable rather than epistemically ideal. It does not establish that preference tuning is the primary cause of context-insensitive convergence.
Sycophancy is useful as a contrast case.
DEFAULT-BASIN FAILURE
underweights decisive local context
SYCOPHANCY
can overweight the user's stated preference
relative to evidence or truth
They are not perfect mirror images, but both show that conditioning can be weighted incorrectly.
Mechanism claims should be tested separately.
For example:
Hypothesis: prompt ambiguity drives hybrid answers
Test: require one mutually exclusive choice
Hypothesis: evaluator pressure drives non-commitment
Test: change judge rubric to reward decision specificity
Hypothesis: generic training frequency dominates
Test: vary framing while preserving the same decisive constraints
Behavior first. Mechanism second.
7. The temperature illusion
When outputs look repetitive, a common response is:
increase temperature
increase top_p
That may increase surface variation, and it may also change the distribution of decisions.
But neither effect proves that the model has become more context-sensitive.
You can easily obtain:
five different phrasings
five different rationales
one repeated recommendation
or, at higher randomness:
more decision variance
without better coupling to the scenario
So the relevant comparison is not:
low-temperature text diversity
vs
high-temperature text diversity
It is:
Does the conditional decision distribution move correctly when the decisive context changes?
Changing decoding parameters can be useful experimentally. It is not a substitute for a context-utilization test.
8. Context ablation should test epistemic behavior, not only decision change
One simple test is context ablation.
Start with:
FULL SCENARIO
company losing money
3 months runway
market contracting
bank covenant near breach
Then remove the decisive details:
ABLATION
A company wants strategic advice.
What should it prioritize?
A weak contract might demand:
decision must change
But that is too rigid. Removing decisive evidence can legitimately produce several outcomes:
same tentative action with lower confidence
more conditional language
request for missing information
abstention
broader recommendation
a different decision
A stronger ablation contract is therefore typed:
context_ablation = {
"removed_fields": [
"runway",
"cash_flow",
"market_direction",
"covenant_risk",
],
"expected_relation": {
"confidence": "decrease",
"conditionality": "increase",
"specificity": "not_increase",
"information_request_or_abstention": "allowed",
"unqualified_same_decision": "suspicious",
},
}
This seeds the central distinction of Chapter 11:
GENERIC COLLAPSE
strong evidence exists but is ignored
EPISTEMIC RESTRAINT
evidence is insufficient, so commitment decreases
Those are very different behaviors.
9. Counterfactual inversion measures decision uptake
A stronger test reverses a decisive field:
runway: 18 months โ 3 months
market: growth โ contraction
budget: $10M โ $100k
deadline: 1 year โ 1 week
risk tolerance: high โ near-zero
The perturbation contract specifies what should change:
contract = {
"name": "runway_inversion",
"intervention": "18_months -> 3_months",
"protected_invariants": ["product", "industry", "team_size"],
"expected_relation": {
"cash_preservation": "increase",
"optional_investment": "decrease",
"planning_horizon": "shorten",
},
}
Three different observations matter.
Context acknowledgment
Did the response notice the changed fact?
Decision effect
Did the substantive recommendation move?
Directional fidelity
Did it move as the oracle requires?
This separates four cases:
TOTAL INVARIANCE
fact not meaningfully reflected
COSMETIC ADAPTATION
fact mentioned, decision unchanged
ACCIDENTAL / WRONG-DIRECTION RESPONSIVENESS
decision moves, but incorrectly
APPROPRIATE ADAPTATION
decision and rationale move correctly
A sentence such as:
Given the shorter runway, innovation remains essential...
may pass context acknowledgment while failing decision uptake.
That is why keyword use is not evidence of reasoning.
10. Claimed decisive facts are hypotheses, not proof
A useful recommendation should expose its local dependency:
DECISION
DECISIVE FACTS
TRADE-OFF
REJECTED ALTERNATIVE
REVERSAL CONDITION
For example:
recommendation = {
"decision": "preserve_cash",
"decisive_facts": [
"runway_3_months",
"negative_cash_flow",
"credit_constraint",
],
"tradeoff": "slower_product_expansion",
"rejected_alternatives": [
{
"option": "large_new_product_bet",
"reason": "liquidity_constraint",
}
],
"reversal_condition": "runway_above_18_months_and_positive_cash_flow",
}
But self-reported rationales are not evidence that those facts actually drove the decision.
A model can produce a plausible post-hoc explanation.
So every claimed decisive fact should become an executable hypothesis:
model claims runway=3 months mattered
โ
intervene on runway
โ
regenerate
โ
did the decision move in the declared direction?
This produces a new measurement:
decisive-fact faithfulness
CLAIMED DECISIVE FACT
โ
COUNTERFACTUAL TEST
โ
SUPPORTED ATTRIBUTION
or
UNSUPPORTED RATIONALE
We do not need access to private chain-of-thought. We test the observable policy implied by the explanation.
11. Reversal conditions should be executed
The reversal condition is especially valuable because it is already a falsifiable statement.
Suppose the model says:
I would reverse this recommendation if runway exceeded 18 months
and cash flow became positive.
Construct that scenario.
Regenerate.
If the recommendation does not reverse, then:
declared response policy
โ
observed response policy
Call this:
reversal-condition fidelity
A good system should not merely state the conditions under which it would change its mind. It should actually change when those conditions are instantiated.
This turns explanation into executable regression testing.
12. Hybrid answers need an explicit exclusion contract
When the task requires a real trade-off, the evaluator needs to know which combinations are impossible or inadmissible.
Let:
For a recommendation set $A(Y)$, a task-specific hybrid conflict rate can be written:
This is not a universal AI score. It is useful only where the domain has an explicit exclusion structure.
A simpler benchmark may just report:
A rate
B rate
HYBRID rate
ABSTAIN rate
under a task contract that declares which states are admissible.
The key distinction remains:
LEGITIMATE HYBRID
both actions are jointly feasible and justified
HYBRID EVASION
the task requires commitment under scarcity,
but the model refuses to choose
13. Decision extraction is another sensor
All of the previous measurements assume that we can map free-form output into a structured decision.
That mapping:
A decision extractor therefore needs its own contract:
decision_extractor_contract = {
"input": "free_text_recommendation",
"output_classes": [
"PRESERVE_CASH",
"INVEST_FOR_GROWTH",
"HYBRID",
"ABSTAIN",
"UNKNOWN",
],
"method": "rubric_or_structured_extractor",
"validation": {
"human_audit": True,
"report_accuracy": True,
"report_agreement": True,
},
"known_blind_spots": [
"implicit_decision",
"conditional_recommendation",
"multi-stage_plan",
"rhetorical_paraphrase",
],
}
If $g(Y)$ is wrong, the basin analysis is wrong.
This is another instance of the book’s recurring rule:
Every observable needs a measurement contract.
For production systems, it is often useful to make the generator expose a structured decision payload directly:
from pydantic import BaseModel
from typing import Any
class DecisiveFact(BaseModel):
variable_name: str
observed_value: Any
expected_direction: str | None = None
class ReversalBoundary(BaseModel):
target_variable: str
threshold_condition: str
expected_new_decision: str
class StructuredDecisionPayload(BaseModel):
primary_decision: str
decisive_facts: list[DecisiveFact]
tradeoff: str | None = None
rejected_alternatives: dict[str, str]
reversal_boundaries: list[ReversalBoundary]
confidence: float | None = None
is_hybrid_choice: bool = False
Structured output does not force correctness. It makes the decision and its claimed dependencies inspectable.
14. Basin occupancy is only a marginal diagnostic
For discrete decisions:
But $ฮB$ is only a marginal statistic.
Suppose the oracle says:
scenario 1 โ A
scenario 2 โ A
scenario 3 โ B
scenario 4 โ B
and the model says:
scenario 1 โ B
scenario 2 โ B
scenario 3 โ A
scenario 4 โ A
Both distributions have dominant occupancy $0.5$.
The model is still wrong on every case.
So Chapter 10 needs three separate diagnostics:
MARGINAL CONVERGENCE
How concentrated are model decisions?
โ basin occupancy / entropy
SCENARIO AGREEMENT
Did each scenario receive the required decision?
โ oracle agreement
CONTEXT COUPLING
Does changing the decisive context move the decision distribution?
โ paired intervention response
Never substitute the first for the other two.
15. The scenario ร decision matrix is the primary basin artifact
For scenario families such as:
DISTRESS
STABLE
GROWTH
and decisions:
PRESERVE
OPTIMIZE
INVEST
report the conditional decision matrix:
| Scenario | Preserve | Optimize | Invest |
|---|---|---|---|
| Distress | 82% | 13% | 5% |
| Stable | 25% | 54% | 21% |
| Growth | 11% | 19% | 70% |
A context-insensitive system may instead produce something like:
| Scenario | Preserve | Optimize | Invest |
|---|---|---|---|
| Distress | 12% | 19% | 69% |
| Stable | 10% | 22% | 68% |
| Growth | 11% | 20% | 69% |
These numbers are illustrative, not results from our own experiment.
The real artifact should be compared with the oracle-required matrix and accompanied by uncertainty intervals.
For a deterministic extracted decision, scenario-level agreement is:
16. Mutual information measures context coupling, not correctness
For controlled scenario classes $X$ and extracted decisions $D$, one can estimate:
That makes mutual information a useful context-coupling diagnostic.
One caveat is easy to get wrong. The plug-in estimator of $I(X;D)$ from empirical counts is biased upward, and the bias grows with the number of cells (scenario classes times decision classes) relative to the sample size. With a handful of samples per scenario, a model that ignores context can still post a positive $\hat I(X;D)$ from noise alone. Report a bias-corrected estimate or a permutation baseline โ shuffle the scenario labels, recompute, and check that the observed $\hat I$ sits well above that null distribution โ before reading any coupling into the number.
One may also compare against the oracle:
But this is not a correctness metric.
A model can map each scenario class deterministically to the wrong decision and still have high mutual information.
So:
I(X;D)
โ how much decisions depend on context class
oracle agreement
โ whether that dependence is correct
Context utilization is relational. Dependence without directional correctness is not enough.
17. Match the distance to the decision object
Whole-answer cosine similarity is useful for exploration but often wrong for the actual decision.
Use a distance appropriate to the structured unit:
| Decision object | Example comparison |
|---|---|
| categorical action | exact match / confusion matrix |
| ranking | Kendall’s $\tau$ or rank distance |
| resource allocation | $L_1$ distance over allocation vector |
| risk level | ordinal distance |
| time horizon | interval / ordinal difference |
| accepted/rejected option | exact or set overlap |
| confidence | absolute difference / calibration analysis |
The same prose can encode different decisions, and different prose can encode the same decision.
Measure the smallest object that contains the expected effect.
18. Stochastic control uses decision distributions
For a stochastic generator:
Estimate instead:
The experiment must also report whether the mass moved in the oracle-required direction.
For each intervention family report:
samples per condition
within-condition decision entropy
paired decision-change rate
direction-correct rate
scenario-level oracle agreement
confidence intervals
perturbation validity
That separates true context sensitivity from sampling noise.
19. Context length and context utilization are different properties
A common response to generic output is:
Add more context.
Sometimes that works.
But a 5,000-token prompt is not evidence that the model used the five facts that determine the decision.
A direct test compares, for example:
SHORT
3 months runway
negative cash flow
LONG
5,000-token company description
containing the same decisive facts
Then intervene on the same decisive fact in both conditions.
The question is not:
Did the long prompt produce a longer answer?
It is:
Did the decisive variable have more, less, or the same directional effect on the decision?
The issue is not context length.
The issue is decision leverage.
20. Build a domain-specific anti-genericity benchmark
This chapter does not hand you one universal TrendslopScore.
It gives you the specification required to build a falsifiable benchmark for your domain.
A minimum suite might contain:
Decision-relevant inversions
runway: 18 months โ 3 months
budget: $10M โ $100k
market: growth โ contraction
deadline: 1 year โ 1 week
risk tolerance: high โ near-zero
regulation: permissive โ prohibitive
Context ablations
remove budget
remove runway
remove deadline
remove market direction
remove decisive evidence
Invariance controls
paraphrase
format
entity names
irrelevant ordering
Trade-off tests
A
B
HYBRID
ABSTAIN
with an explicit admissibility contract.
Claimed-rationale tests
For every self-reported decisive fact:
intervene
regenerate
test directional response
Reversal tests
For every declared reversal boundary:
instantiate boundary
regenerate
verify reversal
A compact benchmark design could use:
6 inversion types ร 2 directions = 12 paired conditions
5 ablations
4 invariance controls
3 samples per condition
The exact count depends on cost and domain. What matters is that the transformation family and oracle are frozen before final evaluation.
Required reporting should include:
perturbation validity rate
decision-extractor accuracy/agreement
scenario-level oracle agreement
context acknowledgment rate
paired decision-change rate
direction-correct rate
within-condition entropy
basin occupancy
excess basin occupancy
hybrid/non-choice rate
decisive-fact faithfulness
reversal-condition fidelity
This is enough to turn:
the answer feels generic
into a reproducible measurement problem.
Our earlier internal exploration of the Trendslop hypothesis supplied an initial perturbation taxonomy, semantic-divergence sketches, and response-manifold intuition. It did not produce the controlled book-owned benchmark required by the stricter Chapters 9โ10 protocol.
That is not a reason to invent a number.
It is a specification for the next experiment.
21. Production reality: use the surface selectively
Full counterfactual evaluation is expensive.
A base prompt plus five perturbations and three samples per condition already requires eighteen generations.
That may be appropriate for a benchmark. It may be absurd as a synchronous gate.
Separate:
CHEAP SYNCHRONOUS
containment
structured constraints
runtime verification
SHADOW
counterfactual suites on sampled traffic
TEMPLATE / WORKLOAD LEVEL
representative prompt-family testing
RISK TRIGGERED
high-impact decisions
low decision specificity
high uncertainty
novel scenario family
OFFLINE REGRESSION
known counterfactual families after changes
MODEL / CONFIG RELEASE GATE
full sensitivity suite before promoting
model, prompt, retrieval, or policy versions
A compact asynchronous shadow check can be as simple as:
import asyncio
async def shadow_counterfactual_check(
generate,
extract_decision,
base_prompt,
counterfactual_prompt,
expected_relation,
):
base_text, cf_text = await asyncio.gather(
generate(base_prompt),
generate(counterfactual_prompt),
)
base_decision = extract_decision(base_text)
cf_decision = extract_decision(cf_text)
return {
"base_decision": base_decision,
"counterfactual_decision": cf_decision,
"relation_passed": expected_relation(
base_decision,
cf_decision,
),
}
This is an orchestration sketch, not a production-complete evaluator. Real systems still need repeated sampling, perturbation validation, extractor validation, tracing, and uncertainty estimates.
22. Repair the system, not only the prose
Once context-insensitive convergence is detected, several interventions are possible.
Force explicit trade-offs
Require:
one primary recommendation
one rejected alternative
reason for rejection
Require structured decision output
A typed payload makes:
decision
decisive facts
tradeoff
rejected alternatives
reversal boundaries
observable.
This can act as a useful mechanical wedge against vague hybrid output, but it does not guarantee correct reasoning.
Generate alternatives before selecting
Use the model to widen the option set, then evaluate alternatives under explicit constraints.
Counterfactually verify the recommendation
Intervene on a claimed decisive fact and test whether the decision moves correctly.
Separate proposal from policy
Let the model propose possibilities. Let a later policy or optimization layer decide what is admissible.
Escalate ambiguous trade-offs
If the available evidence cannot distinguish the options, the right response may be review or abstention rather than a forced answer.
The last intervention leads directly to Chapter 11.
23. The safe-but-useless signature in the reliability record
The diagnostic record can now expose a failure that static hallucination checks miss:
containment = PASS
relation_fidelity = PASS
polarity_fidelity = PASS
provenance = VERIFIED
invariance_controls = PASS
context_acknowledgment = PASS
counterfactual_decision = FAIL
directional_fidelity = FAIL
context_ablation = FAIL
decisive_fact_faithfulness = FAIL
reversal_condition_fidelity = FAIL
decision_specificity = LOW
Nothing here says:
hallucination detected
Yet the output should not be trusted as a context-specific recommendation.
The pattern is:
containment = PASS
constraint fidelity = PASS
context sensitivity = FAIL
This is the signature Chapter 9 prepared us to observe.
24. What you should now be able to answer
After this chapter, you should be able to explain:
- Why different wording is not evidence of different decisions.
- Why basin occupancy is useful but cannot replace scenario-level oracle agreement.
- How context ablation differs from counterfactual inversion.
- Why mentioning a decisive fact does not prove the fact influenced the recommendation.
- How to test a model’s stated reversal condition.
- When a hybrid answer is a legitimate composition and when it is choice avoidance.
- Why higher temperature does not establish greater context sensitivity.
- Why more context tokens do not prove more context utilization.
Exercises
Exercise 1 โ Build a two-basin test
Construct ten distress scenarios and ten growth scenarios with clear oracle decisions. Generate multiple responses per scenario, extract one structured decision, and report:
scenario ร decision matrix
oracle agreement
basin occupancy
hybrid rate
direction-correct counterfactual rate
Exercise 2 โ Test a claimed decisive fact
Ask a model to provide a recommendation plus the three facts it considers decisive. Change one claimed fact while holding the others fixed. Does the recommendation move in the predicted direction?
Exercise 3 โ Execute a reversal condition
Require the model to state what condition would reverse its recommendation. Construct that condition and rerun the model. Record reversal-condition fidelity.
Exercise 4 โ Separate diversity from context use
Run the same scenario at several decoding temperatures. Then run two decision-distinguishing scenarios at one fixed temperature. Compare lexical diversity with decision-distribution movement. Which change actually tracks context?
25. The deeper lesson
Generic answers are attractive because they survive many evaluators. They are hard to falsify, sound reasonable, avoid controversial choices, and often contain familiar best practices.
But decision support is not a contest in producing statements that are difficult to disagree with.
It is a task of conditioning recommendations on constraints.
The relevant distinction is:
GOOD GENERAL ADVICE
can be broadly useful across many cases
GENERIC COLLAPSE
persists even when the task contract
says the decision should change
So the engineering rule is:
Do not ask only whether an answer is reasonable. Ask which facts made this answer different from the one the system would have produced anyway.
And then test those facts.
A fact is not demonstrated to be decision-relevant because the model mentions it.
A fact is demonstrated behaviorally when changing it changes the decision in the relation the task requires.
If we cannot observe that dependency, we may not have observed context-sensitive reasoning.
We may have observed a polished default.
Research roots
-
Angelo Romasanta, Llewellyn D. W. Thomas and Natalia Levina, “Researchers Asked LLMs for Strategic Advice. They Got ‘Trendslop’ in Return,” Harvard Business Review, March 16, 2026. Reports persistent strategic biases across seven core business tensions and warns that LLM recommendations can follow fashionable managerial patterns rather than context-specific strategic logic. https://hbr.org/2026/03/researchers-asked-llms-for-strategic-advice-they-got-trendslop-in-return
-
Natalia Levina, “Researchers Asked LLMs for Strategic Advice. They Got ‘Trendslop’ in Return,” NYU Stern Research Highlight, March 16, 2026. Summarizes the study as thousands of simulations in which leading LLMs repeatedly selected trendy strategic options across varied contexts. https://www.stern.nyu.edu/experience-stern/faculty-research/research-highlights/researchers-asked-llms-strategic-advice-they-got-trendslop-return
-
Harvard Business Review, product summary for Romasanta, Thomas and Levina, “Researchers Asked LLMs for Strategic Advice. They Got ‘Trendslop’ in Return,” 2026. Practical guidance includes using LLMs to expand options rather than make final strategic choices, counteracting biases, avoiding the hybrid trap, and not assuming that context alone removes the bias. https://store.hbr.org/product/researchers-asked-llms-for-strategic-advice-they-got-trendslop-in-return/H093GG
-
Mrinank Sharma et al., “Towards Understanding Sycophancy in Language Models,” 2023. Finds sycophancy across several state-of-the-art assistants, shows that human preference data can favor answers matching user views, and reports that preference-model optimization can sometimes trade truthfulness for sycophancy. https://arxiv.org/abs/2310.13548
-
Ethan Perez et al., “Discovering Language Model Behaviors with Model-Written Evaluations,” Findings of ACL 2023, pp. 13387โ13434. Uses model-written evaluations to identify behaviors including sycophancy and reports examples where RLHF exacerbated undesirable behaviors. https://aclanthology.org/2023.findings-acl.847/
Next: Knowing When Not to Answer
The safe-but-useless model exposes one more ambiguity.
Suppose the model does not adapt strongly to the context.
Why?
Possibility one:
the model ignored decisive facts
But there is another possibility:
the evidence genuinely does not justify a decisive answer
In that case, refusing to make a sharp recommendation may be correct.
We therefore need another question:
Does the system have enough evidence to answer at all?
That is not containment. It is not consistency. It is not sensitivity.
It is epistemic adequacy.
Chapter 11 turns abstention from an embarrassing non-answer into a measurable system capability.