← Agent Architectures: Advanced Strategies for Intelligent LLM Systems

Appendix B — Evidence and Claim Language

The evidence ladder and wording rules behind every evidence box in this book.

The ladder

E0  conceptual / illustrative — a sketch. Says "consider", never "measured".
E1  sourced external claim — a primary paper or doc, cited. Says "X reports".
E2  executable example — code or schema that ran once. Says "we ran".
E3  measured experiment — protocol written and frozen before the run, qualified observer, frozen result.
E4  replicated result — E3 repeated across conditions, with negative arms kept.
E5  mechanically verified property — deterministic check over a fixture set.

Wording rules

  • A critic preferring a candidate is not “better”. The required form is: under rubric R, candidate B scored higher than A in N seeded trials.
  • Agreement is not verification. Retrieval is not exposure; exposure is not influence; influence is not utility. Each link needs its own evidence.
  • Revision is a candidate until adjudicated. Scores are reported with disagreement, never as a single accuracy number.
  • Allowed shorthands: observed (E2+), measured (E3+ with numbers and budget), supported (E3+ beating a baseline), suggests (single-condition E3), consistent with (alternatives open), did not distinguish, inconclusive, not tested, speculative (E0, labeled).
  • Forbidden without E3 or higher: improves, helps, better, verifies, proves, preserves intent, earns overhead.

Failure is a result

Inconclusive and negative runs are retained, not re-rolled. Every excluded run carries a category (target / non-target / agent / environment / observer / evaluation failure, protocol violation, inconclusive), a reason, and an evidence pointer. Infrastructure failure is never classified as model failure.