<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Evaluation Integrity on Programmer.ie</title>
    <link>http://programmer.ie/tags/evaluation-integrity/</link>
    <description>Recent content in Evaluation Integrity on Programmer.ie</description>
    <generator>Hugo</generator>
    <language>en-us</language>
    <lastBuildDate>Wed, 26 Aug 2026 10:00:00 +0100</lastBuildDate>
    <atom:link href="http://programmer.ie/tags/evaluation-integrity/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Evidence and Verification</title>
      <link>http://programmer.ie/books/agents-from-first-principles/10-chapter/</link>
      <pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate>
      <guid>http://programmer.ie/books/agents-from-first-principles/10-chapter/</guid>
      <description>&lt;p&gt;The agent says:&lt;/p&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;Done.&lt;/p&gt;&lt;/blockquote&gt;&#xA;&lt;p&gt;That is a claim.&lt;/p&gt;&#xA;&lt;p&gt;It is not evidence.&lt;/p&gt;&#xA;&lt;p&gt;Every mechanism in this book so far has made the agent better at deciding what to do, and none of them establishes that the user&amp;rsquo;s goal was achieved. A planner can produce a coherent plan for the wrong problem. A tool can return exit code zero without producing the intended effect. A search can select the highest-scoring branch when every branch is wrong. A memory system can retrieve a perfectly relevant fact that stopped being true in March.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Building the Complete Agent</title>
      <link>http://programmer.ie/books/agents-from-first-principles/11-chapter/</link>
      <pubDate>Wed, 26 Aug 2026 10:00:00 +0100</pubDate>
      <guid>http://programmer.ie/books/agents-from-first-principles/11-chapter/</guid>
      <description>&lt;p&gt;Every mechanism in this book was argued against a problem chosen to isolate it.&lt;/p&gt;&#xA;&lt;p&gt;That isolation was deliberate, and it was also a form of protection. The memory chapter picked a task where recall was the bottleneck, held everything else still, and measured the one thing it came to measure. The result is a clean explanation and a weak claim. Nothing in it establishes that the same retrieval policy behaves when a search controller is expanding forty nodes, or when a verifier insists that every piece of evidence carry a state identity the search controller has never heard of.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
