<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>the wizard · dave8172</title><link>https://quirkyagents.com/wizard/blog/</link><description>Benchmarks and teardowns from real AI builds.</description><atom:link href="https://quirkyagents.com/wizard/blog/rss.xml" rel="self" type="application/rss+xml"/><item><title>Building and using MCP servers: 4 things I learned</title><link>https://quirkyagents.com/wizard/blog/building-mcp-servers-lessons/</link><guid>https://quirkyagents.com/wizard/blog/building-mcp-servers-lessons/</guid><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><description>How MCP works, one tool call step by step, and four things I learned building and using MCP servers, from tool descriptions to where secrets live.</description><category>MCP</category><category>Agents</category><category>Claude Code</category><category>Python</category><content:encoded><![CDATA[<p>I use Upwork&#39;s official MCP server from Claude Code, and doceval, my open-source eval tool, ships an MCP server of its own. This is how MCP works underneath, and four things I learned along the way. Where I tested something, I say how.</p>
<hr>
<h2 id="mcp-in-one-paragraph">MCP in one paragraph</h2>
<p>MCP (Model Context Protocol) is an open standard for connecting AI apps to outside tools and data. A service wraps itself once as an MCP server, and any app that speaks MCP (Claude, ChatGPT, Cursor and many more) can connect to it without a custom integration.</p>
<p>If every app had to write its own integration for every service, 20 apps and 1,000 services would need <strong class="spell">20,000 integrations.</strong> With MCP each side builds its half once: 20 clients plus 1,000 servers, 1,020 in all. It&#39;s the USB-C trick: one port, any charger. Like USB-C, the port says nothing about how good the charger is.</p>
<p><img src="https://quirkyagents.com/wizard/blogassets/mcp-why-it-exists-tall.1db16dad.png" alt="Without MCP, 3 apps and 3 services need 9 custom integrations; with MCP, 3 clients and 3 servers. At 20 apps and 1,000 services: 20,000 integrations, or 1,020 with MCP" width="1200" height="1500" loading="lazy" decoding="async"></p>
<h2 id="three-roles">Three roles</h2>
<ul>
<li><strong>Host</strong>: the app running the model, like Claude Code, ChatGPT or an agent you build yourself.</li>
<li><strong>Client</strong>: a connector inside the host, one per server. It passes JSON back and forth.</li>
<li><strong>Server</strong>: wraps one service and offers what it can do as tools.</li>
</ul>
<p>Usually the host&#39;s model decides which tool to call. The server can be plain code (get a request, do the thing, return the result) or run its own model inside. The host can&#39;t tell the difference: a request goes in, a result comes out.</p>
<p>I saw this while building <a href="https://quirkyagents.com/wizard/projects/upsweep/">upsweep</a>, an open-source agent skill that finds Upwork jobs through Upwork&#39;s official MCP server. Claude, running inside Claude Code, reads Upwork&#39;s tool list, decides what to call and makes sense of what comes back. My own subscription pays for that model. Upwork pays nothing for it.</p>
<h2 id="one-call-step-by-step">One call, step by step</h2>
<p><img src="https://quirkyagents.com/wizard/blogassets/mcp-one-tool-call-tall.72e4f933.png" alt="One MCP tool call in five steps: connect, discover with tools/list, the model chooses a tool, call with tools/call, use the result" width="1200" height="1635" loading="lazy" decoding="async"></p>
<ol>
<li><strong>Connect, once.</strong> <code>claude mcp add --transport http upwork https://mcp.upwork.com/mcp</code>. Upwork&#39;s server needs an account, so the first request is refused with &quot;sign in first&quot;, the client finds Upwork&#39;s login server, and I sign in with OAuth. The token works only for that server. A server running locally on your own machine usually needs no sign-in.</li>
<li><strong>Discover.</strong> The client asks <code>tools/list</code>. The server replies with each tool&#39;s name, a plain-English description and the shape of its inputs (a JSON Schema).</li>
<li><strong>Choose.</strong> I say &quot;find LLM eval jobs&quot;. The model reads the descriptions, picks the search tool and fills in the inputs.</li>
<li><strong>Call.</strong> The client sends <code>tools/call</code> with the tool name and arguments. Upwork runs the search.</li>
<li><strong>Use the result.</strong> The result goes into the model&#39;s context, and the model writes me an answer.</li>
</ol>
<p>The messages are JSON-RPC: a method name plus parameters.</p>
<figure class="code"><figcaption>json</figcaption><pre><code class="hljs language-json"><span class="hljs-punctuation">{</span>
  <span class="hljs-attr">&quot;jsonrpc&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;2.0&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;id&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-number">7</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;method&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;tools/call&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;params&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span>
    <span class="hljs-attr">&quot;name&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;find_jobs&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;arguments&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span> <span class="hljs-attr">&quot;query&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;LLM evaluation&quot;</span> <span class="hljs-punctuation">}</span>
  <span class="hljs-punctuation">}</span>
<span class="hljs-punctuation">}</span></code></pre></figure>
<p>(Argument names simplified.)</p>
<h2 id="1-the-description-is-the-whole-interface">1. The description is the whole interface</h2>
<p>Step 3 surprised me most. In a normal agent loop, no code decides which tool gets called. The model decides, by reading names and descriptions. <strong>It never sees your code.</strong></p>
<p>So write each description the way you&#39;d brief a new colleague: what the tool does, when to use it, what comes back. Tool picking gets harder as the list grows and descriptions overlap. How many is too many depends on the model and the host. <a href="https://www.pagerduty.com/eng/lessons-learned-while-building-pagerduty-mcp-server/">PagerDuty&#39;s engineers</a> call 20 to 25 the sweet spot, and their own server ships over 20. <a href="https://quirkyagents.com/wizard/projects/doceval/">doceval</a>&#39;s needs two. One scores a single extraction against the expected values, the other runs a full eval over a labelled dataset.</p>
<p>Keep results short too. A result goes into the model&#39;s context and usually stays there for the rest of the conversation. A tool that returns 50 pages makes the agent slower, dearer and more likely to lose the thread.</p>
<h2 id="2-a-local-server-lives-and-dies-with-your-session">2. A local server lives and dies with your session</h2>
<p>Upwork&#39;s server runs on Upwork&#39;s machines. A local server is different: a program on your own computer that the app starts for you. I wanted to know exactly when it starts and stops, so I tested it with doceval&#39;s server in Claude Code: a test config passed to <code>claude -p</code>, and the process list checked every quarter of a second.</p>
<ul>
<li><strong>It starts once, when the session starts.</strong> My prompt used no tools, and the server still started, exactly once.</li>
<li><strong>It is the app&#39;s child process.</strong> The server&#39;s parent process was Claude.</li>
<li><strong>It is reused for every call.</strong> In a session that called doceval&#39;s tool twice, the server started once.</li>
<li><strong>It stops when the session ends.</strong> Claude shut it down before exiting itself.</li>
<li><strong>It stops after a crash too.</strong> I killed Claude with <code>kill -9</code>, which gives it no chance to clean up. The server was gone within two seconds.</li>
</ul>
<p>The link between them is a pair of pipes: the app writes requests into the server&#39;s input and reads replies from its output. When I closed a running server&#39;s input by hand, it exited on its own. That&#39;s how it learns the session is over, whether the app quit cleanly or crashed.</p>
<p>Between calls the server just waits, and each session starts its own copy. On a small machine, three open sessions means three copies of every local server in memory.</p>
<h2 id="3-your-mcp-config-can-hold-secrets-in-plain-text">3. Your MCP config can hold secrets in plain text</h2>
<p>A local server often needs a password or an API key, and the usual place for it is the <code>env</code> block in the app&#39;s MCP config. That&#39;s a plain-text file. Anything running as your user can read it, including a coding agent with file access, and a project&#39;s config file can end up in git.</p>
<p>Two things that help:</p>
<ul>
<li><strong>Reference the secret instead of writing it.</strong> Claude Code fills in <code>${VAR}</code> from your environment. I tested it with a config file passed to <code>claude --mcp-config</code>: a config with <code>&quot;SECRET&quot;: &quot;${MY_TEST_SECRET}&quot;</code> handed the server the value from my shell. The config can then be shared, and the secret lives in one place outside it.</li>
<li><strong>Prefer servers that sign you in.</strong> Remote servers that use OAuth keep no secret in the config. My entries for Upwork, Notion and Todoist hold a type and a URL, and nothing else.</li>
</ul>
<p>Stronger options exist, like the operating system&#39;s keychain or a secret manager that starts the server for you. I haven&#39;t tested those yet, so I&#39;ll stop at naming them.</p>
<h2 id="4-what-a-tool-returns-is-untrusted-so-approve-the-actions">4. What a tool returns is untrusted, so approve the actions</h2>
<p>A job post, an email or a web page can contain text written to steer the model. That&#39;s prompt injection, and whatever a tool returns goes into the model&#39;s context.</p>
<p>So I approve every Upwork proposal myself. Upwork&#39;s MCP has a preview built in: the agent shows me the proposal, I confirm, and the preview expires after 15 minutes. Claude Code also asks before running a tool, unless you&#39;ve allowed it.</p>
<p>I got this wrong at first. I shipped upsweep with a rule that said &quot;drafts only&quot;, because I read Upwork&#39;s terms as forbidding agent submission. They allow submitting, one approval per proposal, never scripted around. I corrected it in public: <a href="https://quirkyagents.com/wizard/blog/can-ai-agent-submit-upwork-proposals/">Can an AI agent submit Upwork proposals?</a></p>
<h2 id="checklist">Checklist</h2>
<ul>
<li>Describe each tool like a brief to a colleague, with no two that overlap.</li>
<li>Return the smallest result that answers the question.</li>
<li>Expect one copy of every local server per open session.</li>
<li>Keep secrets out of the config file: reference them with <code>${VAR}</code>, or use servers that sign you in.</li>
<li>Treat what a tool returns as untrusted, and approve anything that changes something.</li>
</ul>
<p>The model never sees your code, only your tool names, descriptions and schemas. Write them like they are the code.</p>
<p>For what happens underneath, local and remote, step by step: <a href="https://quirkyagents.com/wizard/blog/local-vs-remote-mcp-servers/">Local vs remote MCP servers: how each connects</a>.</p>
<h2 id="the-tools-in-this-article">The tools in this article</h2>
<ul>
<li><a href="https://quirkyagents.com/wizard/projects/doceval/">doceval</a>: measures extraction accuracy field by field, with an MCP server for coding agents. <code>pip install &quot;doceval[mcp]&quot;</code></li>
<li><a href="https://quirkyagents.com/wizard/projects/upsweep/">upsweep</a>: finds Upwork jobs that match your filters, through Upwork&#39;s official MCP. <code>npx skills add dave8172/upsweep</code></li>
</ul>
<p>Both are free and MIT licensed.</p>
]]></content:encoded></item><item><title>Local vs remote MCP servers: how each connects, step by step</title><link>https://quirkyagents.com/wizard/blog/local-vs-remote-mcp-servers/</link><guid>https://quirkyagents.com/wizard/blog/local-vs-remote-mcp-servers/</guid><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><description>How a local MCP server lives and dies with your session, and how remote MCP sign-in works: OAuth, PKCE and tokens, tested on Upwork, Notion and Todoist.</description><category>MCP</category><category>Agents</category><category>OAuth</category><category>Claude Code</category><content:encoded><![CDATA[<p>An MCP server reaches your AI app in one of two ways. A <strong>local</strong> server is a program on your own machine. A <strong>remote</strong> server lives at a URL and signs you in with OAuth. I wanted to know exactly what happens in each case, so I tested it: a local server (<a href="https://quirkyagents.com/wizard/projects/doceval/">doceval</a>&#39;s) started by Claude Code, and the three remote servers I use (Upwork, Notion and Todoist). Where something comes from the MCP spec instead of a test, I link the spec.</p>
<p>For what MCP is and how one tool call works, see my <a href="https://quirkyagents.com/wizard/blog/building-mcp-servers-lessons/">previous post</a>.</p>
<hr>
<h2 id="local-a-child-process-on-two-pipes">Local: a child process on two pipes</h2>
<p>The app starts the server as a program on your machine and talks to it through two pipes: it writes requests into the server&#39;s input and reads replies from its output. I ran doceval&#39;s server under Claude Code (<code>claude -p</code> with a test config) and checked the process list every quarter of a second:</p>
<p><img src="https://quirkyagents.com/wizard/blogassets/mcp-local-life.134e1804.png" alt="A local MCP server's life: started once when the session starts, reused for every call, stopped when the session ends, and gone within two seconds after a crash" width="1200" height="1331" loading="lazy" decoding="async"></p>
<ol>
<li><strong>Session starts.</strong> Claude Code starts the server as its own child process, once, even when the prompt uses no tools.</li>
<li><strong>Every call reuses it.</strong> A session that called the tool twice started the server once.</li>
<li><strong>Session ends.</strong> Claude stops the server before exiting itself.</li>
<li><strong>Crash.</strong> After <code>kill -9</code> on Claude, which allows no clean-up, the server was gone within two seconds. Its input pipe closed, and it read that as &quot;the session is over&quot;.</li>
</ol>
<p>Each session starts its own copy, so three open sessions run three copies of every local server.</p>
<p><strong>Credentials come from the environment.</strong> The <a href="https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization">MCP spec</a> says local (stdio) servers should take credentials from the environment instead of using OAuth. In practice that&#39;s the <code>env</code> block of the app&#39;s MCP config, a plain-text file. Claude Code fills in <code>${VAR}</code> from your shell: a config with <code>&quot;SECRET&quot;: &quot;${MY_TEST_SECRET}&quot;</code> handed the server my shell&#39;s value, so the secret doesn&#39;t have to sit in the file.</p>
<h2 id="remote-five-parties">Remote: five parties</h2>
<table>
<thead>
<tr>
<th>Who</th>
<th>In my case</th>
</tr>
</thead>
<tbody><tr>
<td><strong>You</strong></td>
<td>the person signing in</td>
</tr>
<tr>
<td><strong>Browser</strong></td>
<td>where you type your password</td>
</tr>
<tr>
<td><strong>Client</strong></td>
<td>Claude Code</td>
</tr>
<tr>
<td><strong>MCP server</strong></td>
<td><code>mcp.upwork.com/mcp</code>: runs the tools, checks tokens</td>
</tr>
<tr>
<td><strong>Sign-in server</strong></td>
<td>issues tokens. The spec calls it the <em>authorization server</em></td>
</tr>
</tbody></table>
<p>The MCP server and the sign-in server can belong to the same company and still be separate roles. Todoist&#39;s MCP server lives at <code>ai.todoist.net</code>, and its sign-in server at <code>todoist.com</code>.</p>
<h2 id="remote-the-first-connection-step-by-step">Remote: the first connection, step by step</h2>
<p><img src="https://quirkyagents.com/wizard/blogassets/mcp-remote-signin.42784831.png" alt="The remote MCP sign-in flow in nine steps across Claude Code, the MCP server, the sign-in server and your browser" width="1200" height="2519" loading="lazy" decoding="async"></p>
<p><strong>1. Claude calls the server with no token, and gets a 401.</strong> I sent all three servers a request with no token. Each one answered <code>401 Unauthorized</code>, with a <code>www-authenticate</code> header pointing to a public file. Upwork&#39;s points to <a href="https://mcp.upwork.com/.well-known/oauth-protected-resource/mcp">mcp.upwork.com/.well-known/oauth-protected-resource/mcp</a>.</p>
<p><strong>2. That file names the sign-in server.</strong> For Upwork it&#39;s <code>https://mcp.upwork.com</code>, and for Todoist it&#39;s <code>https://todoist.com</code>.</p>
<p><strong>3. The sign-in server&#39;s own public file lists its addresses.</strong> These are where you log in and where codes get swapped for tokens. Upwork&#39;s login page and token address are on <code>www.upwork.com</code>. All three list PKCE with <code>S256</code> (more in step 5).</p>
<p><strong>4. Claude identifies itself with a URL.</strong> In my stored credentials, the app ID Claude used with all three servers is the same: <code>https://claude.ai/oauth/claude-code-client-metadata</code>. That&#39;s a public description of the Claude Code app, which the sign-in server fetches. The spec prefers this method and marks the older automatic registration as deprecated. The description says <code>&quot;token_endpoint_auth_method&quot;: &quot;none&quot;</code>: Claude Code has no app secret, and every install shares one public ID.</p>
<p><strong>5. Claude makes a one-time secret: PKCE.</strong> With no app secret, something else has to prove that the app finishing the login is the one that started it. Claude makes a random <em>verifier</em>, keeps it, and puts only its hash (the <em>challenge</em>) into the login link.</p>
<p><strong>6. You log in on the service&#39;s own page.</strong> Claude opens your browser at the sign-in server&#39;s login page. Per the spec, the link carries the app ID, the return address (Claude&#39;s description lists <code>http://localhost/callback</code>), the challenge, and which server the token is for. <strong class="spell">Your password goes only to the service.</strong> Claude never sees it.</p>
<p><strong>7. A one-time code comes back through your browser.</strong> After you approve, the sign-in server sends your browser to Claude&#39;s local return address with a short-lived, single-use code. The spec also requires the client to check that the response came from the sign-in server it expected.</p>
<p><strong>8. Claude swaps the code for tokens.</strong> It sends the code plus the verifier from step 5. The sign-in server hashes the verifier and compares it with the challenge from step 6. Only a match gets tokens, so a stolen code is useless without the verifier. Back come an <strong>access token</strong> (short-lived, used on every request) and a <strong>refresh token</strong> (used only to get new access tokens).</p>
<p><strong>9. Claude stores both.</strong> On my machine they&#39;re in <code>~/.claude/.credentials.json</code>, an owner-only file, with fields for each server: <code>accessToken</code>, <code>refreshToken</code>, <code>expiresAt</code>, <code>clientId</code>, <code>issuer</code>. My MCP config file holds only each server&#39;s type and URL. The tokens are still on disk, readable by anything running as me, but they expire, work for one server, can be revoked, and never contained my password.</p>
<h2 id="remote-every-request-after-that">Remote: every request after that</h2>
<p><strong>The token travels on every request,</strong> as <code>Authorization: Bearer &lt;access token&gt;</code>. The <a href="https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization">spec</a> requires it on every HTTP request, requires the server to check that the token was issued for it, and forbids passing tokens on to anyone else.</p>
<p><strong>How a server checks a token depends on the kind.</strong> There are two:</p>
<ul>
<li><strong>Opaque:</strong> a random string that means nothing by itself. The server looks it up in its own database, or asks the sign-in server (a standard call named <em>introspection</em>).</li>
<li><strong>Signed (a JWT):</strong> the token carries its own contents and a signature, and the server checks it with a public key.</li>
</ul>
<p>All three of my access tokens are opaque to me: single strings, not the three dot-separated parts of a JWT, and none of the three sign-in servers publishes a list of public keys. How each service checks its tokens happens on its side, out of view. Notion&#39;s sign-in server does publish an introspection address, the standard way to look a token up.</p>
<h2 id="how-a-signed-token-is-checked-tested-locally">How a signed token is checked (tested locally)</h2>
<p>None of my three servers uses JWTs, so I built one on my machine to see how the checking works. A JWT is three parts joined by dots: <code>header.contents.signature</code>.</p>
<p><img src="https://quirkyagents.com/wizard/blogassets/mcp-jwt-check.b3c52e2b.png" alt="How a JWT is checked: the server hashes the header and contents, opens the signature with the public key, and compares; changed contents or a different key fail" width="1200" height="1400" loading="lazy" decoding="async"></p>
<ul>
<li><strong>Signing (by the sign-in server):</strong> hash <code>header.contents</code> (anyone can compute a hash; no key needed), then transform the hash with the <strong>private</strong> key. The result is the signature.</li>
<li><strong>Checking (by the MCP server):</strong> open the signature with the <strong>public</strong> key and compare it with your own hash of the text you received. If they match, the token is genuine and unchanged.</li>
<li><strong>My results:</strong> the real token passed. Changing the user ID in the contents failed. Re-signing it with a different key failed too.</li>
<li><strong>Which public key:</strong> the header names the key (<code>kid</code>), and the sign-in server publishes its public keys at a fixed address. Several keys can be published at once, so an old key keeps working while tokens signed with it expire. A careful MCP server fetches keys only from the sign-in server it already trusts, never from an address written inside the token.</li>
<li><strong>One key, many tokens:</strong> a key pair usually signs tokens for every user for a long time, until it&#39;s rotated. What&#39;s new on each login is the token.</li>
</ul>
<p><strong>A JWT is readable by anyone who holds it.</strong> The contents are only encoded, so signing stops tampering without hiding anything. Secrecy comes from HTTPS and from where the token is stored.</p>
<h2 id="refresh-revoke-and-no-sessions">Refresh, revoke, and no sessions</h2>
<ul>
<li><strong>Refresh:</strong> when the access token expires, Claude sends the refresh token to the sign-in server and gets a new access token, with no login.</li>
<li><strong>Revoke:</strong> all three sign-in servers publish a revocation address, and a service can also cancel tokens on its own side. After that the next refresh fails, and you&#39;re back at step 1.</li>
<li><strong>No sessions:</strong> the <a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/">2026-07-28 spec</a> retired the <code>initialize</code> handshake and the <code>Mcp-Session-Id</code> header. Each request now carries its own protocol version and client details, so any copy of a server can answer it. The state lives at the edges: tokens on your disk, and token records at the sign-in server.</li>
</ul>
<h2 id="what-the-tokens-can-39-t-protect-against">What the tokens can&#39;t protect against</h2>
<p>If an attacker steals a sign-in server&#39;s private signing key, they can make tokens the MCP server will accept. That happened in 2023, when a group Microsoft calls Storm-0558 <a href="https://www.microsoft.com/en-us/msrc/blog/2023/07/microsoft-mitigates-china-based-threat-actor-storm-0558-targeting-of-customer-email">forged tokens with a stolen Microsoft signing key</a> to read email. OAuth limits the damage of a leak, with short-lived tokens, tokens bound to one server, revocation and key rotation, but a breach of the service itself is beyond what any token can fix.</p>
<h2 id="who-does-what">Who does what</h2>
<table>
<thead>
<tr>
<th>Piece</th>
<th>Made by</th>
<th>Kept by</th>
</tr>
</thead>
<tbody><tr>
<td>App ID (a URL)</td>
<td>the app&#39;s maker</td>
<td>public</td>
</tr>
<tr>
<td>PKCE verifier</td>
<td>Claude</td>
<td>Claude, until step 8</td>
</tr>
<tr>
<td>One-time code</td>
<td>sign-in server</td>
<td>passes through your browser, used once</td>
</tr>
<tr>
<td>Access token</td>
<td>sign-in server</td>
<td>Claude, sent on every request</td>
</tr>
<tr>
<td>Refresh token</td>
<td>sign-in server</td>
<td>Claude, plus a record at the sign-in server</td>
</tr>
<tr>
<td>Your password</td>
<td>you</td>
<td>the sign-in server only</td>
</tr>
<tr>
<td>Local server secrets</td>
<td>you</td>
<td>the server&#39;s environment</td>
</tr>
</tbody></table>
]]></content:encoded></item><item><title>How I calibrated an LLM judge to grade like me, 25× cheaper</title><link>https://quirkyagents.com/wizard/blog/calibrate-llm-judge-evals/</link><guid>https://quirkyagents.com/wizard/blog/calibrate-llm-judge-evals/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><description>Evals on 188 answers from ChatGPT, Claude, Chatbase and my own pipeline over product datasheets, and a hybrid judge with Jev that grades like I do.</description><category>AI</category><category>Evals</category><category>LLM-as-judge</category><category>Benchmarking</category><category>TypeSafe</category><category>Jev</category><content:encoded><![CDATA[<p>Businesses that sell technical products answer the same kind of question every
day. <em>&quot;What&#39;s the accuracy on this range?&quot;</em> <em>&quot;Can it measure through a coating, and
how thick?&quot;</em> <em>&quot;Does it come with a calibration certificate?&quot;</em> The answers sit in
product datasheets and manuals.</p>
<p>I took 27 of those PDFs, three of them scans with no text layer, and wrote 47
questions the way customers actually ask them. Then I had four AI systems answer
every question and graded all 188 answers.</p>
<p>Running the systems took an afternoon. Building a judge I could trust took the
rest of the work, and that is what this post is about.</p>
<h2 id="the-result">The result</h2>
<p><img src="https://quirkyagents.com/wizard/blogassets/llm-judge-results.3c6cb3a8.png" alt="Correct answers out of 47: my pipeline 46, Claude app 45, ChatGPT app 35, Chatbase 33" width="800" height="680" loading="lazy" decoding="async"></p>
<table>
<thead>
<tr>
<th>System</th>
<th>Correct</th>
<th>Wrong</th>
<th>Made up</th>
</tr>
</thead>
<tbody><tr>
<td>My pipeline</td>
<td><strong>46</strong></td>
<td>1</td>
<td>0</td>
</tr>
<tr>
<td>Claude app</td>
<td><strong>45</strong></td>
<td>1</td>
<td>0</td>
</tr>
<tr>
<td>ChatGPT app</td>
<td>35</td>
<td>10</td>
<td>0</td>
</tr>
<tr>
<td>Chatbase</td>
<td>33</td>
<td>11</td>
<td>0</td>
</tr>
</tbody></table>
<p>My pipeline is Claude Sonnet over the API with page citations. The Claude app ran
Opus with the documents in a project, ChatGPT ran with thinking off, and Chatbase
was on the free plan with its default model.</p>
<p>Every system got the same PDFs and the same instructions. One run each, so I read
a one-question gap as a tie.</p>
<p><strong>No system invented a spec value.</strong> I checked every claim twice: a claim checker
read each one against the page it came from, and I went through the ones it was
unsure about by hand. The weaker systems failed more quietly. They said an answer
wasn&#39;t in the documents when it was.</p>
<h2 id="where-the-misses-came-from">Where the misses came from</h2>
<p><img src="https://quirkyagents.com/wizard/blogassets/llm-judge-by-type.882a241c.png" alt="Correct answers by question type for each system" width="800" height="730" loading="lazy" decoding="async"></p>
<p><strong>Scanned pages.</strong> Systems that only read the text layer saw blank pages. Every
question answered only by a scan came back &quot;not specified&quot;.</p>
<p><strong>Answers that need a second look.</strong> A product advertises one range on the front
page. A different mode of the same product stops at a tenth of it, and the
customer&#39;s question is about that mode. The same pattern showed up in accuracy
tables split by range and in specs that change with the material.</p>
<p><strong>Unit conversions.</strong> The customer asks in one unit. The datasheet lists the
other unit in one row and a separate, lower limit in the customer&#39;s unit in
another. Converting the first number gives a confident wrong answer.</p>
<p><strong>The website and the datasheet disagreeing.</strong> The product page said one value
and the datasheet said half of it. My rule: the datasheet wins, and the reply says
the page may be wrong. Two systems missed it.</p>
<h2 id="rules-that-live-in-no-pdf">Rules that live in no PDF</h2>
<p>Before running anything I reviewed every question. I dropped three that no
customer would ask, corrected two gold answers, and turned the &quot;not in the
documents&quot; cases into rules:</p>
<ul>
<li><strong>Datasheet beats product page.</strong> Give the datasheet value and say the page may
have an error.</li>
<li><strong>Never a bare &quot;not in our documents&quot;.</strong> Say it isn&#39;t in the published
documents, and give an email to ask.</li>
<li><strong>Certifications only if the document says so.</strong> Otherwise, email for
confirmation.</li>
<li><strong>Canned answers</strong> for the two most common questions.</li>
</ul>
<p>After grading I added one more: <strong>answer what was asked.</strong> Extra conditions
confuse buyers and invite more questions.</p>
<p>That short list did more for answer quality than any prompt trick.</p>
<h2 id="why-the-judge-needs-a-judge">Why the judge needs a judge</h2>
<p>Grading 188 answers by hand takes hours, so the usual move is an LLM judge: give
a model the question, the gold answer and the answer, and ask for correct, partial
or wrong. You only know the judge is any good if it agrees with the person whose
standard matters. Here, that&#39;s me.</p>
<p>So I graded 20 answers blind: no system names, five per system, with some likely
mistakes mixed in so the judge would be tested on errors too. Then I compared.</p>
<p><strong>First try: 13 of 20.</strong> Most of the gap was in my own answer key. Three gold
answers were wrong. Grading real answers showed me that I answer from the
datasheet <em>as printed</em>, even when a conversion on it looks off, and flag the
document separately. My key hadn&#39;t been written that way. Once I fixed it,
agreement jumped.</p>
<p>Plain agreement flatters a judge when most answers are correct. A judge that
marks everything &quot;correct&quot; would agree 80% of the time here and understand
nothing. <strong>Cohen&#39;s kappa</strong> corrects for that.</p>
<p><img src="https://quirkyagents.com/wizard/blogassets/llm-judge-kappa.d0dc379e.png" alt="Cohen's kappa: 67 of 100 agreement expected by luck, 28 from skill, 5 missed; kappa = 28 / 33 = 0.85" width="800" height="680" loading="lazy" decoding="async"></p>
<p>Kappa asks how far above luck the judge got, as a share of how far above luck it
could have got. 0 means no better than random; 1 means it matched me every time.</p>
<h2 id="two-judges">Two judges</h2>
<p><img src="https://quirkyagents.com/wizard/blogassets/llm-judge-judges.a785a712.png" alt="Agreement with my grades: Sonnet 35 of 40, hybrid 36 of 40. Cost for 188 answers: $0.80 vs $0.03" width="800" height="780" loading="lazy" decoding="async"></p>
<table>
<thead>
<tr>
<th>Judge</th>
<th>Agrees</th>
<th>Cost, 188</th>
<th>Speed</th>
</tr>
</thead>
<tbody><tr>
<td>Sonnet alone</td>
<td>35 / 40</td>
<td>~$0.80</td>
<td>seconds</td>
</tr>
<tr>
<td><strong>Hybrid</strong></td>
<td><strong>36 / 40</strong></td>
<td><strong>~$0.03</strong></td>
<td>under 1 s</td>
</tr>
</tbody></table>
<p>The hybrid gives each part the job it does best.</p>
<p><img src="https://quirkyagents.com/wizard/blogassets/llm-judge-hybrid-judge.d55153cf.png" alt="The hybrid judge: code compares numbers, Jev answers typed questions, code applies the policy; Sonnet with the PDF only for unsure claims" width="800" height="860" loading="lazy" decoding="async"></p>
<p><strong>Code</strong> pulls the numbers out of the gold answer and the answer being graded,
and records which match. It never confuses ±(1.2% + 5) with ±(0.8% + 5).</p>
<p><strong><a href="https://typesafe.ai">Jev</a></strong> handles meaning. It&#39;s a System One model from
TypeSafe: it reads natural language and returns typed answers with
probabilities, with no prose to parse. One request per answer asks four
questions at once:</p>
<figure class="code"><figcaption>json</figcaption><pre><code class="hljs language-json"><span class="hljs-punctuation">{</span>
  <span class="hljs-attr">&quot;verdict&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span><span class="hljs-attr">&quot;type&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;choice&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;instructions&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;Grade `answer` to `customer_question` against `gold` and `notes`. Use `code_checks` for which gold numbers the answer contains.&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;criteria&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span><span class="hljs-attr">&quot;correct&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;...&quot;</span><span class="hljs-punctuation">,</span> <span class="hljs-attr">&quot;partial&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;...&quot;</span><span class="hljs-punctuation">,</span> <span class="hljs-attr">&quot;wrong&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;...&quot;</span><span class="hljs-punctuation">}</span><span class="hljs-punctuation">}</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;brush_off&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span><span class="hljs-attr">&quot;type&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;noul&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;instructions&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;Does `answer` brush the customer off, such as a bare &#x27;not in our documents&#x27; with no next step?&quot;</span><span class="hljs-punctuation">}</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;unasked_extra&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span><span class="hljs-attr">&quot;type&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;noul&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;instructions&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;Does `answer` add specs, conditions or other models that `customer_question` didn&#x27;t ask about?&quot;</span><span class="hljs-punctuation">}</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;rule_followed&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span><span class="hljs-attr">&quot;type&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;noul&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;instructions&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;Does `answer` do what `rule` says?&quot;</span><span class="hljs-punctuation">}</span>
<span class="hljs-punctuation">}</span></code></pre></figure>
<p>A Choice picks one of the grades and returns a probability for each. A Noul is
the probability that a statement holds. The judgment comes back as numbers, so
the policy stays in code where I can see it and change it. One real answer from
the run:</p>
<figure class="code"><figcaption>text</figcaption><pre><code class="hljs">verdict        correct 0.98 · partial 0.02 · wrong 0.00
unasked_extra  0.91  → at or above 0.9: &quot;too much information&quot; flag</code></pre></figure>
<p>About 2,000 tokens per answer, under a second, and all 188 answers for about three
cents.</p>
<p><strong>Sonnet with the original PDF</strong> only sees the few claims Jev isn&#39;t sure about.</p>
<h2 id="keeping-myself-honest">Keeping myself honest</h2>
<p>I wrote the pass line down before I graded a second set of 20: <strong>at least 16 of
20, kappa at least 0.6.</strong> That second set got used once, as a test, and never for
tuning.</p>
<p>It caught a mistake. On the first 20, I had tuned a penalty: if Jev was 90% sure
an answer added unasked details, the grade dropped to partial. It looked perfect
there. On the fresh set it caused two of the three disagreements. The judge
passed the line (17 of 20, kappa 0.63), but barely.</p>
<p>Reading those two answers showed why. I grade substance and length separately.
An answer can be right and too long. The detector itself was right: it flagged
exactly the four answers across both sets that I had found too long, and no
others. The penalty was the wrong policy. So it became a flag beside the grade.
Without it the judge agrees with me 19 of 20, kappa 0.85. That change came after
I&#39;d seen the second set, so the next fresh set is where it has to hold.</p>
<h2 id="the-loop-i-39-d-run-every-time">The loop I&#39;d run every time</h2>
<p><img src="https://quirkyagents.com/wizard/blogassets/llm-judge-eval-loop.d3c5ffd8.png" alt="The eval loop: golden set, all systems answer, I grade 20 blind, the judge grades the same 20 and I fix each disagreement, a fresh 20 against a line set in advance, then grade everything" width="800" height="850" loading="lazy" decoding="async"></p>
<p>I did it in a worse order: ran the systems, graded everything with an untested
judge, reported numbers, then calibrated. Next time:</p>
<ol>
<li>Build the golden set: questions as asked, reviewed gold answers.</li>
<li>All systems answer, same documents, same instructions.</li>
<li>I grade 20 answers blind.</li>
<li>The judge grades the same 20. I read every disagreement and fix the cause,
which is usually the answer key.</li>
<li>Write the pass line down, then test on a fresh 20. Use it once.</li>
<li>Only then grade everything and report.</li>
</ol>
<p>A golden set isn&#39;t finished until a person has checked real answers against it.</p>
<h2 id="what-it-means">What it means</h2>
<p>On product datasheets, a frontier model with the PDFs attached is already very
good: my pipeline and the Claude app tied. Accuracy is the starting point. The
hard part is getting that accuracy into the place where buyers ask, across a
catalogue too big for one prompt, while datasheets change and the website drifts
away from them. Then proving it on the business&#39;s own questions, with a judge
calibrated to the person who answers them. That is what I&#39;m building now.</p>
<p>The whole run cost about $8 in API calls. Most of that went on mistakes I won&#39;t
repeat. Now I estimate every paid run from a one-question trial, counting every
path that can call a paid model.</p>
]]></content:encoded></item><item><title>Can an AI agent submit Upwork proposals? Yes, with your OK</title><link>https://quirkyagents.com/wizard/blog/can-ai-agent-submit-upwork-proposals/</link><guid>https://quirkyagents.com/wizard/blog/can-ai-agent-submit-upwork-proposals/</guid><pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate><description>Upwork's official MCP lets Claude, Codex or Cursor send a proposal for you. What it shows before sending, what stays on upwork.com, and what I got wrong.</description><category>Upwork</category><category>MCP</category><category>Agents</category><category>Claude Code</category><content:encoded><![CDATA[<p><strong>Short answer: yes.</strong> Upwork&#39;s official MCP, launched in August 2026, lets an AI agent submit a proposal on your behalf. It sends nothing until you approve a preview of that exact proposal.</p>
<p>I tested it today from Claude Code, on a real job, with my own account.</p>
<hr>
<h2 id="what-happens-when-you-say-send">What happens when you say &quot;send&quot;</h2>
<ol>
<li><strong>The agent checks for an invite or an earlier proposal</strong> on that job. If the client already invited you, it accepts the invite instead of applying again.</li>
<li><strong>It builds a preview.</strong> You see the cover letter, your rate, the Connects it will cost, what you will have left, and the portfolio items attached.</li>
<li><strong>It asks about extras:</strong> files to attach, and whether to boost. A boost is a Connects bid for a top slot, and the preview shows the real competing bids, so you know the smallest amount that gets in.</li>
<li><strong>You approve, or change something.</strong> The preview expires after 15 minutes, so a forgotten one never goes out later.</li>
<li><strong>It submits.</strong> The proposal shows as Pending on Upwork, like any other.</li>
</ol>
<p>One approval covers one proposal. There is no &quot;apply to all ten&quot;.</p>
<h2 id="what-stays-on-upwork-com">What stays on upwork.com</h2>
<p>Per <a href="https://support.upwork.com/hc/en-us/articles/55446516654611-How-to-use-Upwork-with-AI-agents-through-MCP">Upwork&#39;s own help page</a>, offers, milestone payments and accepting a contract complete on upwork.com, outside the agent. The agent can bring you an offer; you take it on the site.</p>
<h2 id="the-rules-i-follow">The rules I follow</h2>
<ul>
<li><strong>Approve every proposal yourself,</strong> after reading the preview. Upwork&#39;s <a href="https://www.upwork.com/legal#apimcpterms">API &amp; MCP Terms</a> forbid scripting around its confirmation step.</li>
<li><strong>Activity on your login is yours.</strong> If your agent sends something wrong, it went from your account.</li>
<li><strong>Sweep on demand.</strong> Searching on a schedule counts as monitoring, which the Terms rule out.</li>
</ul>
<h2 id="what-i-got-wrong">What I got wrong</h2>
<p>I shipped <strong class="spell">upsweep</strong>, my open-source skill for finding Upwork jobs with an agent, with a rule that said <em>drafts only, never submit</em>. I had read the Terms as forbidding agent submission. They don&#39;t, and Upwork&#39;s own tooling is built around a submit-with-confirmation flow. The rule now reads: submit only after the user approves the preview, one proposal at a time. The fix is a <a href="https://github.com/dave8172/upsweep/commit/ee6785f">public commit</a>.</p>
<h2 id="try-it">Try it</h2>
<figure class="code"><figcaption>text</figcaption><pre><code class="hljs">npx skills add dave8172/upsweep</code></pre></figure>
<p>Say &quot;upsweep&quot; for jobs that pass your filters, &quot;draft 2&quot; for a proposal draft, and &quot;send&quot; when you want the preview. Details on the <a href="https://quirkyagents.com/wizard/projects/upsweep/">project page</a>.</p>
<p>upsweep is unofficial and not affiliated with Upwork.</p>
]]></content:encoded></item><item><title>Ask twice: seven measurements from building with Jev</title><link>https://quirkyagents.com/wizard/blog/jev-ask-twice/</link><guid>https://quirkyagents.com/wizard/blog/jev-ask-twice/</guid><pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate><description>A coordinate lookup scored 3/24 where handing the model the naming job scored 22/24. Then two claims I had argued rather than measured turned out to be wrong.</description><category>AI</category><category>Reliability</category><category>Calibration</category><category>TypeSafe</category><category>Jev</category><content:encoded><![CDATA[<p>Jev is a System One model from TypeSafe. It reads natural language like any LLM,
and instead of writing a reply it returns a probability distribution over options
you define. No prose, no reasoning trace, no JSON to repair.</p>
<p>I spent a day building <strong class="spell">waif</strong>, which reads a piece of
text and names the feeling in it. That job has no right answer, which makes it an
unusually honest test rig: nothing can be graded against a key, so every design
decision has to be either argued or measured.</p>
<p>I measured five, argued two, and the two I argued were both wrong. Sections 1
and 7 are those, kept in place with the measurement underneath rather than
quietly edited out.</p>
<hr>
<h2 id="1-a-choice-splits-its-own-vote-between-synonyms-but-less-than-i-claimed">1. A Choice splits its own vote between synonyms — but less than I claimed</h2>
<p>Here is the argument I built on, and I am leaving it in its original words
because the correction under it is the useful part.</p>
<p><em>Annoyed</em>, <em>irritated</em> and <em>frustrated</em> are one feeling in three wordings. A
Choice divides probability between them, so a text the model read perfectly
clearly comes back split three ways, and confidence collapses for a reason that
has nothing to do with the input. The distribution is telling you about your
option list, not about the text. Therefore: never put sixty near-synonyms in one
Choice.</p>
<p><strong>Then I measured it, and the effect is real but far smaller than the argument
needs.</strong> One Choice over all 62 words, same glosses, 23 texts:</p>
<table>
<thead>
<tr>
<th></th>
<th>Mean P(chosen word)</th>
<th>Score vs acceptable words</th>
</tr>
</thead>
<tbody><tr>
<td>Two Choices — family, then shade within it</td>
<td>0.689</td>
<td>43 / 46</td>
</tr>
<tr>
<td><strong>One Choice over all 62 words</strong></td>
<td><strong>0.812</strong></td>
<td><strong>44 / 46</strong></td>
</tr>
</tbody></table>
<p>The one-Choice version is <em>more</em> certain, not less. It picks the same word 21
times out of 23. On a plainly frustrated text it returns <em>frustration</em> at 0.98.
Stripping the glosses and asking with 62 bare words barely moves it: 0.79.</p>
<p>So where does the splitting go? It shows up exactly where you would want it to:</p>
<figure class="code"><figcaption>text</figcaption><pre><code class="hljs">shipped   pride 0.38  →  excitement 0.36
scan      worry 0.54  →  anxiety    0.39
sentit    guilt 0.60  →  embarrassment 0.26</code></pre></figure>
<p>Every one of those lost ground to a near-synonym — that <em>is</em> vote splitting. But
the two-Choice design flattens on <strong>the same four texts</strong> (0.33, 0.53, 0.35,
0.32). The splitting is not caused by the option list being long. It is caused
by the text genuinely sitting between two words, and both designs report it
because both are calibrated. <em>&quot;I snapped at her in front of the kids&quot;</em> really is
guilt and regret at once.</p>
<p><strong>A long option list does not flatten a clear text.</strong> That is the part I had
backwards, and the rule has to be narrower than I wrote it:</p>
<blockquote>
<p>A Choice splits its vote when two options are <strong>the same answer in different
wordings</strong>. The fix is criteria that separate them — not fewer options.</p>
</blockquote>
<p>Sixty-two words each carrying a gloss that distinguishes it (<em>annoyance: a small
thing, quickly over</em> against <em>resentment: an old grievance still carried</em>) are
sixty-two alternatives, not sixty-two synonyms. The count was never the problem.</p>
<hr>
<h2 id="2-ask-twice-anyway-for-a-reason-that-survived">2. Ask twice anyway — for a reason that survived</h2>
<p>Section 1 was the reason I split the question in two, and section 1 did not
hold. The split stayed, and it is worth being precise about what is now holding
it up, because &quot;it scored the same and I had already built it&quot; is not a reason.</p>
<p>Families are genuinely different answers — anger, fear, sadness and shame are
not wordings of each other. And once the family is fixed, so are its shades:
<em>which shade of anger</em> is a fair question, because the context has already ruled
out the fifty-four words that were never in the running.</p>
<p>What that buys, and one Choice over 62 words cannot, is <strong>two separate
uncertainty signals</strong>. Family confidence and word confidence are different
doubts: <em>&quot;I do not know whether this is sadness or affection&quot;</em> is not the same
failure as <em>&quot;it is clearly shame, but guilt or embarrassment?&quot;</em>. The page says
different things in each case. One Choice gives you one number that cannot tell
them apart.</p>
<p>The cost is honest too: a wobble in the family answer corrupts the word, because
once family says sadness, <em>nostalgia</em> is not on the ballot. That cost me one
text out of 23 — <em>&quot;drove past the old house, the tree we planted is taller than
the roof&quot;</em> came back <strong>sorrow</strong>.</p>
<p>So the naming became two Choices in sequence: the first picks the family, the
second picks the shade from that family alone. TypeSafe&#39;s docs say a second
request is warranted when an earlier answer determines the next question&#39;s
options. This is exactly that, and here is what it bought — scored on 24 texts
against a list of acceptable answers for each:</p>
<table>
<thead>
<tr>
<th>Design</th>
<th>Score</th>
</tr>
</thead>
<tbody><tr>
<td>Nearest word in the whole vocabulary, by published valence / arousal / dominance</td>
<td>3 / 24</td>
</tr>
<tr>
<td>Nearest word within a family, by rank on the axis that separates that family</td>
<td>9 / 24</td>
</tr>
<tr>
<td>Nearest word within a family, by distance in those published ratings</td>
<td>12 / 24</td>
</tr>
<tr>
<td><strong>Family chosen by the model, then the shade chosen by the model</strong></td>
<td><strong>22 / 24</strong></td>
</tr>
</tbody></table>
<p>Both of the two failures were the wrong <em>family</em>. Given the family, the shade was
right every single time.</p>
<p>The alternative was speculative fan-out: ask all eleven within-family Choices in
the first request, each stating its own family as a premise, and keep only the
answer belonging to the family that won. That keeps a reading to one request, at
roughly 3–4k input tokens against 1,250 + 444 for two. I reasoned about it and
rejected it: two requests won on cost, and on a latency story I could explain.</p>
<p>That paragraph was wrong when I published it. It is still here because
<a href="#7-i-reasoned-where-i-should-have-measured">section 7</a> is the part of this post
I would keep if I could only keep one.</p>
<hr>
<h2 id="3-knowing-which-half-to-give-the-model-is-a-decision-you-can-measure">3. <strong>Knowing which half</strong> to give the model is a decision you can measure</h2>
<p>The job is: read a text, name the feeling. That job splits between the code I
write and the model I call, and the only real decision is where the line falls —
how much of the work do I hand over?</p>
<p>The first three rows of that table are me keeping most of it.</p>
<p>I had a reason. Human ratings exist for exactly the dimensions I was measuring:
Warriner, Kuperman &amp; Brysbaert scored 13,915 English words for valence, arousal
and dominance by asking people. So the model&#39;s job shrinks to placing the text
on those three scales, and my code finishes the job — look up the nearest word
in the table, return it. Research-backed coordinates instead of somebody&#39;s
guesses. It looked like the serious version.</p>
<p>It scored <strong>3/24</strong>. Handing the model the whole naming job scored 22/24.</p>
<p>Three reasons, worth knowing before anyone else reaches for an emotion lexicon:</p>
<ul>
<li><strong>Three scales were really about two.</strong> Across these emotion words, valence
and dominance move together at <strong>+0.87</strong> — a word that reads pleasant almost
always reads in-control. The third scale is close to a copy of the first, so
it separates far less than the theory promises.</li>
<li><strong>The unpleasant words are crammed into one corner.</strong> Fear, frustration,
worry, terror, jealousy and embarrassment land in a ball small enough that
taking the nearest one is close to picking at random. The coordinates are
fuzzy to begin with: each is an average over about twenty raters who disagreed
by around 1.7 on a 1–9 scale.</li>
<li><strong>Rating a word on its own is not the same measurement as reading a
sentence.</strong> People rate the bare word <em>&quot;gratitude&quot;</em> as far more activated than
an actual grateful message reads. The two sets of numbers were never on the
same ruler.</li>
</ul>
<p>I tried to fix that last mismatch by re-centring both sides against a sample of
texts. It got worse, and instructively: the sample leaned negative, so the
correction leaned negative, and a plainly warm text landed below the middle and
got named from the sad half of the space. That is straightening a bent ruler
with a bent ruler.</p>
<p>The lesson is not <em>don&#39;t use lexicons</em>. It is that <strong>where the line falls
between what code owns and what the model owns is a design decision with a
number attached</strong>, and my intuition about it was wrong by a factor of seven.</p>
<p>The norms kept the one job they are good at. A Score is not given &quot;rate this 1
to 5&quot; — it is given a rubric, a written description of what each level means,
the way a grading key spells out what a B looks like. Every level of mine now
names words whose ratings were measured, so <em>&quot;as activated as rage or panic&quot;</em> is
a claim a reader can argue with. <em>&quot;Very aroused&quot;</em> is only a word getting louder.</p>
<hr>
<h2 id="4-confidence-is-peakedness-and-peakedness-misses-a-coin-toss">4. Confidence is peakedness — and peakedness misses a coin toss</h2>
<p>A Score&#39;s confidence measures how bunched together the answer is. It does not
measure how likely the model is to be right, which is the trap everyone hits
first.</p>
<p>Those two sound like the same thing until a reading like this one turns up.
Confidence came back at <strong>0.72</strong>, comfortably above any threshold I would set,
and here is what was underneath it:</p>
<table>
<thead>
<tr>
<th>Where the answer sat</th>
<th>Share</th>
</tr>
</thead>
<tbody><tr>
<td>The winning level</td>
<td><strong>50%</strong></td>
</tr>
<tr>
<td>The level right next to it</td>
<td><strong>49%</strong></td>
</tr>
<tr>
<td>The other three</td>
<td>1%</td>
</tr>
</tbody></table>
<p>The 0.72 is not lying. Ninety-nine percent of the weight really is in two
buckets and there is nothing anywhere else — that is about as bunched as an
answer gets.</p>
<p>It is also a coin toss. Choosing between the top two is 50 against 49, and my
page announced the winner in exactly the voice it uses at 0.98.</p>
<p>Confidence cannot catch this, because <em>bunched into two neighbours</em> is still
bunched. What I actually wanted was the <strong>gap between first and second place</strong>,
which is a different number entirely: 98 against 1 is a winner, 50 against 49 is
a tie with a rounding error. Confidence says the same thing about both — and the
gap is the one a reader cares about.</p>
<hr>
<h2 id="5-audit-your-questions-one-of-mine-fired-on-21-of-24-inputs">5. Audit your questions: one of mine fired on 21 of 24 inputs</h2>
<p>A question set grows by accretion. Each addition looks free, and none of them
announce that they have stopped saying anything.</p>
<p>So run them over a corpus and look at the spread. Mine had a Noul asking <em>is more
than one feeling present</em>. It returned ≥0.6 on <strong>twenty-one of twenty-four</strong>
texts. That is not a signal about the input, it is a property of writing — a
question paying tokens to tell you something you already knew.</p>
<p>Two more went for a different reason. They were informative, but nothing
downstream consumed them beyond appending a line to the output. A question whose
entire effect is an occasional footnote costs a reader more attention than it
returns.</p>
<p>The check is cheap: for every question, the min, max and spread of its answer
across a representative corpus. Anything that barely moves is either a gate or a
mistake.</p>
<hr>
<h2 id="6-two-small-things-that-cost-real-time">6. Two small things that cost real time</h2>
<p><strong>A rubric level must not contain a word from another axis.</strong> My control rubric
had a level reading <em>&quot;Overwhelmed: struggling to keep any grip on it&quot;</em>. That
primes the model with a feeling while asking it about agency, and it labels the
output with a word that is not a position on a control scale at all. Every level
of an axis has to be a point on that axis.</p>
<p><strong>Input is dominated by your rubrics, not by your input.</strong> Criteria are sent on
every call, so a one-line text costs almost exactly what a paragraph does. At
this size requests are the scarce resource and tokens are not — which is the
whole argument for batching every independent question into one call.</p>
<p>I stopped one clause too early, and wrote that a <em>dependent</em> second call is
therefore a real cost rather than a rounding error. Read it again: if requests
are scarce and tokens are not, the conclusion goes the other way.</p>
<hr>
<h2 id="7-i-reasoned-where-i-should-have-measured">7. I reasoned where I should have measured</h2>
<p>Someone read section 2, noticed I had talked myself out of the fan-out, and told
me to go and time both. It took forty minutes and a dollar&#39;s worth of nothing.</p>
<p>Run the same 24 texts through both designs, twice each, back to back on every
text so neither gets the warmer connection:</p>
<table>
<thead>
<tr>
<th></th>
<th>Requests</th>
<th>Latency</th>
<th>Input tokens</th>
<th>Cost / 1,000 readings</th>
</tr>
</thead>
<tbody><tr>
<td>Two requests, the second dependent</td>
<td>2</td>
<td><strong>762ms</strong></td>
<td>1,698</td>
<td>$0.071</td>
</tr>
<tr>
<td>One request, eleven speculative Choices</td>
<td>1</td>
<td><strong>398ms</strong></td>
<td>2,803</td>
<td>$0.118</td>
</tr>
</tbody></table>
<p>One request was faster on <strong>46 of 46</strong> pairs. It named the same family <strong>48 out
of 48 times</strong>, and the same shade 45 of 48 — where all three misses were texts
under 0.41 confidence that the page already reports as sitting between two
words, and where the <em>sequential</em> design disagreed with itself between rounds on
one of them. The eleven extra questions barely moved the six that were already
there: mean drift of 0.022 on a 0–4 axis score, and identical intent 48 times
out of 48.</p>
<p>So the fan-out is 1.9× faster for <strong>five hundredths of a cent</strong> a thousand
readings, and it halves the request count against the limit that actually binds.</p>
<p>Two things I had, and did not put together:</p>
<p><strong>Jev &quot;ingests the state once and evaluates every question against it in
parallel.&quot;</strong> That sentence is in TypeSafe&#39;s own model card, and I had quoted the
half of it that suited me. Question count is nearly free; a round trip is not.
Ten wasted questions cost less than one extra wait.</p>
<p><strong>Jev charges for input only — output tokens are free.</strong> The fan-out triples the
output, returning eleven probability distributions instead of one, and that is
worth exactly nothing on the bill.</p>
<p>What I actually got wrong is narrower than &quot;I didn&#39;t measure it&quot;, and more
useful. The docs say a second request is warranted when the first answer
determines the second question&#39;s <em>options</em>. It does — and I read that as
settling the matter. But <em>determined</em> is not <em>unknown</em>: there were only ever
eleven option sets, all of them written down in my own source file, so every one
of them could be asked on spec. The rule is about what the options are. It says
nothing about when you are allowed to ask.</p>
<p>The tell was in my own sentence. &quot;A latency story I could explain&quot; is not a
latency number. Anywhere a design note says <em>presumably</em>, <em>roughly</em>, or <em>I could
explain</em>, there is a measurement someone is about to make for you, and it is
cheaper to make it yourself.</p>
<hr>
<p>waif is live at <a href="https://quirkyagents.com/wizard/projects/waif/play/">/waif</a>. Every number in this post is on the page, under
<em>How it works</em>, next to the vocabulary it was scored against.</p>
]]></content:encoded></item><item><title>Choice, Score and Noul: four mistakes with Jev's primitives</title><link>https://quirkyagents.com/wizard/blog/jev-primitives-four-mistakes/</link><guid>https://quirkyagents.com/wizard/blog/jev-primitives-four-mistakes/</guid><pubDate>Sat, 19 Sep 2026 00:00:00 GMT</pubDate><description>I built a Magic 8 Ball on Jev, TypeSafe's calibrated-decision model, and misread its three primitives four times. The measured numbers, and what each mistake cost.</description><category>AI</category><category>Reliability</category><category>Calibration</category><category>TypeSafe</category><category>Jev</category><content:encoded><![CDATA[<p>Jev is a System One model from TypeSafe. It reads natural language like any LLM,
and instead of writing a reply it returns a probability distribution over
options you supply. No prose, no reasoning trace, no JSON to repair — a number
per option, summing to 1.</p>
<p>I spent a week building a <strong class="spell">Magic 8 Ball</strong> on it, which
sounds like a toy and turned out to be a decent test rig: every answer is a
judgment with no ground truth, which is where calibrated probabilities either
earn their keep or embarrass you. It&#39;s live at
<a href="https://quirkyagents.com/wizard/projects/jevball/play/">/jevball</a>.</p>
<p>I got the primitives wrong four times. Each mistake produced working code that
looked right, so they&#39;re worth writing down.</p>
<h2 id="the-three-primitives-in-one-paragraph-each">The three primitives in one paragraph each</h2>
<p><strong>Choice</strong> takes a map of named options and returns a probability for each, plus
the highest-scoring one. Routing a ticket to a department, classifying a
document, picking a handler.</p>
<p><strong>Score</strong> takes an ordered array of level descriptions — two to ten — and returns
a position along them that can land between levels. Severity, frustration, skill
level. Anything on a spectrum you can describe in words.</p>
<p><strong>Noul</strong> takes a yes/no statement and returns one number: the probability it&#39;s
true. No confidence field, because the probability <em>is</em> the answer.</p>
<p>All three take the same two inputs: <code>state</code>, the data being judged, and
<code>instructions</code>, the question asked about it. That&#39;s the whole surface.</p>
<h2 id="mistake-1-reading-the-score-as-a-winner">Mistake 1: reading the score as a winner</h2>
<p>A Score returns <code>score</code>, a position on your levels. My ball rounded it to pick a
bucket. Someone asked it <em>&quot;is black a color?&quot;</em> and got this:</p>
<figure class="code"><figcaption>text</figcaption><pre><code class="hljs">bars   : 0=2%  1=6%  2=19%  3=16%  4=56%
score  : 3.18</code></pre></figure>
<p><code>Math.round(3.18)</code> is 3 — a bucket holding <strong>16%</strong> — while 56% of the mass sat
on level 4 beside it.</p>
<p>The score is a <strong>probability-weighted average</strong>. That 19% on level 2 dragged the
average down across a bucket boundary. An average of a skewed distribution
points at a place the distribution isn&#39;t.</p>
<p>TypeSafe defines a Choice&#39;s answer as the option with the highest probability.
Score has no equivalent field, and I quietly substituted rounding for one. The
fix is to take the argmax of <code>probabilities</code> yourself:</p>
<figure class="code"><figcaption>js</figcaption><pre><code class="hljs language-js"><span class="hljs-keyword">const</span> probs = <span class="hljs-title class_">Array</span>.<span class="hljs-title function_">from</span>({ <span class="hljs-attr">length</span>: <span class="hljs-number">5</span> }, <span class="hljs-function">(<span class="hljs-params">_, i</span>) =&gt;</span> v.<span class="hljs-property">probabilities</span>[<span class="hljs-title class_">String</span>(i)] ?? <span class="hljs-number">0</span>);
<span class="hljs-keyword">const</span> level = probs.<span class="hljs-title function_">indexOf</span>(<span class="hljs-title class_">Math</span>.<span class="hljs-title function_">max</span>(...probs));</code></pre></figure>
<p>It agrees with rounding on every peaked distribution, which is why it survived
my whole test suite. It differs exactly when the distribution is skewed, which
is when it matters.</p>
<h2 id="mistake-2-reading-low-confidence-as-doubt-about-the-answer">Mistake 2: reading low confidence as doubt about the answer</h2>
<p>Every Choice and Score comes back with <code>confidence</code>, 0 to 1. I built a &quot;reply
hazy&quot; branch on the assumption that a vague question would produce a low one.</p>
<p>Then I measured it:</p>
<table>
<thead>
<tr>
<th>question</th>
<th>score</th>
<th>confidence</th>
</tr>
</thead>
<tbody><tr>
<td>Should I quit my job?</td>
<td>1.97</td>
<td><strong>0.98</strong></td>
</tr>
<tr>
<td>Will it rain next Tuesday?</td>
<td>1.99</td>
<td><strong>0.98</strong></td>
</tr>
<tr>
<td>Is black a color?</td>
<td>3.18</td>
<td><strong>0.33</strong></td>
</tr>
</tbody></table>
<p>The first two are as uncertain as a question gets, and confidence is 0.98.
Nearly all the mass sat on the level that <em>means</em> &quot;could go either way&quot; — the
distribution was a sharp spike on a level labelled uncertainty.</p>
<p>Confidence measures how <strong>peaked</strong> the distribution is. It answers &quot;did the
options separate cleanly&quot;, and it will happily report 0.98 for a confident
coin-flip. My hazy branch would have fired on precisely the wrong questions.</p>
<p>Low confidence has three different causes, and they need different fixes: the
question is genuinely contested, the options overlap each other, or <code>state</code>
doesn&#39;t contain enough to decide. Only the first one is the model being honest.</p>
<p>I tried to work out the formula by fitting twelve real responses. <code>max(p)</code> gets
closest at 0.04 mean error and matches several exactly, but not all —
<em>&quot;will humans land on Mars before 2050?&quot;</em> reported 0.58 against a peak of 0.50,
higher than the peak. The docs are upfront that it&#39;s a convenience statistic and
hand you the full <code>probabilities</code> so you can compute your own. If a decision
rides on it, do that.</p>
<h2 id="mistake-3-offering-options-that-split-their-own-vote">Mistake 3: offering options that split their own vote</h2>
<p>The classic 8 Ball has twenty answers. My first design asked Jev to pick one, as
a Choice.</p>
<p>Ten of those twenty mean &quot;yes&quot;. &quot;It is certain&quot;, &quot;Without a doubt&quot;, &quot;Yes
definitely&quot; — the same claim in different words. Here&#39;s what happens, on one
question with four different option sets:</p>
<table>
<thead>
<tr>
<th>options offered</th>
<th>answer</th>
<th>winner&#39;s share</th>
</tr>
</thead>
<tbody><tr>
<td><code>yes</code> / <code>no</code></td>
<td>yes</td>
<td>69%</td>
</tr>
<tr>
<td><code>yes</code> / <code>definitely</code> / <code>certainly</code> / <code>no</code></td>
<td>yes</td>
<td><strong>48%</strong></td>
</tr>
</tbody></table>
<p>The total yes-mass is identical: 70% in the second row, spread across three
near-synonyms. Nothing changed about the question. The options ate each other,
and confidence fell from 0.38 to 0.31 reporting a disagreement that existed only
in my option list.</p>
<p>So the ball asks Jev for one of <strong>five ordered buckets</strong>, and the code picks
which of that bucket&#39;s phrasings to show. One judgment with a right answer goes
to the model; the theatre stays in software.</p>
<h2 id="mistake-4-forgetting-that-the-options-are-the-prompt">Mistake 4: forgetting that the options are the prompt</h2>
<p>Your question IDs never reach the model. The <code>instructions</code> and every entry in
<code>criteria</code> do — names and descriptions both. The option list isn&#39;t a filter you
apply to a result. It&#39;s part of what you&#39;re asking.</p>
<p>Same question, same model, options changed:</p>
<table>
<thead>
<tr>
<th>options offered</th>
<th>answer</th>
</tr>
</thead>
<tbody><tr>
<td><code>yes</code> / <code>no</code></td>
<td><strong>yes</strong> (69%)</td>
</tr>
<tr>
<td><code>yes_everyday</code> / <code>no_physics</code> / <code>depends</code></td>
<td><strong>depends</strong> (51%)</td>
</tr>
</tbody></table>
<p>&quot;Depends&quot; won — an answer that simply didn&#39;t exist in the first row. Black is a
colour in everyday use and the absence of light in physics, so &quot;depends&quot; is
arguably the best answer available. It was unreachable until I offered it.</p>
<p>The docs put it plainly: the model cannot choose an omitted value. Leave an
option out and you haven&#39;t biased the result, you&#39;ve made it impossible.</p>
<h2 id="what-this-is-actually-for">What this is actually for</h2>
<p>Every judgment above cost about <strong>$0.000014</strong>. Jev charges $0.042 per million
input tokens and nothing for output, against $1.00/$5.00 for the cheapest
frontier model I&#39;d otherwise reach for — roughly 27× on a like-for-like
classification, before an LLM&#39;s format instructions and reasoning tokens widen
it further.</p>
<p>That price only matters because of what you give up. Jev writes no prose,
explains nothing, and can&#39;t answer anything whose answer space you can&#39;t
enumerate. For summarising, drafting, coding or open-ended reasoning it&#39;s the
wrong tool entirely.</p>
<p>What it&#39;s good at is the narrow decision a pipeline makes ten thousand times a
day, where you need a number you can threshold on. The documented pattern is to
put it in front of the expensive work: score everything, route the confident
cases to deterministic code, escalate the rest to a model or a person. You pay
fourteen dollars a million to decide, and frontier prices only on the slice that
earned it.</p>
<p>One caveat worth keeping. Calibration is a property of <em>groups</em> of predictions —
across many answers, the ones marked 0.8 should be right about 80% of the time.
It guarantees nothing about any single answer, and it&#39;s measured on TypeSafe&#39;s
data, not yours. Validate it in your own domain before trusting a threshold.</p>
<hr>
<p>The ball is at <a href="https://quirkyagents.com/wizard/projects/jevball/play/">/jevball</a>. Ask it something you actually want to know
and open the panel underneath — it shows the full distribution, the confidence,
and which of the four questions produced the answer.</p>
]]></content:encoded></item><item><title>AI turns emails into spreadsheet rows. Did it get them right?</title><link>https://quirkyagents.com/wizard/blog/measure-ai-reading-inbound-emails/</link><guid>https://quirkyagents.com/wizard/blog/measure-ai-reading-inbound-emails/</guid><pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate><description>Eight real inbound emails through Claude Haiku: 61 of 64 fields correct, every miss on one field. What a failure taxonomy shows that an accuracy score hides.</description><category>AI</category><category>Email Automation</category><category>Data Entry</category><category>Reliability</category><category>Evaluation</category><content:encoded><![CDATA[<p>Somebody in your office opens the sales inbox every morning and types what they find into a spreadsheet. Company, contact, what they want, how many, by when, budget. It takes a few minutes an email, it happens every day, and it is exactly the kind of job people now hand to an LLM.</p>
<p>The handover usually goes well for about a week. The model returns valid JSON every time. Every field is filled. Nothing throws an error.</p>
<p>Then someone notices the follow-up went to the wrong person.</p>
<hr>
<h2 id="wrong-output-looks-exactly-like-right-output">Wrong output looks exactly like <strong class="spell">right output</strong></h2>
<p>An extraction model does not fail the way code fails. There is no stack trace and no empty response — you get a complete, confident, well-formed row, with no signal about which of its eight fields you can trust.</p>
<p>So I wrote down the correct answer for eight real inbound enquiries and measured it properly. Here is what I expected to break.</p>
<h3 id="the-forwarded-thread">The forwarded thread</h3>
<figure class="code"><figcaption>text</figcaption><pre><code class="hljs">From: Dan &lt;dan@brightline.io&gt;
Subject: Fwd: quick question

---------- Forwarded message ----------
From: Marcus Ellery &lt;m.ellery@brightline.io&gt;

Dan - can you ask them whether they do anodised finish? We'd need
about 250 brackets, the aluminium ones we discussed. Budget is
around 4k GBP, maybe a bit more if the finish is good.</code></pre></figure>
<p>The lead is Marcus. The <code>From:</code> header says Dan. Read the header first and you get the product right, the quantity right, the budget right — and address your quote to the wrong human.</p>
<h3 id="the-email-that-is-not-a-lead">The email that is not a lead</h3>
<figure class="code"><figcaption>text</figcaption><pre><code class="hljs">A tender matching your registered categories has been published.
Reference: TN-2026-0884
Closing: 2026-09-12
Do not reply to this address.</code></pre></figure>
<p>No company, no contact, no quantity, no budget. The correct extraction is mostly <code>null</code> — but models are obliging, and asking for eight fields tends to get you eight fields, invented where the text runs out.</p>
<h3 id="dates-that-only-exist-relative-to-the-email">Dates that only exist relative to the email</h3>
<blockquote>
<p>&quot;required before 25th aug&quot; · &quot;by the first week of September&quot; · &quot;before the end of October&quot;</p>
</blockquote>
<p>None of those are dates until they are resolved against the email&#39;s own <code>Date:</code> header. Resolve them against today instead and you are quietly a year out every January.</p>
<h3 id="the-enquiry-in-italian">The enquiry in Italian</h3>
<figure class="code"><figcaption>text</figcaption><pre><code class="hljs">Vorrei un preventivo per 600 maniglie in ottone.
Consegna a Bologna entro fine settembre. Budget indicativo 3.500 euro.</code></pre></figure>
<p>600 brass handles, Bologna, end of September, €3,500. Easy to get <em>almost</em> right: <code>3.500</code> is three thousand five hundred, not three and a half.</p>
<h3 id="the-polite-no">The polite no</h3>
<blockquote>
<p>&quot;We are only collecting information for a project starting next year. I cannot give quantities yet and there is no budget approved.&quot;</p>
</blockquote>
<p>Every field sales cares about is genuinely absent. An extractor that fills them in has not saved anyone time; it has created work.</p>
<hr>
<h2 id="what-actually-happened">What actually happened</h2>
<p>I ran all eight through Claude Haiku with a plain extraction prompt. Eight emails, eight fields each, sixty-four fields total. It cost 0.6 cents.</p>
<figure class="code"><figcaption>text</figcaption><pre><code class="hljs">Overall: 61/64 fields correct (95.3%)
Successful: 8  Failed: 0

Failure modes:
  wrong_value : 3</code></pre></figure>
<p><strong>Every one of the five cases above came back correct.</strong> It took Marcus over Dan on the forwarded thread. It returned nulls for the tender notice instead of inventing a company. It resolved &quot;before 25th aug&quot; against the email&#39;s own date. It read <code>3.500 euro</code> as 3500 and translated <em>maniglie in ottone</em> into brass handles. It left the polite no almost entirely empty, which was the right answer.</p>
<p>All three misses were the same field — <code>product</code> — and they looked like this:</p>
<table>
<thead>
<tr>
<th>Expected</th>
<th>Extracted</th>
</tr>
</thead>
<tbody><tr>
<td><code>316 stainless sheet, 2mm</code></td>
<td><code>316 stainless steel sheet, 2mm</code></td>
</tr>
<tr>
<td><code>anodised aluminium brackets</code></td>
<td><code>aluminium brackets with anodised finish</code></td>
</tr>
<tr>
<td><code>M10 x 40 hex bolts; M10 nuts</code></td>
<td><code>M10 x 40 hex bolts, M10 nuts</code></td>
</tr>
</tbody></table>
<p>The model is not wrong in any of those. <strong>I am</strong> — or rather, exact string match is the wrong check for a free-text field, and a semicolon is not a fact about the world.</p>
<hr>
<h2 id="that-is-what-the-taxonomy-is-for">That is what the taxonomy is for</h2>
<p>If all I had was <strong>95.3%</strong>, I would be shopping for a better model right now. I would be wrong, and I would have paid for it.</p>
<p>The number that mattered was not the percentage. It was that all three failures carried the same tag on the same field. Four tags, four different bugs:</p>
<table>
<thead>
<tr>
<th>Failure mode</th>
<th>What happened</th>
<th>What it means</th>
</tr>
</thead>
<tbody><tr>
<td><code>missed_field</code></td>
<td>The email stated it; the model left it blank</td>
<td>The prompt is under-specified</td>
</tr>
<tr>
<td><code>hallucination</code></td>
<td>The email stated nothing; the model filled it in</td>
<td>The prompt does not permit nulls</td>
</tr>
<tr>
<td><code>wrong_format</code></td>
<td>Right value, wrong shape — <code>3.500</code> read as 3.5</td>
<td>You need normalisation, not a bigger model</td>
</tr>
<tr>
<td><code>wrong_value</code></td>
<td>Genuinely a different answer</td>
<td>Either the model is wrong, or your check is</td>
</tr>
</tbody></table>
<p>Three <code>hallucination</code> tags scattered across six fields is a prompt problem. Three <code>wrong_value</code> tags stacked on one free-text field is a <em>measurement</em> problem. Averaged into a single percentage, those two situations are indistinguishable — and they have nothing in common.</p>
<p>The fix here is not a better model. It is deciding what &quot;correct&quot; means for <code>product</code>: normalise before comparing, or score that field with a judge instead of <code>==</code>, or stop scoring it at all and accept that a human reads it anyway.</p>
<hr>
<h2 id="doing-it-yourself">Doing it yourself</h2>
<p>Write down the correct answer for a few dozen real emails, once. That labelled set is the asset; everything else is plumbing.</p>
<p>Then run your extractor against it. I maintain a small open-source harness for exactly this — <a href="https://github.com/dave8172/doceval">doceval</a> — which takes any extraction function and any schema:</p>
<figure class="code"><figcaption>bash</figcaption><pre><code class="hljs language-bash">pip install doceval

doceval run \
  --docs    ./emails \
  --labels  ./labels \
  --extractor my_module:extract</code></pre></figure>
<p>Your extractor is any Python function that takes bytes and returns a dict:</p>
<figure class="code"><figcaption>python</figcaption><pre><code class="hljs language-python"><span class="hljs-keyword">def</span> <span class="hljs-title function_">extract</span>(<span class="hljs-params">doc_bytes: <span class="hljs-built_in">bytes</span>, filepath: <span class="hljs-built_in">str</span></span>) -&gt; <span class="hljs-built_in">dict</span>:
    email_text = doc_bytes.decode(<span class="hljs-string">&quot;utf-8&quot;</span>, errors=<span class="hljs-string">&quot;replace&quot;</span>)
    <span class="hljs-comment"># call whatever model you like</span>
    <span class="hljs-keyword">return</span> {<span class="hljs-string">&quot;company&quot;</span>: <span class="hljs-string">&quot;Brightline Systems Ltd&quot;</span>, <span class="hljs-string">&quot;contact_name&quot;</span>: <span class="hljs-string">&quot;Marcus Ellery&quot;</span>, ...}</code></pre></figure>
<p>A label is the correct answer, <code>null</code> included, because the nulls are half the point:</p>
<figure class="code"><figcaption>json</figcaption><pre><code class="hljs language-json"><span class="hljs-punctuation">{</span>
  <span class="hljs-attr">&quot;company&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-literal"><span class="hljs-keyword">null</span></span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;contact_name&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-literal"><span class="hljs-keyword">null</span></span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;product&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;galvanised steel conduit&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;quantity&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-literal"><span class="hljs-keyword">null</span></span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;target_date&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;2026-09-12&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;budget&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-literal"><span class="hljs-keyword">null</span></span>
<span class="hljs-punctuation">}</span></code></pre></figure>
<p>The eight emails above ship with it as a runnable example, labels included, if you want something to point it at before labelling your own.</p>
<p><strong>One honest caveat about the numbers here:</strong> eight emails, one model, and I wrote both the labels and the prompt. That is a demonstration of the method, not a benchmark — and if I had run it on eighty emails from a real inbox, the interesting failures would almost certainly be somewhere else. Which is the argument for running it on yours.</p>
<hr>
<h2 id="the-part-worth-keeping">The part worth keeping</h2>
<p>The labelled set outlives every other decision. Models change, prompts change, providers change — and each time, the only thing that tells you whether the change helped is a set of emails where you already wrote down the right answer.</p>
<p>An afternoon of labelling buys that permanently. Skipping it means finding out from a customer.</p>
]]></content:encoded></item><item><title>98.3% vs 96.3%: the cost of an auditable AI extraction pipeline</title><link>https://quirkyagents.com/wizard/blog/invoice-extraction-cost-accuracy-benchmark/</link><guid>https://quirkyagents.com/wizard/blog/invoice-extraction-cost-accuracy-benchmark/</guid><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate><description>105 labeled invoices and receipts: Claude Opus vs Haiku+Sonnet, a 2% accuracy gap and a 12× cost gap, plus the eval set and audit log behind the numbers.</description><category>AI</category><category>Document Extraction</category><category>Reliability</category><category>Benchmarks</category><category>Cost Analysis</category><content:encoded><![CDATA[<p>Most AI extraction pipelines are built to pass a demo: run a clean invoice through, get a clean JSON back, ship it. The problems show up later — in the reconciliation diff, the failed audit, the support ticket that says &quot;the totals are wrong.&quot;</p>
<p>I built the layer that catches this before it happens: labeled eval sets, accuracy reports, confidence scoring, audit logs. Along the way I also benchmarked the two obvious model choices against the same data, so you don&#39;t have to guess which one your pipeline needs. Here&#39;s what 105 labeled documents taught me.</p>
<hr>
<h2 id="why-extraction-pipelines-quietly-fail">Why extraction pipelines <strong class="spell">quietly fail</strong></h2>
<p>Extraction looks easy because it looks right. The model fills in all the fields, the JSON is valid, nothing throws an error. The failure modes are invisible:</p>
<ul>
<li><strong>Silent wrong answers.</strong> The model reads &quot;132,30&quot; (Swedish decimal notation) and returns <code>13230</code>. No error. The invoice is just wrong.</li>
<li><strong>Hallucinated fields.</strong> A blurry scanned receipt gets a subtotal that doesn&#39;t exist anywhere on the document.</li>
<li><strong>Inconsistent formatting.</strong> Currency as <code>$</code> on some invoices, <code>USD</code> on others, nothing on scanned ones — your downstream code breaks on whichever one you didn&#39;t test.</li>
</ul>
<p>Without a labeled eval set and systematic accuracy measurement, you find these problems in production. That&#39;s too late.</p>
<hr>
<h2 id="the-eval-set">The eval set</h2>
<p>The foundation is 105 labeled documents: 100 digital invoices and 5 scanned thermal receipts. Each document has a ground truth JSON with the correct values for every field — vendor, document number, date, currency, line items, totals.</p>
<p>A label looks like this:</p>
<figure class="code"><figcaption>json</figcaption><pre><code class="hljs language-json"><span class="hljs-punctuation">{</span>
  <span class="hljs-attr">&quot;document_id&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;invoice_Shahid_Shariari_30140&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;filename&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;invoice_Shahid Shariari_30140.pdf&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;processing_mode_expected&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;embedded_text&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;difficulty&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;easy&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;ground_truth&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span>
    <span class="hljs-attr">&quot;vendor&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;SuperStore&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;document_number&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;30140&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;date&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;2012-11-15&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;currency&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;USD&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;total&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;748.36&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;shipping&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;73.00&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;line_items&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">[</span>
      <span class="hljs-punctuation">{</span> <span class="hljs-attr">&quot;description&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;Safco 3-Shelf Cabinet, Traditional&quot;</span><span class="hljs-punctuation">,</span> <span class="hljs-attr">&quot;quantity&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;2&quot;</span><span class="hljs-punctuation">,</span> <span class="hljs-attr">&quot;unit_price&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;337.68&quot;</span><span class="hljs-punctuation">,</span> <span class="hljs-attr">&quot;amount&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;675.36&quot;</span> <span class="hljs-punctuation">}</span>
    <span class="hljs-punctuation">]</span>
  <span class="hljs-punctuation">}</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;eval_notes&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span> <span class="hljs-attr">&quot;line_items_evaluated&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-literal"><span class="hljs-keyword">true</span></span> <span class="hljs-punctuation">}</span>
<span class="hljs-punctuation">}</span></code></pre></figure>
<p>The eval runner loads each label, runs the extractor, and compares every field in the ground truth against what the model returned — handling numeric equivalence (<code>&quot;$132.30&quot;</code> <strong> <code>&quot;132.30&quot;</code>), European decimal formats (<code>&quot;132,30&quot;</code> </strong> <code>&quot;132.30&quot;</code>), and date normalization (<code>&quot;Dec 27 2012&quot;</code> == <code>&quot;2012-12-27&quot;</code>).</p>
<hr>
<h2 id="two-configs-same-eval-set">Two configs, same eval set</h2>
<p>Two questions come up the moment you build an LLM extraction pipeline: which model, and at what cost? &quot;It depends&quot; isn&#39;t an answer, so I ran both configurations against the same 105 documents and measured accuracy, failure modes, and cost.</p>
<table>
<thead>
<tr>
<th>Config</th>
<th>Model(s)</th>
<th>Image resolution</th>
<th>Cost per run</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Cheap</strong></td>
<td>Haiku 4.5 (text) + Sonnet 4.6 (vision)</td>
<td>1500px max</td>
<td>~$0.08</td>
</tr>
<tr>
<td><strong>Expensive</strong></td>
<td>Opus 4.8</td>
<td>7500px max</td>
<td>~$1.00</td>
</tr>
</tbody></table>
<p>Cost ratio: <strong>12.5×.</strong></p>
<table>
<thead>
<tr>
<th>Config</th>
<th>Overall accuracy</th>
<th>Perfect documents</th>
<th>Digital</th>
<th>Scanned</th>
</tr>
</thead>
<tbody><tr>
<td>Haiku+Sonnet · 1500px</td>
<td><strong>96.3%</strong> (918/953 fields)</td>
<td>77.1% (81/105)</td>
<td><strong>97.9%</strong></td>
<td><strong>69.8%</strong></td>
</tr>
<tr>
<td>Opus · 7500px</td>
<td><strong>98.3%</strong> (937/953 fields)</td>
<td>88.6% (93/105)</td>
<td><strong>99.2%</strong></td>
<td><strong>83.0%</strong></td>
</tr>
</tbody></table>
<p>A <strong>2-point accuracy gap</strong> for 12.5× the cost. Worth knowing where that 2% actually lives before deciding it&#39;s worth paying for.</p>
<p><strong>On digital invoices:</strong> 97.9% vs 99.2% — a 1.3-point gap. On 1,000 invoices, that&#39;s the difference between ~21 errors and ~8.</p>
<p><strong>On scanned documents:</strong> 69.8% vs 83.0% — a 13-point gap. Scanned receipts are where Opus earns its price: OCR quality, resolution, and degraded thermal printing all compound, and the extra resolution genuinely helps here.</p>
<p>If your document mix is mostly digital PDFs, Haiku+Sonnet is nearly indistinguishable from Opus. If you&#39;re processing scanned documents at volume, the gap is real.</p>
<h3 id="field-by-field">Field by field</h3>
<p>Not all fields are equal:</p>
<table>
<thead>
<tr>
<th>Field</th>
<th>Haiku+Sonnet</th>
<th>Opus</th>
<th>Gap</th>
</tr>
</thead>
<tbody><tr>
<td><code>shipping</code></td>
<td>100.0%</td>
<td>100.0%</td>
<td>0</td>
</tr>
<tr>
<td><code>date</code></td>
<td>99.0%</td>
<td>100.0%</td>
<td>1.0%</td>
</tr>
<tr>
<td><code>total</code></td>
<td>99.0%</td>
<td>100.0%</td>
<td>1.0%</td>
</tr>
<tr>
<td><code>currency</code></td>
<td>99.0%</td>
<td>100.0%</td>
<td>1.0%</td>
</tr>
<tr>
<td><code>subtotal</code></td>
<td>97.1%</td>
<td>100.0%</td>
<td>2.9%</td>
</tr>
<tr>
<td><code>discount</code></td>
<td>99.0%</td>
<td>99.0%</td>
<td>0</td>
</tr>
<tr>
<td><code>vendor</code></td>
<td>98.1%</td>
<td>99.0%</td>
<td>0.9%</td>
</tr>
<tr>
<td><code>document_number</code></td>
<td>95.2%</td>
<td>95.2%</td>
<td>0</td>
</tr>
<tr>
<td><code>line_items</code></td>
<td>80.6%</td>
<td>92.2%</td>
<td><strong>11.6%</strong></td>
</tr>
<tr>
<td><code>tax</code></td>
<td>80.0%</td>
<td>80.0%</td>
<td>0</td>
</tr>
</tbody></table>
<p><code>line_items</code> carries almost the entire gap — reading and counting structured table rows is hard on lower-resolution scans, and it closes considerably on digital invoices. <code>document_number</code> is identical across both configs for a simpler reason: scanned receipts carry several reference numbers (order number, transaction ID, loyalty card), and both models struggle equally to pick the right one.</p>
<h3 id="resolution-matters-more-than-you-39-d-expect">Resolution matters more than you&#39;d expect</h3>
<p>Scaling from 7500px to 1500px per page cuts API cost roughly 25× with barely any accuracy loss on digital documents — invoice text is legible at 1500px, and the model doesn&#39;t gain much from pixels it doesn&#39;t need. Higher resolution earns its keep only on scans, up to the point where the text becomes readable.</p>
<h3 id="which-one-to-use">Which one to use</h3>
<p><strong>Haiku+Sonnet at 1500px, if:</strong> your documents are mostly digital PDFs, you&#39;re processing at volume, and you can route low-confidence extractions to a human.</p>
<p><strong>Opus at 7500px, if:</strong> you&#39;re processing scanned documents at meaningful volume, every field has to be right, or you&#39;re pulling line items from messy layouts.</p>
<p>The math: Haiku+Sonnet costs ~$0.00076/doc, Opus ~$0.0095/doc. At 10,000 documents a month, that&#39;s $7.60 vs $95.</p>
<hr>
<h2 id="the-accuracy-floor-and-how-it-closed">The accuracy floor, and how it closed</h2>
<p>Before any tuning, the first run came in at <strong>79.3%.</strong> That sounds bad. Turns out, it was informative.</p>
<p>Categorizing every mismatch turned up four root causes:</p>
<table>
<thead>
<tr>
<th>Issue</th>
<th>Example</th>
<th>Fix</th>
</tr>
</thead>
<tbody><tr>
<td>Currency format</td>
<td>Model returned <code>$</code> instead of <code>USD</code></td>
<td>Prompt: use ISO 4217 codes</td>
</tr>
<tr>
<td>Verbosity</td>
<td><code>Visa ending in 0627</code> instead of <code>Visa</code></td>
<td>Prompt: card network name only</td>
</tr>
<tr>
<td>European decimals</td>
<td><code>&quot;132,30&quot;</code> compared as not equal to <code>&quot;132.30&quot;</code></td>
<td>Eval: locale-aware numeric parser</td>
</tr>
<tr>
<td>Line item parsing</td>
<td>Descriptions mangled at comma splits</td>
<td>Prompt: strip after dash separator only</td>
</tr>
</tbody></table>
<p>Fixing all four got both configs to the headline numbers above. The lesson: the first accuracy number tells you where the problems are, not where the ceiling is.</p>
<hr>
<h2 id="confidence-scoring-the-model-knows-what-it-doesn-39-t-know">Confidence scoring: <strong>the model knows what it doesn&#39;t know</strong></h2>
<p>Every extraction includes a self-reported confidence level — high, medium, or low — plus a list of uncertain fields and a short note explaining why.</p>
<p>It predicts actual accuracy well:</p>
<table>
<thead>
<tr>
<th>Confidence</th>
<th>Field accuracy</th>
<th>Documents</th>
</tr>
</thead>
<tbody><tr>
<td>High</td>
<td><strong>97.7%</strong></td>
<td>101 / 105</td>
</tr>
<tr>
<td>Medium</td>
<td><strong>71.9%</strong></td>
<td>3 / 105</td>
</tr>
<tr>
<td>Low</td>
<td><strong>50.0%</strong></td>
<td>1 / 105</td>
</tr>
</tbody></table>
<p>The one low-confidence document — a dense 80-item Dollarstore thermal receipt — also had the worst accuracy. The model flagged it correctly, without being told what &quot;correct&quot; was.</p>
<p>The practical use: <strong>route by confidence, not by document type.</strong> Send high-confidence extractions straight to processing. Route the 4% that aren&#39;t to a human. You&#39;re not reviewing 105 documents — you&#39;re reviewing 4, and the effective accuracy on the auto-processed 96% comes out to 97.7%, at Haiku+Sonnet pricing.</p>
<hr>
<h2 id="the-audit-log">The audit log</h2>
<p>Every extraction writes a structured log entry before anything downstream sees the data:</p>
<ul>
<li><strong>What went in:</strong> filename, processing mode (digital vs scanned), input size</li>
<li><strong>What came out:</strong> all extracted fields</li>
<li><strong>How confident:</strong> overall level, uncertain fields, the model&#39;s own explanation</li>
<li><strong>Operational metadata:</strong> which model, how long the call took, UTC timestamp</li>
</ul>
<p>PII redaction runs on every string value before the log is written — card numbers, emails, phone numbers, SSNs. The log is <strong>safe to store and share</strong> even when the source documents aren&#39;t.</p>
<p>A real entry looks like this:</p>
<figure class="code"><figcaption>json</figcaption><pre><code class="hljs language-json"><span class="hljs-punctuation">{</span>
  <span class="hljs-attr">&quot;id&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;b1d01383-42fd-4eac-acc0-f333d31a9409&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;timestamp&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;2026-06-15T13:43:51Z&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;filename&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;invoice_Shahid Shariari_30140.pdf&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;mode&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;embedded_text&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;model&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;claude-haiku-4-5&quot;</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;duration_s&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-number">2.13</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;confidence&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span>
    <span class="hljs-attr">&quot;overall&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;high&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;uncertain_fields&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">[</span><span class="hljs-punctuation">]</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;notes&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;All key fields are clearly visible and unambiguous.&quot;</span>
  <span class="hljs-punctuation">}</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;input_summary&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span> <span class="hljs-attr">&quot;char_count&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-number">386</span> <span class="hljs-punctuation">}</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;extraction&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-punctuation">{</span>
    <span class="hljs-attr">&quot;vendor&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;SuperStore&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;document_number&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;30140&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;date&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;2012-11-15&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;currency&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;USD&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;total&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;748.36&quot;</span><span class="hljs-punctuation">,</span>
    <span class="hljs-attr">&quot;shipping&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-string">&quot;73.00&quot;</span>
  <span class="hljs-punctuation">}</span><span class="hljs-punctuation">,</span>
  <span class="hljs-attr">&quot;pii_redacted&quot;</span><span class="hljs-punctuation">:</span> <span class="hljs-literal"><span class="hljs-keyword">false</span></span>
<span class="hljs-punctuation">}</span></code></pre></figure>
<p>This is what handing a pipeline to an auditor looks like in practice — not a demo showing the happy path, but a log of exactly what happened, for every document, with the model&#39;s own confidence attached.</p>
<hr>
<h2 id="known-limitations-the-honest-ones">Known limitations <strong>(the honest ones)</strong></h2>
<ul>
<li><strong>Scanned documents:</strong> 69.8% vs 97.9% for digital. OCR quality and resolution drive the gap — dense thermal receipts in poor light are genuinely hard.</li>
<li><strong>Non-English locales:</strong> the eval set includes Swedish receipts. Decimal separators, currency conventions, and vendor formats vary; prompt tuning per locale helps.</li>
<li><strong>Template concentration:</strong> 100 of the 105 digital invoices share one template. Real invoice variance will affect accuracy — these numbers are a ceiling on a homogeneous dataset, not a floor on a diverse one.</li>
</ul>
<hr>
<h2 id="methodology">Methodology</h2>
<ul>
<li><strong>Dataset:</strong> 105 labeled documents — 100 digital invoices (single vendor template) + 5 scanned thermal receipts (Swedish, mixed vendors)</li>
<li><strong>Fields evaluated:</strong> vendor, document_number, date, currency, total, subtotal, shipping, discount, tax, payment_method, line_items</li>
<li><strong>Comparison:</strong> numeric normalization across currency symbols and locale formats; date normalization across ISO and abbreviated formats</li>
<li><strong>Eval tool:</strong> <a href="https://github.com/dave8172/doceval">doceval</a> — open-source, schema-agnostic field-level accuracy harness</li>
</ul>
<hr>
<p>If your extraction pipeline needs the same treatment — eval set, benchmark, audit log, the works — <a href="mailto:hello@quirkyagents.com">reach out</a>.</p>
]]></content:encoded></item></channel></rss>
