Memories
Writing and benchmarks from real builds: what broke, what it cost, and the number that settled it.
Writing · 5 Oct 2026
Building and using MCP servers: 4 things I learned
How MCP works, one tool call step by step, and four things I learned building and using MCP servers, from tool descriptions to where secrets live.
Writing · 5 Oct 2026
Local vs remote MCP servers: how each connects, step by step
How a local MCP server lives and dies with your session, and how remote MCP sign-in works: OAuth, PKCE and tokens, tested on Upwork, Notion and Todoist.
Writing · 3 Oct 2026
How I calibrated an LLM judge to grade like me, 25× cheaper
Evals on 188 answers from ChatGPT, Claude, Chatbase and my own pipeline over product datasheets, and a hybrid judge with Jev that grades like I do.
Writing · 2 Oct 2026
Can an AI agent submit Upwork proposals? Yes, with your OK
Upwork's official MCP lets Claude, Codex or Cursor send a proposal for you. What it shows before sending, what stays on upwork.com, and what I got wrong.
Writing · 20 Sep 2026
Ask twice: seven measurements from building with Jev
A coordinate lookup scored 3/24 where handing the model the naming job scored 22/24. Then two claims I had argued rather than measured turned out to be wrong.
Writing · 19 Sep 2026
Choice, Score and Noul: four mistakes with Jev's primitives
I built a Magic 8 Ball on Jev, TypeSafe's calibrated-decision model, and misread its three primitives four times. The measured numbers, and what each mistake cost.
Writing · 4 Sep 2026
AI turns emails into spreadsheet rows. Did it get them right?
Eight real inbound emails through Claude Haiku: 61 of 64 fields correct, every miss on one field. What a failure taxonomy shows that an accuracy score hides.
Writing · 17 Jun 2026
98.3% vs 96.3%: the cost of an auditable AI extraction pipeline
105 labeled invoices and receipts: Claude Opus vs Haiku+Sonnet, a 2% accuracy gap and a 12× cost gap, plus the eval set and audit log behind the numbers.
Benchmark · 31 May 2026
Self-Hosted Whisper Benchmark
A self-hosted meeting-transcription pipeline, plus a benchmark of Whisper model sizes on a real 32-minute recording to find the accuracy/speed sweet spot.
Benchmark · 12 Feb 2025
GPT-4o vs. Veryfi Extraction Benchmark
GPT-4o against Veryfi on irregular financial scans, across several document types — where each one wins, and what each extraction costs.