# Memories (writing and benchmarks)

Writing and benchmarks from real builds: what broke, what it cost, and the number that settled it. By dave8172. RSS: https://quirkyagents.com/wizard/blog/rss.xml

- [Building and using MCP servers: 4 things I learned](https://quirkyagents.com/wizard/blog/building-mcp-servers-lessons/index.md): How MCP works, one tool call step by step, and four things I learned building and using MCP servers, from tool descriptions to where secrets live. (2026-10-05)
- [Local vs remote MCP servers: how each connects, step by step](https://quirkyagents.com/wizard/blog/local-vs-remote-mcp-servers/index.md): How a local MCP server lives and dies with your session, and how remote MCP sign-in works: OAuth, PKCE and tokens, tested on Upwork, Notion and Todoist. (2026-10-05)
- [How I calibrated an LLM judge to grade like me, 25× cheaper](https://quirkyagents.com/wizard/blog/calibrate-llm-judge-evals/index.md): Evals on 188 answers from ChatGPT, Claude, Chatbase and my own pipeline over product datasheets, and a hybrid judge with Jev that grades like I do. (2026-10-03)
- [Can an AI agent submit Upwork proposals? Yes, with your OK](https://quirkyagents.com/wizard/blog/can-ai-agent-submit-upwork-proposals/index.md): Upwork's official MCP lets Claude, Codex or Cursor send a proposal for you. What it shows before sending, what stays on upwork.com, and what I got wrong. (2026-10-02)
- [Ask twice: seven measurements from building with Jev](https://quirkyagents.com/wizard/blog/jev-ask-twice/index.md): A coordinate lookup scored 3/24 where handing the model the naming job scored 22/24. Then two claims I had argued rather than measured turned out to be wrong. (2026-09-20)
- [Choice, Score and Noul: four mistakes with Jev's primitives](https://quirkyagents.com/wizard/blog/jev-primitives-four-mistakes/index.md): I built a Magic 8 Ball on Jev, TypeSafe's calibrated-decision model, and misread its three primitives four times. The measured numbers, and what each mistake cost. (2026-09-19)
- [AI turns emails into spreadsheet rows. Did it get them right?](https://quirkyagents.com/wizard/blog/measure-ai-reading-inbound-emails/index.md): Eight real inbound emails through Claude Haiku: 61 of 64 fields correct, every miss on one field. What a failure taxonomy shows that an accuracy score hides. (2026-09-04)
- [98.3% vs 96.3%: the cost of an auditable AI extraction pipeline](https://quirkyagents.com/wizard/blog/invoice-extraction-cost-accuracy-benchmark/index.md): 105 labeled invoices and receipts: Claude Opus vs Haiku+Sonnet, a 2% accuracy gap and a 12× cost gap, plus the eval set and audit log behind the numbers. (2026-06-17)
- [Self-Hosted Whisper Benchmark](https://quirkyagents.com/wizard/projects/whisper-transcription-benchmark/index.md): A self-hosted meeting-transcription pipeline, plus a benchmark of Whisper model sizes on a real 32-minute recording to find the accuracy/speed sweet spot. (2026-05-31)
- [GPT-4o vs. Veryfi Extraction Benchmark](https://quirkyagents.com/wizard/projects/pdf-excel-ai/index.md): GPT-4o against Veryfi on irregular financial scans, across several document types — where each one wins, and what each extraction costs. (2025-02-12)
