# the wizard

dave8172's open source, live AI tools and writing, on Quirky Agents (https://quirkyagents.com/).

## Spellbook (open source)

Open source: code you can read, run and check.

- [doceval](https://quirkyagents.com/wizard/projects/doceval/index.md): Point it at your extractor and a labeled dataset. Get field-level accuracy, a failure taxonomy and per-document cost, without writing eval infrastructure.
- [tally-aiagent](https://quirkyagents.com/wizard/projects/tally-aiagent/index.md): Tally invents any master name it doesn't recognise, then reports success. This refuses names the books don't hold, and verifies every entry after posting.
- [upsweep](https://quirkyagents.com/wizard/projects/upsweep/index.md): Say "upsweep" to your AI agent. It lists the Upwork jobs that pass your filters and sums up what those clients keep asking for.

## Wands (live tools)

Live tools: wave one and see what it does.

- [affterms](https://quirkyagents.com/wizard/projects/affterms/index.md): Affiliate terms for 471 dev and SaaS tools, each figure sourced and dated. A calibrated model says which pages changed; a person reads them before any figure moves.
- [Magic Jev Ball](https://quirkyagents.com/wizard/projects/jevball/index.md): A Magic 8 Ball that reads the question first. The verdict is a calibrated probability from Jev, not a random pick from twenty phrases — and it shows its working.
- [waif](https://quirkyagents.com/wizard/projects/waif/index.md): What am I feeling? Text in, a feeling and an intent out — named in two stages by a calibrated model, with the eval that picked the design printed on the page.

## Memories (writing and benchmarks)

Writing and benchmarks from real builds: what broke, what it cost, and the number that settled it.

- [Building and using MCP servers: 4 things I learned](https://quirkyagents.com/wizard/blog/building-mcp-servers-lessons/index.md): How MCP works, one tool call step by step, and four things I learned building and using MCP servers, from tool descriptions to where secrets live.
- [Local vs remote MCP servers: how each connects, step by step](https://quirkyagents.com/wizard/blog/local-vs-remote-mcp-servers/index.md): How a local MCP server lives and dies with your session, and how remote MCP sign-in works: OAuth, PKCE and tokens, tested on Upwork, Notion and Todoist.
- [How I calibrated an LLM judge to grade like me, 25× cheaper](https://quirkyagents.com/wizard/blog/calibrate-llm-judge-evals/index.md): Evals on 188 answers from ChatGPT, Claude, Chatbase and my own pipeline over product datasheets, and a hybrid judge with Jev that grades like I do.
- [Can an AI agent submit Upwork proposals? Yes, with your OK](https://quirkyagents.com/wizard/blog/can-ai-agent-submit-upwork-proposals/index.md): Upwork's official MCP lets Claude, Codex or Cursor send a proposal for you. What it shows before sending, what stays on upwork.com, and what I got wrong.
- [Ask twice: seven measurements from building with Jev](https://quirkyagents.com/wizard/blog/jev-ask-twice/index.md): A coordinate lookup scored 3/24 where handing the model the naming job scored 22/24. Then two claims I had argued rather than measured turned out to be wrong.
- [Choice, Score and Noul: four mistakes with Jev's primitives](https://quirkyagents.com/wizard/blog/jev-primitives-four-mistakes/index.md): I built a Magic 8 Ball on Jev, TypeSafe's calibrated-decision model, and misread its three primitives four times. The measured numbers, and what each mistake cost.
- [AI turns emails into spreadsheet rows. Did it get them right?](https://quirkyagents.com/wizard/blog/measure-ai-reading-inbound-emails/index.md): Eight real inbound emails through Claude Haiku: 61 of 64 fields correct, every miss on one field. What a failure taxonomy shows that an accuracy score hides.
- [98.3% vs 96.3%: the cost of an auditable AI extraction pipeline](https://quirkyagents.com/wizard/blog/invoice-extraction-cost-accuracy-benchmark/index.md): 105 labeled invoices and receipts: Claude Opus vs Haiku+Sonnet, a 2% accuracy gap and a 12× cost gap, plus the eval set and audit log behind the numbers.
- [Self-Hosted Whisper Benchmark](https://quirkyagents.com/wizard/projects/whisper-transcription-benchmark/index.md): A self-hosted meeting-transcription pipeline, plus a benchmark of Whisper model sizes on a real 32-minute recording to find the accuracy/speed sweet spot.
- [GPT-4o vs. Veryfi Extraction Benchmark](https://quirkyagents.com/wizard/projects/pdf-excel-ai/index.md): GPT-4o against Veryfi on irregular financial scans, across several document types — where each one wins, and what each extraction costs.
