# doceval — Extraction Eval Harness

Open-source eval harness for LLM extraction. Point it at your extractor and a labeled dataset: field-level accuracy, a failure taxonomy and cost per document.

Open source, by dave8172. Page: https://quirkyagents.com/wizard/projects/doceval/

## Overview

**In plain terms:** if you're using AI to pull data out of invoices, contracts, forms — or emails, web pages and call transcripts — doceval measures how accurate it is. A number you can show a client, an auditor, or your own team, instead of "it looks about right."

For the technical side: you've built an LLM-based extractor. It seems to work. But how accurate is it? Which fields fail most — and why? Did accuracy change when you updated the prompt?

Without answers, "seems to work" is all you have. That's not good enough for production.

**doceval** is an open-source eval harness that gives you those answers. Point it at your extraction function and a labeled dataset, and it produces:

- **Field-level accuracy** — precision per field, not just overall
- **Failure taxonomy** — every mismatch classified as `missed_field`, `hallucination`, `wrong_format`, or `wrong_value`
- **Cost tracking** — optional per-document cost reporting when your extractor returns it
- A shareable Markdown report

Works with any extractor (Claude, GPT, regex, rules-based) and any schema. It also works on any *input* — what it scores is whether the extracted fields are right, which never depended on the source being a document. PDFs and scans, yes; equally emails, HTML pages, transcripts or feed dumps.

→ [View on GitHub](https://github.com/dave8172/doceval)

---

## How It Works

Write an extractor — a Python function that takes `(doc_bytes, filepath)` and returns a dict:

```python
def extract(doc_bytes: bytes, filepath: str) -> dict:
    # call Claude, GPT, or any extraction logic
    return {"vendor": "Acme", "total": "1234.56", "date": "2026-01-15"}
```

Add one JSON label file per document:

```json
{
  "vendor": "Acme Corp",
  "total": "1234.56",
  "date": "2026-01-15"
}
```

Run the eval:

```bash
pip install doceval

doceval run \
  --docs    ./dataset/docs \
  --labels  ./dataset/labels \
  --extractor my_module:extract
```

---

## Failure Mode Taxonomy

Every mismatch is classified into one of four modes:

| Mode | Meaning |
|------|---------|
| `missed_field` | Label has a value; extractor returned empty |
| `hallucination` | Extractor returned a value; label is empty |
| `wrong_format` | Both non-empty; numeric or date values differ |
| `wrong_value` | Both non-empty; string values differ |

doceval handles numeric normalisation (`$1,234.56` = `1234.56` = `1.234,56`) and date normalisation (`Nov 15 2012` = `2012-11-15`) before comparison, so format differences don't inflate your error count.

---

## Optional Cost Tracking

Return a `(dict, cost_usd)` tuple from your extractor and doceval tracks cost automatically:

```python
def extract(doc_bytes: bytes, filepath: str) -> tuple[dict, float]:
    response = client.messages.create(...)
    cost = response.usage.input_tokens / 1e6 * 0.80
    return result_dict, cost
```

---

## Try the Example

The repo includes a working 20-document invoice dataset with labels and a Claude Haiku extractor you can run immediately:

```bash
git clone https://github.com/dave8172/doceval
cd doceval
pip install -e ".[examples]"
export ANTHROPIC_API_KEY=sk-ant-...

doceval run \
  --docs    examples/invoices/docs \
  --labels  examples/invoices/labels \
  --extractor examples.invoices.extractor:extract
```

### The same harness on something that isn't a document

`examples/leads/` runs over eight inbound sales emails — a forwarded thread where the lead is the original sender rather than the forwarder, a terse RFQ, a web-form dump, an automated tender notice with no contact details, and one enquiry in Italian.

```bash
doceval run \
  --docs    examples/leads/docs \
  --labels  examples/leads/labels \
  --extractor examples.leads.extractor:extract
```

Three of those emails carry null-heavy labels on purpose. An extractor that invents a quantity or a budget for them scores *worse*, not better — which is the point of measuring fields rather than output shape.

---

## MCP Server — Use It From an AI Agent

doceval also ships as an [MCP](https://modelcontextprotocol.io) server, so an AI coding agent (Claude Code, Claude Desktop, Cursor) can score extraction accuracy directly, mid-conversation, without shelling out to the CLI.

```bash
pip install "doceval[mcp]"
```

Two tools:

- **`score_extraction`** — score one extraction result against expected values inline, no filesystem needed. What an agent reaches for mid-development: "here's what I extracted, here's what it should be, how'd I do?"
- **`run_eval`** — run a full eval over a docs/labels dataset on disk and get back the structured result plus a rendered Markdown report, same as the CLI.

Point any MCP client at the `doceval-mcp` command over stdio:

```json
{
  "mcpServers": {
    "doceval": { "command": "doceval-mcp" }
  }
}
```

---

## Tech Stack

- **Language:** Python 3.10+
- **CLI:** Click
- **Agent integration:** MCP server (`doceval-mcp`)
- **Packaging:** Hatchling / pyproject.toml
- **Example extractor:** Anthropic Claude Haiku
- **Supported inputs:** documents — PDF, PNG, JPG, JPEG, TIFF, WEBP · text — TXT, MD, JSON, JSONL, NDJSON, CSV, TSV, HTML, HTM, XML, EML, MSG, VTT, SRT
