Tracefield

Every value proves where it came from.
Or it says so.

Tracefield shows you, on the page, exactly where every number came from — and flags the ones it can't prove.

Proof, not promises.

Every extraction tool claims accuracy. The question your auditor will actually ask is: how do you know this number is real? With Tracefield, the answer is one click.

01EXTRACT

It reads the document, not the vibes.

Every value is pulled from the words on the page — and the model must hand back the exact words it read, so there is nothing to take on trust.

  • OCR word stream with coordinates for every word
  • The model returns each value with its verbatim quote
  • Quotes are matched back against the page, not assumed
  • A quote that isn't on the page fails the field
app.tracefield.io
02INVESTIGATE

Red claims get a detective.

When a value fails verification, the Review Agent works the case before a human sees it — searches the word stream, reads the page around the label, cross-checks the document's own arithmetic, ranks the reference data — then recommends.

  • A bounded tool loop — every call lands in a persisted trace
  • One agent, two engines: raw loop and LangGraph, evaluated 4/4 each
  • Spend is checked before every model call, not reported after
  • It recommends; a person decides, under their own name
app.tracefield.io
03REVIEW

What it can't prove, it won't pass.

A value with no source on the page is routed to your review queue — clearly marked — instead of flowing silently into your ERP. That is the whole product.

  • Unverified never auto-passes — the database refuses the row
  • A person approves, corrects, or rejects, by name
  • Every decision lands in the audit trail
  • Corrections feed back into future extractions
app.tracefield.io
04EXPORT

Only what it can prove leaves the building.

The spreadsheet your team actually works from is filled by verified values and human decisions — and by nothing else. A cell nobody could stand behind exports empty, with the reason beside it.

  • A cell is an auto-verified claim or a person's decision
  • Withheld cells are counted, named, and announced in the response
  • Fields marked redact never leave, in any format
  • CSV, JSON, or straight into an agent over MCP
app.tracefield.io

Measured, not asserted.

A tool that is 95% right but can't tell you which 5% is wrong makes you check everything by hand. What matters is whether it knows which answers to stand behind — so that is what we measure.

0
Wrong answers vouched for
across every test run — the number the whole system is built to keep at zero
100%
Right when it says verified
measured against documents designed to fool it
1 click
From any value to its proof
highlighted on the original document
Minutes
To your first verified document
upload, extract, click through the evidence
the real product · sample data

The rest of the workday.

A verification engine nobody can run all day is a demo. These are the parts that make it a tool: batches, corrections, catalogues, and a way out that is not a screenshot.

Your document types, your fields

Name a type, define what to pull from it, and the fields become the question every extraction answers. A type with no fields cannot be uploaded against — the product refuses to extract nothing and call it success.

Batch upload

Drop in a folder's worth. Three run at a time, each reports its own outcome, and 'seven of nine worked' is an answer — a single spinner over nine files is not.

Review at keyboard speed

j and k walk the fields, a approves, e opens the correction, r marks it absent — and a decision moves you to the next thing waiting. The keys can do exactly what the buttons can do, and nothing more.

Reference catalogues

Upload your supplier list or chart of accounts and a value found on the page can be corroborated against what you already know. Corroborated — never substituted: a match is not evidence that the words were there.

Change the schema, ask again

Add a field in month three and the documents from month one can be run again from their stored originals — priced in the sentence above the button, with every previous answer and review left intact.

Agents, over MCP

Connect Claude and it reads each field with the page's own words behind it. There is no tool for recording a review: an agent can investigate and recommend, but the decision carries a person's name.

Simple pricing, honest meter.

Priced per document processed. Every plan includes the full verification engine — we don't sell trust as an add-on.EARLY ACCESS — free plan only for now

Starter
$0/ month

For trying it on your own documents.

  • 50 documents / month
  • 1 document schema
  • Click-to-proof review screen
  • Community support
Start free
MOST POPULAR
Growth
$49/ month

For teams processing invoices every day.

  • 1,000 documents / month
  • Unlimited schemas
  • Review queue with audit trail
  • Signed webhooks & API access
  • Priority support
Start 14-day trial
Scale
Custom

For volume, compliance, and control.

  • Volume document pricing
  • SSO & role-based access
  • Self-hosting option
  • SLA & dedicated support
Talk to us

EARLY ACCESS — EVERY PLAN IS FREE WHILE WE ONBOARD DESIGN PARTNERS. BILLING (VIA PADDLE) ARRIVES WITH GA.

FOR THE ENGINEER YOUR TEAM SENDS TO CHECK OUR CLAIMS

Built to be verified, not believed.

The invariant is enforced by the database, not by good intentions — a verified verdict is unreachable without evidence, by constraint. The runtime is MIT-licensed, the eval suite gates every merge, and seven acceptance proofs kill workers mid-run to check the claims that only matter when demonstrated.

constraint passed_requires_evidence
check (verdict <> 'passed'
or evidence_type is not null)
$ npm run quickstart  # 23s, measured, no API key

Our own CI lied to us once

The pipeline printed success after a step failed — the exact failure this product exists to prevent, inside the tool built to prevent it. The decision log records it, because the bugs that matter produce plausible output instead of errors.

The baseline gate was wrong first

A gate assumed an extractor that says nothing scores zero. It scores exactly the share of fields that are genuinely absent. It failed a working system, and the fix was to compute the floor instead of assuming it.

The data flywheel lost, on purpose

Retrieving past corrections as few-shot examples moved nothing: +0.0 points, every category — measured on a held-out split with a leakage guard that fails the suite rather than inflating it. Negative results get published here.

BUILT TO BE CHECKED

Stop spot-checking.
Start tracing.

Fifty documents free, no card. Every field proves itself — or says so.