datum

Insights · 2026-08-01 · 6 min read

How we benchmark an AI claims engine against court judgments

Every AI vendor claims accuracy. Almost none can prove it without asking you to trust their own grading. We wanted a harder test: an answer key written by someone with no stake in our product looking good.

Published court judgments turn out to be exactly that. In a construction delay case like Walter Lilly v Mackay [2012] EWHC 1773 (TCC), the judge spends hundreds of paragraphs establishing what happened and when: possession on 12 July 2004, a 23-week lead time flagged in November 2005, the instruction that came too late. That chronology is adjudicated ground truth - and nobody at datum wrote it.

The method has four steps. First, we reconstruct the project record strictly from the judgment's findings of fact: the letters, instructions, notices and minutes the court describes, re-created with the court's dates, parties and quoted words, every page marked as a reconstruction. Second, the engine ingests that record blind - OCR, classification, indexing, event extraction - exactly as it would a client's record. It never sees the judgment. Third, we score strictly: a court-established event counts as found only if the engine surfaced it on the right date with verbatim evidence carrying the operative phrase. Fourth, we publish - including the low scores, with their reasons.

The current board: eleven judgments, 153 documents, 100% document dating, 86% of court-established events found on the strictest matching. The hardest document shapes - NCR chains and claim letters recounting months of history - are stated on the page rather than hidden.

Two honest limits. Reconstruction corpora are real-shaped, not real: they rehearse the engine, they do not replace pilot validation on a genuinely real record with a practitioner's review. And a 12-event answer key means a single miss moves a score by eight points - which is why we publish the per-case detail, not just the average.

The full board, every number reproducible, is at datumclaims.com/benchmarks.

See it on your own matter. Pilots run on a real record, with your reviewers in command.

Request a pilot