BENCHMARKS

How we measure
ourselves.

We benchmark Kendraa against an established review platform on a public corpus. This page is the method. The numbers are a conversation away.

Why benchmark at all

Two platforms can quietly disagree about how many documents exist. You want to find that before a hearing, not during one.

The corpus

The public Enron email corpus
Real corporate mail, in the public domain, obtainable by anyone who wants to repeat the run.
Whole mailboxes, per custodian
As published, per custodian.

What we measure

The processing funnel

Every stage from container to reviewable document, compared stage by stage rather than only on the end total.

Deduplication behaviour

What collapses across custodians, and whether attachments are kept and flagged rather than removed.

Search semantics

Proximity, noise words, hyphenation — where tokenisation returns a different set for the same words.

Chat and mobile messages

Reviewable units, conversation continuity, and how deleted-message provenance is represented on each side.

How a run works

Same input
The identical published mailbox goes into both systems.
Item-level bridging
Each document is matched to its counterpart, so a difference traces to items.
Classified differences
Design difference, accepted divergence, or defect. Defects get fixed.
Repeatable
Scripted and re-run when the pipeline changes, so a regression shows up.

Ask us for the data

Want it run on your own file types and languages? Ask.

We will walk your team through a full run, stage by stage. Our security posture covers where it all runs.

Request the benchmark data