How Mentis is evaluated
Evidence before benchmark theatre.
Diligence software should be measured on whether it recovers the facts that matter and keeps them faithful to the source. This is the methodology Mentis uses to test document extraction and question answering.
Mentis does not publish a headline accuracy claim from a small or sensitive corpus. Public results will follow after the anonymized set is large enough, the answer key has been independently sampled, and limitations can be reported alongside the score.
The evaluation chain
From source document to defensible result
Start with representative diligence documents
The evaluation set spans the formats that make real data rooms difficult: pitch decks, financial models, cap tables, contracts, and scanned files.
Define atomic, verifiable facts
Each document is reviewed into a frozen set of critical facts: a number, date, entity, table cell, or clause that a useful extraction must preserve.
Test the system customers actually use
Mentis runs through the same parsing, OCR, retrieval, and answer pipeline used in the product rather than a simplified benchmark implementation.
Measure recall and faithfulness
Recall asks whether critical facts were recovered. Precision checks whether extracted claims remain faithful to the source instead of becoming garbled or unsupported.
Inspect failures before publishing a number
A score is only useful when the ground truth, judging process, limitations, and failure cases can withstand customer scrutiny.
Publication standard
What must be true before a result goes public
- Source documents are anonymized before entering the benchmark corpus.
- A human reviews the ground-truth facts; model-generated candidates are not accepted blindly.
- The same task and scoring rules apply to every system under comparison.
- Sample size, judge model, failure cases, and known limitations are disclosed.
See the product workflow