How accurate is it?
Everyone in this field promises reliable answers. Almost nobody publishes a number. Here is ours — with the setup, the failures, and what it does not prove.
The result
One run over 149 probes against a controlled evidence base. Every probe carries a predetermined expected behaviour: answer where the documents support it, abstain where they do not.
Run date: 14 August 2026. Thresholds: minimum evidence similarity 0.6, retrieval match threshold 0.5.
How it was measured
The setup decides whether a number like this means anything. So it is here in full.
The eleven failures
What matters more than the rate is the direction a system fails in. Ours fails upward: it abstains too often rather than inventing.
10 abstentions
The answer was in the documents and the system still put it in front of a human. That costs time and is irritating — but it exposes nobody to a false statement towards their customer.
1 wrong content
An answer supported by the corpus but wrong on the substance. That is the failure that actually hurts, which is why it is listed on its own rather than folded into a remainder.
0 fabrications
The failure this product cannot afford did not occur — across all 38 cases where inventing would have been possible.
What this number does not prove
A measurement without stated limits is advertising. This one has four.
The corpus is constructed. It is internally consistent, cleanly formatted and thematically controlled — real company documentation is none of those things.
It is one company, one sector, one size. A corpus with contradictory legacy documents, or three generations of policy, behaves differently.
The number describes the state on 14 August 2026. Any change to retrieval, model or thresholds moves it, and we have watched it move: the same corpus gave 89.9 per cent two days earlier.
What was measured is draft quality, not the outcome of a review. Whether a buyer accepts your answers rests on your evidence, not on our hit rate.
Reproducibility
The corpus, the fact sheet and all 149 probes live in the product's source; the run is a single command. Anyone who doubts the number can recompute it — that is the point of publishing one.
Put this process on autopilot
Upload your evidence and your next questionnaire — get cited drafts to review, not answers to blindly trust. Free for your first questionnaire.