Measurement

How accurate is it?

Everyone in this field promises reliable answers. Almost nobody publishes a number. Here is ours — with the setup, the failures, and what it does not prove.

The result

One run over 149 probes against a controlled evidence base. Every probe carries a predetermined expected behaviour: answer where the documents support it, abstain where they do not.

Probes149111 expect an answer, 38 expect an abstention.
Correct13892.6 per cent.
Fabricated answers0Not one answer without support in the corpus — across all 38 abstention cases.
Wrong content1One answer with support in the corpus but a wrong statement.
Abstained although answerable10The documents held the answer; the system put it in front of a human instead of writing it.
Citations correct101 / 101Wherever a source attribution could be checked, it pointed at the right passage.

Run date: 14 August 2026. Thresholds: minimum evidence similarity 0.6, retrieval match threshold 0.5.

How it was measured

The setup decides whether a number like this means anything. So it is here in full.

The evidence base is invented: eleven policies and reports belonging to a fictitious Nordwind Cloud GmbH. A real customer corpus could be neither shared nor checked.
A canonical fact sheet is the single source of truth. No document may contradict it or introduce a fact it does not carry — otherwise the test drifts into measuring the corpus rather than the system.
Six topics are deliberately absent, among them an ISO 27001 certification of its own and recovery objectives. For those topics the fact sheet carries a list of forbidden words that may appear nowhere.
Adjacent material is explicitly wanted: penetration testing next to the missing bug bounty, backups next to the missing recovery objectives. That adjacency is what makes an abstention hard — and therefore worth measuring.
Every probe carries its expected behaviour and, where an answer is expected, the strings that must appear in it. Scoring is mechanical, not impressionistic.

The eleven failures

What matters more than the rate is the direction a system fails in. Ours fails upward: it abstains too often rather than inventing.

10 abstentions

The answer was in the documents and the system still put it in front of a human. That costs time and is irritating — but it exposes nobody to a false statement towards their customer.

1 wrong content

An answer supported by the corpus but wrong on the substance. That is the failure that actually hurts, which is why it is listed on its own rather than folded into a remainder.

0 fabrications

The failure this product cannot afford did not occur — across all 38 cases where inventing would have been possible.

What this number does not prove

A measurement without stated limits is advertising. This one has four.

The corpus is constructed. It is internally consistent, cleanly formatted and thematically controlled — real company documentation is none of those things.

It is one company, one sector, one size. A corpus with contradictory legacy documents, or three generations of policy, behaves differently.

The number describes the state on 14 August 2026. Any change to retrieval, model or thresholds moves it, and we have watched it move: the same corpus gave 89.9 per cent two days earlier.

What was measured is draft quality, not the outcome of a review. Whether a buyer accepts your answers rests on your evidence, not on our hit rate.

Reproducibility

The corpus, the fact sheet and all 149 probes live in the product's source; the run is a single command. Anyone who doubts the number can recompute it — that is the point of publishing one.

Put this process on autopilot

Upload your evidence and your next questionnaire — get cited drafts to review, not answers to blindly trust. Free for your first questionnaire.