Most tools that answer questions about your data will not tell you how often they are wrong. This page does, on questions the system had never seen — with every question published, including the ones it failed.
On 60 questions written and sealed before the code that answers them was changed:
| Measure | Result |
|---|---|
| Answered | 35% |
| Refused | 65% |
| Correct, when it answered | 100% |
| Wrong, when it answered | 0 |
| Refused something it could have answered | 1 |
It refuses about two questions in three. That is the product working, not failing. Most of what it declines are questions it genuinely cannot answer correctly — a metric nobody has defined, a breakdown the data does not support, a period it has no history for — and the alternative to refusing is a plausible number in your board deck.
A perfect score on questions you tuned against is a statement about your own prompt engineering. Ours was 0% wrong on the development set and 19.2% wrong the first time it met questions it had never seen. Every round is here:
| Question set | Wrong when answered |
|---|---|
| Set 1 — development set, tuned against | 0% (meaningless) |
| Set 2 — first unseen set | 19.2% |
| Set 3 — sealed before the time fix | 13.8% |
| Set 4 — sealed before the latest fixes | 0% |
Two real defects were found this way. A question about a past period was answered with today's figure — a current number wearing another date. And a breakdown the data could not support was silently dropped, returning an unrelated total at high confidence. Both are now refused, and each fix was written only after the next set of questions had been sealed, so it could not be tuned to the test.
Precision matters more than the number. Four things this measurement does not say:
A number without the questions cannot be checked. Here they are, with what each was expected to do.
Written after the tuned set, scored once. This is the set that found 8.7% wrong answers, including a year-to-date question answered with today's figure.
Written and committed before the fix existed, so the fix could not be tuned to it. Contains a present-tense bucket specifically to catch the fix over-firing.
The current headline number comes from this set.
Kept for completeness and NOT used for any claim. The planner was tuned against these three times, so its perfect score means nothing.
Measured 8 September 2026 on the demo dataset. Questions: hello@truacta.ai · See also Trust & security.