Measuring accuracy with an answer key
An answer key is a set of records whose right values a person has confirmed, one per dataset. Build one by promoting values people already decided in review or by hand, by labelling records, or from a CSV. A second person can confirm a safety label; the one who labelled it can't. A label whose record text has changed since is marked stale and left out of scores. Publishing a key freezes it, so “accuracy on key version 2” always means the same thing; the next version starts from it.
- 1
Measured, not guessed
An evaluation runs the real pipeline over the key and writes nothing. Each value gets its recall with a range, and below 30 labelled records a value says it is too few to judge rather than printing a confident number. Missed safety values come first.
- 2
Measured on every change
Saving a contract, publishing a key or changing the model measures today's fields again, automatically.
- 3
Compared record by record
Two evaluations side by side: what broke, what got fixed, what changed. A safety value found before and not now leads the list.
- 4
Checked before publishing
With a published key, a release is refused if a safety value would be found less often than in the last release, or if today's fields haven't been measured, unless a person gives a reason, which the release keeps. Only safety values stop a release; overall accuracy is shown, never gating.
Callers are told. A release keeps how it measured, and the deployed tool's description says so, so an agent can tell its user how far to trust an answer.