Measuring accuracy with an answer key

    An answer key is a set of records whose right values a person has confirmed, one per dataset. Build one by promoting values people already decided in review or by hand, by labelling records, or from a CSV. A second person can confirm a safety label; the one who labelled it can't. A label whose record text has changed since is marked stale and left out of scores. Publishing a key freezes it, so “accuracy on key version 2” always means the same thing; the next version starts from it.

    1. 1

      Measured, not guessed

      An evaluation runs the real pipeline over the key and writes nothing. Each value gets its recall with a range, and below 30 labelled records a value says it is too few to judge rather than printing a confident number. Missed safety values come first.

    2. 2

      Measured on every change

      Saving a contract, publishing a key or changing the model measures today's fields again, automatically.

    3. 3

      Compared record by record

      Two evaluations side by side: what broke, what got fixed, what changed. A safety value found before and not now leads the list.

    4. 4

      Checked before publishing

      With a published key, a release is refused if a safety value would be found less often than in the last release, or if today's fields haven't been measured, unless a person gives a reason, which the release keeps. Only safety values stop a release; overall accuracy is shown, never gating.

    Callers are told. A release keeps how it measured, and the deployed tool's description says so, so an agent can tell its user how far to trust an answer.