Field note · 6 min read

    An answer key for AI enrichment

    A prompt, model or rule change can make a catalog better overall and worse exactly where it matters. How a golden set catches the release that loses a safety value.

    6 min read5 sectionsWritten from shipped code

    01Overall accuracy hides the miss

    Change a prompt and overall accuracy goes from 97% to 98%. It looks like progress. But the same change might find dairy in fewer dishes than before, and no overall number will say so: a few lost safety values disappear inside thousands of correct ordinary ones.

    So the question to ask of every change is narrower: did any safety value get found less often?

    02The answer key

    A golden set is an answer key for one dataset: records and fields whose values a person confirmed, including a confirmed "none". Once published it is frozen, so "accuracy on golden v2" always means the same thing. A new version is a new set.

    • A label on a safety field needs a second person to confirm it.
    • When a record's text changes after it was labelled, the item is marked stale and left out of scores until someone checks it again.
    • Items can come from values people already confirmed in review, from records chosen and labelled, or from a CSV of known values.

    03Evaluate without writing anything

    An evaluation is a dry run of a candidate, the contracts and the model as they would be, over the answer key. It writes nothing to records, receipts or the review queue, and it decides each item as if the record had no value yet, so a hand edit can't make a contract look better than it is.

    Scores are per value, with recall front and centre. An item the candidate didn't answer still counts: its expected safety values are counted as missed.

    04Compare item by item

    Two evaluations are compared record by record: what broke (right before, wrong now), what was fixed, what changed while still wrong. A lost safety value, one given before and not now, always leads the list, with the text and the rule hits behind both answers.

    05The gate on every release

    With a published answer key, publishing a release is refused if today's contracts and model haven't been evaluated, or if any safety value's recall falls below the last stamped release. A person can still publish, but only by giving a reason, and the release keeps it.

    Every release carries its accuracy stamp, and the served tool's description includes it, so an agent's builder can see what the data was measured against. It is one of the controls on security and trust.

    Bring your own model, without trusting it

    Teams want their own model for cost, contracts or data residency. What stays enforced whichever model answers, and the one thing we disclose instead of solving.

    Read next