Field note · 4 min read

    Test the tool, not just the data

    An answer key says the values are right. It doesn't say the search an agent sends still answers right. Scenarios check that, and hold back a release that would break one.

    4 min read3 sectionsWritten from shipped code

    01The gap between right data and a right answer

    An answer key checks that values are correct: this dish has dairy, that one doesn't. But an agent never reads values one by one. It sends a search, "nut-free mains", and gets back whatever the exclusions, the filters and the ranking make of the data.

    A release can have every value right and still answer that search wrong: a dish withheld by mistake, a field renamed, a filter that no longer matches. Nothing about the values says so.

    02A scenario is a search with expectations

    A scenario is the request an agent would send, saved with what must hold about the answer:

    • Records it must never return: the kaju katli, for a nut-free search.
    • Records it must return: the dal tadka, which is nut-free and should be offered.
    • Optionally, at least so many results, so an exclusion that suddenly empties the menu is caught too.

    Every page of the answer is read, so a forbidden record can't hide past the first one.

    03A release that fails waits

    Once a tool has a scenario, a new release reaches it only after passing them all. One that fails is held back: the tool keeps answering from the last release that passed, so agents never see the broken one, and the owners get an email saying which scenario failed and why.

    When the fix is in the scenario rather than the data, a person checks the held release again and it goes through. Adding the first scenario never takes a tool offline: releases published before it keep being served.

    Why we took the names out of our embeddings

    Similar-sounding exercises kept ranking next to each other. The fix was to stop letting names speak for meaning.

    Read next