Guide · 7 min read

    AI data for safety-critical products: unknown is never safe

    When an AI agent answers "is this nut-free?" or "is this safe in pregnancy?", the data behind it decides whether someone gets hurt. Five rules for building data an agent can be trusted with.

    7 min read7 sectionsWritten from shipped code

    01What makes product data safety-critical

    Safety-critical usually brings to mind flight software or medical devices, with their own standards and certification. This is a different and much more common case: ordinary product data that an AI agent now reads aloud to people who act on it. A menu's allergens, a medicine's pregnancy advice, a hotel's step-free access, an exercise's load on a bad knee.

    None of that data was written for an agent. It was written for a person who would read the whole page, notice what was missing, and ask. An agent does none of that. It is asked a yes-or-no question and gives a confident answer, and a wrong yes is the one mistake that matters.

    • A wrong "yes, it's nut-free" can put someone in hospital.
    • A wrong "no" only loses a sale.
    • So the two errors are not equal, and the data has to be built for the first one.

    02Rule 1: unknown is never safe

    Most search and filtering treats missing data as a pass. Ask for dishes without dairy, and every dish whose dairy field is empty comes back, because nothing says it contains dairy. For an agent answering a safety question, that is exactly backwards.

    The rule that works is to fail closed. When a caller excludes a value, a record is returned only if it is known not to have it. A dish nobody has checked for dairy is left out of a dairy-free answer, even though nothing says it is unsafe. The answer gets shorter, and every row in it can be defended.

    03Rule 2: safety values can be added, never quietly removed

    Data drifts. A re-import, a new menu, an edited spreadsheet, and an allergen tag disappears from a dish without anyone deciding it should. On an ordinary field that is a small error. On a safety field it turns a dish that was correctly excluded into one that is served as safe.

    So a safety field should only grow on its own. Adding a value is automatic; removing one waits for a person, who sees what is being removed and why, and whose name goes on the change. We wrote up how that works in a catalog that can't forget an allergen.

    04Rule 3: every value carries a receipt

    When an agent says a product is safe, someone will eventually ask how it knew. "The model said so" is not an answer anyone can check. Every value should carry where it came from, which is what data provenance for AI agents means in practice:

    • How it was decided: a rule matching a known word, a model, or a person.
    • How confident the decision was, and the threshold it had to clear.
    • The exact words it was read from, checked to be in the source, and the page when the source is a document.
    • Which version of the field's rules it was decided under, and who approved it.

    Receipts are also what make the other rules workable: a person reviewing a removed allergen, or an unsure answer, needs to see what the decision was based on.

    05Rule 4: keep the safety check outside the model

    A model is good at reading messy text and bad at being a guarantee. Use it for what code can't do, such as reading "finished with a splash of cream" as dairy, and keep the safety decision itself in code: allowed values the model can't step outside, a confidence threshold below which nothing is stored, and the exclusion enforced in the query. The longer argument is in don't make the model your safety system.

    Rules before models helps twice. Known words settle most records for free and the same way every time, and the model only sees what the rules couldn't settle, which is where its judgement is actually needed.

    06Rule 5: measure before every release

    A change that improves overall accuracy can still lose a safety value somewhere. The only way to know is an answer key: records whose values a person has confirmed, run through the pipeline before a release goes out.

    The number to watch is recall on safety values, not overall accuracy. If a new prompt, model or rule finds fewer of the allergens it found last time, the release should stop and a person should have to say why it is going out anyway.

    07How CloudCrane builds this

    CloudCrane is built on these five rules. Each field has a contract: the values it may take, the words that settle it without AI, and whether it is a safety field. Safety fields fail closed in the query and only grow on their own, and once a golden set is published, every release is gated on their recall against it. Every value has a receipt, and the result is a tool an agent calls over MCP or REST. The controls behind each of these are listed on security and trust.

    08Questions

    What is safety-critical product data?
    Product data where a wrong answer can hurt someone: allergens on a menu, a medicine's advice in pregnancy, step-free access at a hotel, an exercise's load on an injured joint. It is ordinary catalog data, not certified safety software, but an AI agent now answers questions from it.
    Why shouldn't an AI agent treat missing data as safe?
    A missing allergen tag usually means nobody checked, not that the allergen is absent. If unknown counts as safe, a confident "yes, it's nut-free" rests on nothing. Failing closed returns only records known not to have the excluded value.
    Can a prompt make an AI agent safe for allergens?
    Not reliably. A prompt asks the model to be careful; it doesn't stop an unchecked record reaching it. Put the exclusion in the query the agent calls, so unknown records are never in the answer. See don't make the model your safety system.
    How do you measure AI data accuracy on safety fields?
    With a golden set: records a person has confirmed, run through the pipeline before each release. Track recall on safety values rather than overall accuracy, and stop a release that finds fewer of them than the last one.

    Allergen data for AI agents: stop a chatbot saying "nut-free" when it isn't

    Restaurant and grocery chatbots get allergen questions every day, and menu data was never written to answer them. Where the errors come from, and how to structure allergen data an agent can rely on.

    Read next