01Why an agent's answer needs a trail
When an AI agent recommends a product, someone will eventually ask why: a customer who had a reaction, a support team chasing a complaint, a buyer doing a security review. If the only record is that a model decided, nobody can check the decision, fix the cause, or show it was reasonable.
Provenance answers that question for every value, and for safety-critical product data it is not optional: where it came from, how it was decided, and who stood behind it.
02What a useful receipt holds
- The method: a rule that matched a known word, a model, or a person.
- The evidence: the exact words relied on, and the document and page when there is one.
- The confidence, and the threshold the field required.
- The version of the field's rules it was decided under.
- The person who approved it, when one did.
Each of these answers a different question later. The evidence shows whether the source said it. The version shows whether a later rule change would decide it differently. The person shows who made the call.
03Check the quote, not just the answer
A model asked to cite its source will sometimes cite words that aren't there. So a quote should be checked before a value is stored: the words must appear in the source, and they must actually say the value. A number read from a spec sheet has to be found in the spec sheet.
When the check fails, the value isn't stored. It goes to a person instead, which is the right outcome for any answer whose evidence can't be found.
04Name the people, and lock what they decide
Some values will always need a person: the dish the model was unsure of, the allergen a re-import tried to remove, the record where two sources disagree. When a person decides, the receipt should name them, and their decision should stay put. A later automated run should not quietly overwrite it.
05Freeze what the agent saw
Data changes after an answer is given. To answer "what did the agent see on Tuesday?", the data an agent reads should be published as releases that never change, each with its receipts. Then a complaint about Tuesday's answer can be checked against Tuesday's data, not today's.
06How CloudCrane does it
Every value CloudCrane stores has a receipt: method, evidence with its page, confidence, the contract version and who approved it. A fact read from text, like a number or a date, is stored only once its quote is found in the source and says the value. Hand decisions lock, and an agent reads from an immutable release, so any answer can be traced back to the exact data and decision behind it. How each part is enforced is on security and trust.
07Questions
- What is data provenance for AI?
- A record of where each value an AI system uses came from: its source, how it was decided, how confident that decision was, and who approved it.
- Isn't the model's own explanation enough?
- No. A model's explanation can't be checked against anything, and a model asked for a source will sometimes cite words that aren't there. Check the quote against the source before storing the value.
- What is an immutable release?
- A published snapshot of the data an agent reads, with its receipts, that never changes afterwards. A new version is a new release, so any past answer can be checked against the data it was given.