A decision made with the help of an AI system is an event. Like any event that later comes under scrutiny, it can be examined, questioned, defended, or challenged. The question that matters is not what the system decided. It is whether you can show how the decision was reached, and whether that account holds up when someone competent examines it.
Organizations often preserve the wrong thing. They keep the output. They keep the score, the recommendation, the flag, or the generated text. They do not retain the conditions that produced it. When the decision is questioned months later, the organization may hold a conclusion with no reliable way to reconstruct how it was reached.
This article examines what must be preserved to reconstruct an AI-influenced decision, and why the output alone is not enough for that task.
The preservation record at a glance
A defensible record should preserve, in context:
- the effective instruction given to the system, including hidden instructions and retrieved context where available;
- the source data and retrieval records used at the time;
- the model, version, configuration, tools, and relevant system state;
- the complete, unedited output;
- execution, audit, error, and tool-call logs;
- reliable timestamps that allow the records to be aligned;
- the human review, edits, approval, rejection, or override;
- the decision reached and the action taken; and
- the policies, roles, authorities, and accountability records surrounding the decision.
The list is a starting point, not a universal retention schedule. The required scope depends on the system, the consequence of the decision, applicable law and policy, and the purpose for which the evidence may later be needed.
The decision is the evidentiary unit, not the output
In forensic work, you do not defend a conclusion by restating it. You defend it by showing the artifacts it rests on and the reasoning that connects them. The same principle applies here.
An AI output is a single artifact. It sits inside a longer chain. The chain runs from the instruction given to the system, through the data and context it drew on, the state it was in, and the process it executed, to the output it returned. It does not end there. A person exercises judgment on that output. That judgment produces a decision. The decision leads to an action.
Instruction, context, system state, execution, output, human judgment, decision, action. Organizations routinely collapse the last three into one. The output is treated as the decision, and the decision is treated as the action. Separating them is the first discipline. The output is what the system returned. The decision is what a person concluded from it. The action is what the organization then did. Each is a distinct artifact, and each can fail independently.
The unit that must be preserved is not the output. It is the decision, understood as everything required to reconstruct how the output was produced, how a person acted on it, and what followed.
What reconstruction requires
Reconstruction is the test for the record, not for the decision. If you can rebuild the state of the system and the surrounding process as they existed at the moment of the decision, you have the evidentiary basis to assess that decision: to challenge it, to defend it where the evidence justifies defense, and to identify where it failed. If you cannot rebuild that state, you are asserting, not demonstrating.
A reconstructed decision is not automatically a sound one. You may reconstruct a decision completely and find that it was biased, unlawful, outside policy, or based on defective data. Reconstruction does not confer legitimacy. It confers the ability to examine the decision on evidence rather than assertion. That distinction runs through the rest of this article.
Reconstruction requires the following categories of evidence. Each answers a specific question about the decision.
Prompts and the effective instruction context
The prompt records the instruction or objective presented to the system. Where a human supplied it, it may also provide evidence relevant to that person's intent. It does not, by itself, establish the intent of the system executing it. And it is rarely the whole input.
The evidentiary question is not what the user typed. It is what information and instructions were actually available to the model when it generated the response. That effective instruction context can include several layers: the user's prompt, the system instruction, developer instructions, prior conversation history, retrieved documents, tool descriptions, and stored memory. Any of these can shape the output, and several are hidden from the operator. A record that keeps the user's question and omits the rest preserves a fragment of the input and misrepresents the whole.
Source data
Most systems do not reason in isolation. They draw on data. In a retrieval system, that means the specific documents pulled at the time of the query. In an agent, it means the tools called and what those tools returned.
The output is a function of the data available at that moment. Data changes. Records are updated, corrected, or deleted. The version of the data the system used is what matters, not the version that exists when the decision is later reviewed. Preserving current data and treating it as the source is a common and serious error.
In a retrieval system, preserving the documents is a start, not the whole. The retrieval metadata matters as well: the query issued, the documents ranked and returned, the chunks passed to the model, the retrieval scores, any filtering applied, and the version of the index or embedding model in use. Two retrieval systems drawing on the same repository can hand the model different evidence. Without the retrieval metadata, you know what was in the repository but not what the model was shown.
Model version and configuration
The phrase "the model" is misleading. Models change. Weights are updated, system prompts are revised, safety filters are adjusted, and parameters are retuned. A decision produced under one version is not comparable to one produced under another.
Configuration is part of the system's state. Temperature, token limits, tool access, thresholds, and guardrails all shape the output. Preserving the model name without the version and configuration is like citing a document without a date or an edition. It names something, but not the thing that acted.
One caution about what this preservation buys you. Even with the same model, prompt, configuration, and data, some generative systems will not produce the identical output twice. The objective is not deterministic reproduction. It is reconstruction. You preserve the version and configuration so you can rebuild the conditions the decision was made under, not so you can rerun the system and expect the same tokens.
The output
Preserve the complete output, unedited. Not the summary. Not the portion that was acted on. The full response, including anything the system hedged, qualified, or flagged.
This matters because outputs often carry caveats that get stripped when a human acts on them. A model may return a recommendation with stated uncertainty. If the record keeps only the recommendation, it erases the uncertainty and makes the decision look more confident than the system was.
Logs
Logs are the process record. They establish what the system did and in what order: the sequence of tool calls, the retries, the errors, and the intermediate steps where they are captured. For agentic systems this is essential.
A caution belongs here. With generative models, logs establish the observable execution path. They do not establish the model's internal reasoning. You can record what the system received, what it called, and what it returned. You generally cannot recover the internal computation that produced each token, particularly with a proprietary foundation model. This does not weaken the record. It defines what the record is. Logs let you reconstruct the operational sequence and context. They do not open the model's hidden computation, and an account that claims otherwise overstates the evidence.
Even so, logs establish the observable chain between input and output. Without them, you hold two endpoints and no path. Digital-forensic guidance has long treated event, audit, error, installation, and other application logs as potential sources for reconstructing what occurred.
Timestamps
Timestamps record when each thing happened. Their value is not the time itself. It is that timing lets you align the decision with the state of everything else at that moment.
A timestamp connects the output to the model version deployed then, the data available then, and the person on duty then. Timing is the mechanism that lets you reconstruct state. Without reliable timestamps, the other artifacts float free of each other and cannot be assembled into a sequence.
Human interventions
Somewhere in most processes, a person enters the loop. The record must show what that person actually did. Did they review the output, edit it, override it, approve it, or pass it through without examination?
This is where many accounts fail. "Human in the loop" is a claim about design. The evidence is the record of what the human did and when. A record showing one hundred approvals within a minute would call into question whether meaningful review occurred. It does not prove the absence of review on its own. It creates an evidentiary question: how was meaningful review possible at that pace? Policy describes what review should look like. The artifacts show whether it is plausible that it happened.
Establishing where human judgment entered the decision depends on this record. When responsibility for a decision is contested, the question is where human judgment ended and system output began. You cannot answer that from policy. You answer it from evidence of the intervention itself.
The decision and the action taken
The output is not the decision, and the decision is not the action. The record must connect them. When an AI system recommends denying a claim, closing an account, opening an investigation, or escalating an incident, the evidentiary question is what the organization did, and whether that action followed from the recommendation.
Without this link, you can reconstruct the AI interaction but not its significance. You can show that a recommendation existed. You cannot show that it drove anything. Establishing causal significance requires connecting the recommendation to the human judgment that adopted or rejected it, and connecting that judgment to the action the organization executed. This is what turns a preserved interaction into a preserved decision.
Surrounding organizational records
The decision does not exist alone. It sits inside an institution. The policies and standards in effect at the time, the purpose the decision served, the people who relied on it, and the lines of accountability around it all situate the decision and establish what was at stake.
These records answer a different question than the technical artifacts. The technical artifacts show how the output was produced. The organizational records show why it mattered and who was answerable for it. A complete account needs both. NIST's AI Risk Management Framework places the same emphasis on documented roles, responsibilities, system inventories, monitoring, and risk decisions.
The failure modes
Three patterns recur.
The first is the output-only fallacy. The organization keeps the result and discards the conditions. This feels sufficient until the decision is challenged, at which point the absence of context becomes the whole problem.
The second is state drift. The model and the data change continuously. An organization that does not capture the version and the data at decision time loses the ability to reconstruct the decision as more time passes. The record decays even when nothing is deleted.
The third is the human-in-the-loop assumption. The organization documents that a review step exists and treats that as evidence the review occurred. The design of the process is not the same as the record of the process. Only the artifacts show what happened.
Preservation is a design decision, not a response
You cannot preserve an artifact that was never captured. By the time a decision is questioned, the prompt, the retrieved data, the model version, and the intermediate logs may already be gone. Many systems do not retain them by default.
This means preservation is decided before the dispute, not after. It is a property of how the system is built and operated. An organization that waits until a decision is challenged to think about evidence may already have lost much of what it needed.
Current public-sector and regulatory frameworks reflect this principle. Canada's Directive on Automated Decision-Making requires federal departments within its scope to document decisions and assessments made or assisted by automated systems. The EU AI Act requires automatic event logging for high-risk AI systems and links human oversight to the ability to interpret, disregard, override, or reverse system output. These instruments differ in scope and legal effect, but both make the same operational point: traceability cannot be improvised after the event.
When AI acts rather than advises
The decision chain described above assumes a person remains in the path between system output and organizational action. A model produces an output, a person exercises judgment, a decision follows, and an action is taken. Preserving that chain is what lets you reconstruct the decision.
Agentic systems can compress the chain. Where a system can select or execute consequential actions directly, the human judgment step may become thin or disappear. When that happens, two things change. The evidentiary burden shifts from the record of human judgment toward the system's observable action trajectory. And a separate question opens that a decision record does not answer: who or what controlled the system that acted.
That is a different evidentiary problem, and it is the subject of a companion article. This one addresses what must be preserved to reconstruct a decision. The companion addresses what must be preserved to reconstruct control and support attribution when an AI system executes consequential action.
What this establishes, and what it does not
This framework establishes what must exist to reconstruct an AI-influenced decision. It does not, by itself, resolve two further problems. The first is integrity: showing that the preserved artifacts have not been altered, and establishing their provenance. The second is retention and governance: how long the artifacts should exist, who may access them, and under what authority. Preservation, integrity, and governance are three distinct layers. This article addresses the first.
What the first layer establishes is a capability. If you can reconstruct the decision as it occurred, from the effective instruction and data, through the system state and execution, to the output, the human judgment, the decision, and the action, then you can examine it on evidence. You can assess it, challenge it, defend it where the evidence justifies defense, and identify where it failed. Reconstruction gives you that capability. It does not confer legitimacy on its own.
If you cannot reconstruct the decision, none of that is available to you. You are left defending a conclusion you can no longer explain. In any setting where the decision might be examined by a regulator, a court, an auditor, or your own leadership, that is the difference that counts.
Selected authoritative references
- Directive on Automated Decision-Making, Treasury Board of Canada Secretariat. See especially the requirements concerning access to system components, including released versions of proprietary software components, and documentation of decisions and assessments.
- Guide on the Scope of the Directive on Automated Decision-Making, Government of Canada. This clarifies that the directive covers systems that make or assist with administrative decisions, including processes with a human in the loop.
- Algorithmic Impact Assessment tool, Government of Canada. The assessment addresses decision records, system-generated logs, explanations, data, system design, and organizational accountability.
- Artificial Intelligence Risk Management Framework 1.0, National Institute of Standards and Technology. The framework connects systematic documentation with transparency, accountability, monitoring, and human review across the AI lifecycle.
- Guide to Integrating Forensic Techniques into Incident Response, NIST SP 800-86, National Institute of Standards and Technology. The guide describes the forensic value of application data, configuration, documentation, timestamps, and multiple forms of logging.
- Regulation (EU) 2024/1689, Articles 12–14, European Union. These provisions address record-keeping, transparency, and human oversight for high-risk AI systems.
Related Zemi North practice: AI assurance.