
Spend enough time examining the AI assurance landscape and a pattern emerges. There is no shortage of work on model evaluation, red teaming, governance, observability, safety, testing, validation, audit, and risk management. Each discipline contributes something necessary. Yet organizations can accumulate all of these mechanisms without a coherent way to answer the larger question:
What does the available evidence allow us to conclude about whether confidence in an AI system is justified for a specific purpose?
That is an AI assurance question. As AI systems become embedded in consequential processes, the challenge is no longer only whether they can be built or deployed. It is whether the confidence placed in them can be justified.
The Problem Is Not the Absence of Assurance Activities
It would be incorrect to suggest that organizations are doing nothing that resembles AI assurance. Quite the opposite. A substantial ecosystem already exists.
NIST places testing, evaluation, validation, and verification within the broader lifecycle of AI risk management, while the NIST AI Risk Management Framework 1.0 organizes activities through Govern, Map, Measure, and Manage. The UK government describes assurance techniques as mechanisms for measuring, evaluating, and communicating AI risks and for helping demonstrate safety, trustworthiness, and compliance. ISO/IEC 42001:2023 establishes an organizational management system for addressing AI-related risks and opportunities.
So the problem is not that assurance work does not exist. The problem is that much of the work remains distributed across different disciplines, teams, tools, vocabularies, and evidence sources. Consider what an organization might already be doing: LLM and model evaluation, model validation, verification, functional and robustness testing, adversarial testing, red teaming, model risk management, AI safety testing, bias and fairness assessment, observability, security testing, responsible AI governance, algorithmic auditing, compliance assessment, impact assessment, human oversight, and production monitoring.
Each answers a necessary question. But none, by itself, establishes how much confidence the deployed system deserves.
A Test Result Is Not an Assurance Conclusion
Suppose an AI system achieves 94 percent accuracy on an evaluation dataset. That is evidence.
Suppose a red team discovers three prompt-injection paths. Those findings are evidence.
Suppose production monitoring shows a low hallucination rate during the previous 30 days. That is evidence.
Suppose an algorithmic audit finds no statistically significant disparity within the population it examined. That is evidence.
Suppose an AI governance assessment confirms that the system has an assigned owner, documented risk classification, human oversight requirements, and approved usage policy. That is also evidence.
But none of those observations independently establishes that confidence in the deployed system is justified. The moment we make that larger statement, we move from an artifact, measurement, observation, or test result into a claim. That claim requires reasoning.
What population was evaluated? Were the test conditions representative of production? What was excluded? Was the evaluation dataset independent from development? Which model version was tested, and were retrieval components, tools, and agents included? Did authorization controls operate during testing? What happens under distribution shift? Were failure rates materially different for specific scenarios? What evidence conflicts with the positive results? What remains unknown? What level of residual risk has been accepted, and by whom?
These questions illustrate why assurance is larger than evaluation.
So What Is AI Assurance?
A useful way to understand AI assurance is as the discipline that connects claims about an AI system with the evidence needed to determine how much confidence those claims deserve.
This is different from governance. Governance may establish that high-risk systems require independent validation, that certain actions require human approval, that performance thresholds must be met, that sensitive data cannot be used under stated conditions, that systems must be monitored for drift, and that specific security controls must operate. Governance establishes the criteria, responsibilities, and decision boundaries. Assurance examines whether the evidence justifies confidence that they are operating as intended.
The relationship can be stated directly:
AI governance establishes what must be satisfied, by whom, and within which boundaries. AI assurance examines whether the evidence justifies confidence that those requirements are operating as intended.
This does not make assurance superior to governance. They perform different functions. Governance without assurance risks becoming declarative. Assurance without governance lacks the criteria against which confidence should be evaluated. The two need each other.
Assurance Is Also Larger Than the Model
Treating AI assurance as model assurance leaves too much of the deployed system outside the examination boundary. Modern AI systems are rarely just models. Consider a simplified enterprise architecture:
Data → preprocessing → embeddings → retrieval → context construction → model → orchestration → agent → tool/API → authorization → business rules → enterprise system → output/action → monitoring
An assurance failure can occur almost anywhere along this chain. The model might correctly summarize the information it receives while the retrieval system supplies outdated documents. It might produce an appropriate recommendation while the orchestration layer passes incorrect parameters to a downstream tool. An agent might reach the right decision but execute it with authority that should never have been available. A system might generate a factually correct response while exposing confidential information. An evaluation may demonstrate strong performance against a static benchmark while production performance deteriorates because the operating environment has changed. A model may behave correctly while the organization cannot later reconstruct which model version, system prompt, retrieved documents, tool calls, or business rules produced a consequential decision.
In each situation, evaluating the model alone would provide an incomplete picture. The object of assurance must therefore be the AI-enabled system within its operating environment, not merely the model at its centre.
The Assurance Evidence Stack
The disciplines that appear fragmented begin to make sense when their outputs are treated as parts of an evidence stack rather than as isolated activities.
Model and LLM evaluation provide evidence about performance, capability, limitations, and failure patterns. Verification and validation examine whether requirements have been implemented and whether the resulting system is appropriate for its intended purpose. Red teaming searches for weaknesses, unsafe behaviour, misuse paths, and assumptions that fail under challenge. Testing examines functional behaviour, robustness, reliability, security, boundary conditions, and failure modes. Model risk management contributes evidence about materiality, limitations, controls, risk acceptance, validation, and ongoing oversight. Observability supplies production evidence about what the deployed system is actually doing. Algorithmic and compliance audits independently examine controls, processes, outcomes, and obligations. Governance supplies the policies, accountability structures, decision rights, risk tolerances, and requirements against which that evidence must be interpreted.
NIST recognizes the importance of iterative testing and evaluation throughout the AI lifecycle and includes independent evaluation and documented assumptions among its risk-management considerations. These disciplines do not need to be collapsed into one. AI assurance connects their evidence to the claims on which decisions depend.
The Missing Step: From Evidence to Finding
Imagine an organization has performed all of the activities above. It now has hundreds of artifacts: evaluation scores, red-team reports, model cards, risk registers, validation reports, attack simulations, observability dashboards, incident records, audit findings, security assessments, governance approvals, benchmark results, bias testing, monitoring alerts, and human-review statistics.
Someone still has to determine what those artifacts establish. That is an evidentiary problem. Assurance-case approaches already recognize the need to connect claims with structured supporting evidence. The evidence must also be properly scoped, preserved, corroborated, interpreted within its limitations, and subjected to accountable human judgment. This is where the ZEMI Method becomes relevant.
The published ZEMI Method, Version 1.2, is a versioned evidentiary methodology for investigation and assurance. Its governing premise is that a claim relevant to a decision remains a hypothesis until proportionate testing establishes what the available evidence supports. ZEMI does not replace forensic standards, technical evaluation, or domain procedures. It governs the reasoning that connects authority, evidence, analysis, human judgment, and reporting.
This keeps ZEMI in its proper role. It does not replace LLM evaluation, red teaming, model validation, AI testing, the NIST AI RMF, ISO/IEC 42001, algorithmic auditing, model risk management, or security testing. Those mechanisms generate, structure, or govern evidence. ZEMI governs the reasoning used to determine what their combined evidence establishes.
Applying the ZEMI Method to AI Assurance
ZEMI Version 1.2 begins with an Authority and Scope Gate, followed by six phases: Frame, Preserve, Corroborate, Analyze, Adjudicate, and Report. It concludes with a standing defensibility check. Version 1.2 also defines the evidence-to-finding pathway, strengthens the test for independent corroboration, and separates technical execution, control, authorization, intent, and accountability where AI activity is material. Applied to AI assurance, the structure provides a disciplined path from source material to a human-adjudicated finding that can support, but does not make, an organizational decision.
Phase 0: Authority and Scope
Before assuring an AI system, determine what is actually being assured. Which system, which version, which use case, which environment, which users, and which decisions? Which risk classification, and against which policies, standards, laws, contractual obligations, or organizational requirements? Who has the authority to make the resulting assurance determination? Without a bounded scope, statements such as "the model is safe" or "the AI is compliant" are almost meaningless.
1. Frame
Convert broad assertions into testable claims. Instead of asking whether the AI is accurate, ask whether the deployed system maintains the defined performance threshold for this specific task, population, environment, and operating condition. Instead of asking whether the agent is secure, ask whether it can perform consequential actions outside its authorized tool, identity, transaction, or policy boundary. Instead of asking whether the RAG system is grounded, ask whether material factual claims in generated outputs are supported by authoritative retrieved sources under the defined retrieval conditions. Now we have something that evidence can actually test.
2. Preserve
Preserve the evidence necessary to evaluate those claims. For AI systems, that can include model versions, system prompts, configuration, evaluation datasets, retrieval sources, embeddings, tool definitions, agent traces, API calls, policy states, approval records, evaluation results, red-team findings, monitoring telemetry, human overrides, source documents, system logs, and relevant environmental state. Without preservation, assurance quickly becomes retrospective storytelling.
3. Corroborate
This is where fragmentation can become complementary rather than merely disconnected. Different assurance disciplines become independent or complementary evidence sources. An evaluation result can be compared against production telemetry. A model claim can be challenged through red teaming. A vendor assertion can be compared against independent testing. An observability signal can be checked against raw transaction records. An AI-generated explanation can be compared with the actual sources and system state from which the decision was produced.
Corroboration also means looking for disagreement. When two credible evidence sources conflict, the conflict itself becomes something to investigate rather than something to average away.
4. Analyze
Now determine what the combined evidence establishes. Where are the failure modes? Which assumptions remain, which controls operated, and which controls failed? What uncertainty remains? Are results sensitive to model version, prompt variation, retrieval state, population, language, environment, or adversarial conditions? Do supposedly independent tests actually share the same underlying failure mode? This is where assurance moves beyond dashboards and scores into reasoning.
5. Adjudicate
The ZEMI Method maintains a mandatory human adjudication role. Automation can assist with retrieval, correlation, summarization, pattern proposal, and drafting, but the determination of the finding remains human.
This boundary becomes especially important when AI is used to assure AI. An AI system may help analyze thousands of test results, identify anomalous behaviour, correlate incidents, summarize red-team findings, and propose an assurance conclusion. But an analytical instrument should not silently become the authority that determines whether confidence in another system is justified. Human accountability must remain identifiable.
6. Report
Finally, communicate what the evidence supports, together with its limitations. ZEMI classifies findings in three categories:
Known. Supported by evidence within the stated authority, scope, and limits.
Assumed. Relied upon for a stated and limited purpose but not directly established.
Undetermined. Not resolvable on the available evidence after proportionate inquiry.
Imagine applying that discipline to an AI assurance report. Instead of saying "the system is safe," we might say:
Known: Under the tested conditions, unauthorized execution outside the defined tool permission boundary was not observed across the evaluated scenarios.
Assumed: The production identity configuration remains equivalent to the configuration independently validated during testing.
Undetermined: Resistance to previously unseen multi-agent privilege-escalation paths has not been established.
That is a much more defensible statement.
Assurance Should Produce Bounded Confidence, Not Absolute Trust
The goal of AI assurance should not be to declare that an AI system is simply "trustworthy." Trustworthiness is contextual. A system can be reliable enough for summarizing internal meeting notes and completely unsuitable for independently approving financial transactions. The same model can require different levels of assurance depending on what surrounds it, what authority it receives, what data it accesses, what decisions it influences, and what happens when it fails.
Assurance therefore needs to produce bounded confidence: confidence in this system, for this purpose, under these conditions, supported by this evidence, subject to these limitations, for this period of time.
The time boundary prevents assurance from becoming a one-time declaration. Models, prompts, retrieval corpora, tools, business rules, threats, data distributions, and vendor services change. Users discover behaviours the original evaluators did not anticipate. Operational evidence must therefore continue to inform the assurance finding.
From Governance to Evidence to Defensible Finding
The relationship can be expressed as an evidence and decision cycle.
The AI assurance evidence-to-decision cycle
-
01
Governance criteria
Define what must be satisfied.
Obligations, accountability, risk tolerance, decision boundaries, and required controls establish the criteria for examination.
-
02
Assurance activities
Examine the system from multiple directions.
Evaluation, validation, verification, red teaming, testing, safety assessment, observability, security assessment, impact assessment, model risk management, and audit produce evidence.
-
03
Assurance evidence
Preserve the resulting record.
Scores, observations, logs, reports, test results, failures, control records, incidents, exceptions, approvals, and measurements form the evidentiary record.
-
04
ZEMI evidentiary reasoning
Determine what the combined evidence establishes.
- Authority and Scope
- Frame
- Preserve
- Corroborate
- Analyze
- Adjudicate
- Report
-
05
Bounded assurance finding
Classify the conclusion.
- Known
- Assumed
- Undetermined
-
06
Governance and operational decision
Act within the limits of the finding.
- Deploy
- Restrict
- Remediate
- Monitor
- Escalate
- Reassess
- Retire
Operational evidence returns to the assurance record. Monitoring, incidents, changes, and decisions can alter the scope, controls, or confidence previously established.
This is a feedback loop rather than a compliance exercise. Governance establishes what needs to be demonstrated. Technical and organizational activities produce evidence. ZEMI provides a disciplined method for determining what that evidence supports. The resulting findings inform deployment, remediation, monitoring, risk acceptance, and governance. New operational evidence then begins the cycle again.
Where AI Forensics Fits
AI assurance asks what evidence justifies confidence in the system. AI forensics asks what happened when the system, its behaviour, or its consequences became disputed. The two should not be isolated. Every AI incident can expose assumptions that assurance should reconsider, and every mature assurance program should preserve enough evidence for future incidents to be reconstructed.
If an autonomous agent performs an unauthorized transaction, for example, assurance and forensics immediately intersect. What model was operating? What context did it receive, and what retrieved evidence influenced the decision? Which tools were available, and what identity executed the transaction? What authorization checks occurred, and what system state existed immediately before execution? Was human approval required, and was it actually obtained? Could the action be reproduced, and which part of the chain failed? If the organization cannot answer those questions, it has both a forensic problem and an assurance problem.
The Opportunity Ahead
AI assurance does not need another isolated checklist. Evaluation teams, red teams, security teams, model validators, risk teams, and auditors already produce evidence. Governance establishes the requirements. Observability contributes operational evidence. The unresolved task is to connect those sources and determine what is known, what is assumed, what remains undetermined, and what conclusion can be defended.
That is the connective role of AI assurance. It is not a replacement for governance or testing, another name for model evaluation, or a general certification of trustworthiness. It is the discipline through which claims about AI systems are subjected to evidence, challenge, limitation, human judgment, and continuing re-evaluation.
The ZEMI Method provides one evidentiary approach to that problem. Its premise is intentionally simple:
Evidence before conclusion.
As AI moves from generating content to influencing decisions, invoking tools, interacting with enterprise systems, and taking consequential actions, the evidentiary burden grows with it. The future of trustworthy AI will depend not only on what can be built. It will depend on what can be established about the systems placed into use.
References
National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). https://doi.org/10.6028/NIST.AI.100-1
National Institute of Standards and Technology. (n.d.). AI test, evaluation, validation and verification (TEVV). https://www.nist.gov/ai-test-evaluation-validation-and-verification-tevv
UK Government. (2024). Introduction to AI assurance. https://www.gov.uk/government/publications/introduction-to-ai-assurance/introduction-to-ai-assurance
UK Government. (2021). The roadmap to an effective AI assurance ecosystem. https://www.gov.uk/government/publications/the-roadmap-to-an-effective-ai-assurance-ecosystem/the-roadmap-to-an-effective-ai-assurance-ecosystem
International Organization for Standardization. (2023). ISO/IEC 42001:2023: Information technology: Artificial intelligence: Management system. https://www.iso.org/standard/81230.html
Watson, K. V. (2026). The ZEMI Method: An evidentiary methodology for investigation and assurance, Version 1.2. Zenodo. https://doi.org/10.5281/zenodo.22084060