Most teams ask the wrong question first.
They ask: "Is the organisation compliant?" That question triggers a document review. Someone checks for a policy, a completed risk assessment, and an approval record. If all exist, the team moves on.
The harder question matters during an incident, a customer complaint, or a manager asking what happened with that refund. It is different: "Can the team prove what happened on this specific decision, for this specific case, at this specific time?"
That question requires runtime proof, not merely an approval artifact. It matters for every team that uses AI in a workflow with real consequences. That includes not only regulated enterprises but any team where a wrong decision costs money or trust.
Every major AI governance framework published in the last three years recognises this. ISO 42001, the NIST AI Risk Management Framework, the NIST Generative AI Profile (AI 600-1), and the EU AI Act all require both pre-deployment controls and post-deployment monitoring. Even if your team is not subject to these frameworks today, understanding them matters if you sell to enterprise customers or plan to operate in regulated markets.
Knowing where each framework is strong and where each stops lets teams design systems that satisfy all four, rather than optimising for one and leaving gaps in the others.
ISO 42001: A Management System Approach with Strong Approval Gates and No Case-Level Schema
ISO/IEC 42001:2023 is the first international standard for AI management systems. It follows the Annex SL structure of ISO 27001. Teams familiar with information security management systems will recognise the pattern: context, leadership, planning, support, operation, performance evaluation, improvement.
Annex A defines 38 controls for AI-unique risks: bias, transparency, accountability, data quality, lifecycle governance.
Where it is strong on pre-deployment approval:
Clause 8 (Operation) requires gated approval stages. Models must pass defined criteria (bias tests, security assessments, ethical reviews) before production. Clause 5 (Leadership) requires top management to assign accountability for AI decision-making. Clause 6 (Planning) requires documented risk assessments, mitigation strategies, and incident response procedures before deployment.
The effect is a formal approval control plane. Each AI system in scope must have an owner, a risk tier, a documented intended use, and a review history. Changes require evidence of evaluation before release. This is governance-first.
Where it is strong on runtime monitoring:
Clause 9 (Performance Evaluation) mandates continuous monitoring, measurement, analysis, and evaluation of AI system performance. That includes internal audits, management reviews, and ongoing checks that deployed systems still meet their objectives. Industry guidance suggests roughly 30 per cent of AI risk management resources go to post-deployment monitoring for model drift, emerging risks, and performance degradation.
Clause 10 (Improvement) requires corrective action when monitoring reveals problems. That closes the loop between runtime observation and governance response.
Where it stops:
ISO 42001 does not prescribe runtime evidence at the case level. It requires monitoring and management reviews. It does not define a logging schema, a retention period, or a mechanism that links a runtime decision to the approval record that authorised the system's behaviour. The organisation fills that gap.
Two organisations can both be ISO 42001 certified and answer "what happened on this specific case?" differently. One may have a rich decision ledger tied to each case. The other may have system-level dashboards with no case-level drill-down. The standard permits both.
NIST AI RMF: A Risk-Based Feedback Loop without the Plumbing
The NIST AI Risk Management Framework (AI 100-1) appeared in January 2023. It organises AI governance around four core functions: Govern, Map, Measure, and Manage. It is not a certifiable standard. It is a voluntary framework for identifying, assessing, and mitigating AI risks through a continuous feedback loop.
Where it is strong on pre-deployment approval:
Govern defines organisational culture and processes for AI risk management. It sets accountability structures, risk appetite, and the policies that decide whether an AI system can proceed to deployment. Map contextualises the system: intended use, risk profile, stakeholders, deployment environment. Together they build the pre-deployment approval infrastructure (documentation, sign-offs, risk tiers).
NIST recommends approval workflows for high-risk AI applications. It suggests risk-tiered review cadences and an AI Risk Committee on a defined schedule.
Where it is strong on runtime monitoring:
Measure and Manage cover post-deployment. Measure uses quantitative, qualitative, or mixed-method tools to analyse, benchmark, and monitor AI risk over time. Manage allocates resources to respond: updating controls, triggering incident recovery, communicating with stakeholders.
The framework treats AI systems as living systems. Point-in-time testing at deployment is not enough. Controls must adapt as risks evolve. Monitoring signals (performance metrics, drift detection, policy violations) feed back into governance reviews. That is a continuous loop, not a gate-and-forget model.
Where it stops:
The NIST AI RMF does not architecturally separate pre-deployment approval from post-deployment evidence. It treats both as parts of one risk management lifecycle. Approval gates feed policy constraints into operational systems. Runtime signals feed back into governance reviews. The framework assumes this integration but does not prescribe how.
Philosophically, that is a strength: it avoids siloed approval and monitoring systems. Practically, it is a weakness. Most organisations run approvals in one system (a GRC platform, a change management tool, a review board) and runtime monitoring in another (an observability platform, a model registry, an operations dashboard). The integration rarely exists in practice. NIST describes the feedback loop. It does not provide the plumbing.
NIST AI 600-1: The Generative AI Profile That Recommends Logging but No Schema
NIST AI 600-1 appeared in July 2024. It is a profile: an implementation guide that adapts the AI RMF for generative AI systems. The Public Working Group on Generative AI developed it. It focuses on four considerations: governance, content provenance, pre-deployment testing, incident disclosure.
What it adds for pre-deployment:
It addresses risks that the base RMF handles generically and generative AI makes acute. It requires organisations to document roles, responsibilities, and communication lines specific to generative AI risk. It requires pre-deployment validation criteria. The question is not only "does the model work?" It is "does it meet content safety, provenance, and integrity requirements?" It also addresses retrieval-augmented generation (RAG) as a risk mitigation technique and a governance concern. Retrieval introduces risk vectors around data quality and source poisoning.
What it adds for runtime:
AI 600-1 recommends content logging, metadata annotation, watermarking, and source attribution where technically feasible. It calls for consolidated activity logging that captures data interactions, including third-party exchanges, with user IDs, timestamps, and metadata. The Govern 1.5 subcategory requires ongoing monitoring and periodic review of the risk management process, with defined frequencies and roles.
For prompt-related risks, it covers input validation, output safety monitoring, and content filtering. These are guardrails at both input and output boundaries. It does not provide a detailed prompt governance framework. It establishes that prompt-related risks must be assessed, monitored, and mitigated as part of the overall risk profile.
Where it stops:
AI 600-1 operates at the guidance level. It recommends logging but does not prescribe a schema. It recommends provenance tracking but does not define retention requirements. It recommends monitoring but does not specify what runtime evidence is enough for a regulatory inquiry. Teams that follow it faithfully will have better generative AI governance than those that do not. Two faithful implementations can still differ sharply in their ability to produce case-level evidence of a specific runtime decision.
EU AI Act: The Regulatory Enforcement Approach That Requires System Logs but Not Case Records
The EU AI Act entered into force in August 2024. Enforcement is phased through 2027. It is more prescriptive than the other three frameworks on runtime evidence. It is also the only one with enforcement teeth: non-compliance for high-risk AI systems can result in fines up to 3 per cent of annual worldwide turnover or EUR 15 million, whichever is higher.
Where it is strong on pre-deployment approval:
The Act classifies AI systems by risk tier. High-risk systems (Annex III) must meet extensive requirements before deployment: conformity assessments, technical documentation, quality management systems, registration in the EU database. Providers must establish risk management systems that operate throughout the AI system's lifecycle.
For deployers, Article 26 requires competent personnel for human oversight, verification that input data is relevant, and documented evidence of compliance measures.
Where it is strong on runtime evidence:
This is where the EU AI Act goes further than any other framework.
Article 12 mandates that high-risk AI systems technically allow automatic recording of events (logs) over the system's lifetime. Logs must be tamper-resistant and maintained for a period appropriate to the intended purpose, at least six months unless EU or member state law specifies otherwise.
Logs must capture events relevant to identifying when the system may present a risk and facilitating post-market monitoring. They must also support detection of anomalies, dysfunctions, or unexpected performance.
Article 26 requires deployers to monitor system operation based on provider instructions. They must detect anomalies and dysfunctions. They must also assess performance continuously. If a deployer suspects the system presents a risk, they must inform the provider, importer, distributor, and relevant market surveillance authorities without undue delay. Serious incidents must be reported immediately.
For biometric identification systems, logging requirements are more specific: usage periods, reference databases, input data characteristics, identities of individuals verifying results, and annual reports to surveillance authorities.
Where it stops:
The Act mandates system-level logging and monitoring. Its requirements are oriented around the system, not individual case-level decisions. Article 12 requires logs that identify risks and support monitoring. It does not explicitly require each decision or recommendation to be traceable to a specific case record with full context.
An organisation can comply with Article 12 through detailed system-level event logs (model inputs, outputs, performance metrics, anomaly flags). It may still lack the operational context of each decision. That context includes the reviewer, the available and blocked actions, the downstream system's response, and the case outcome.
The Act also creates tension with GDPR's right to erasure. Tamper-resistant logs retained for six months or longer may contain personal data that a data subject has the right to delete. The Act does not fully resolve this. Organisations must design retention architectures that satisfy both requirements.
The Comparison Matrix: Where Each Framework Stops
| Dimension | ISO 42001 | NIST AI RMF | NIST AI 600-1 | EU AI Act |
|---|---|---|---|---|
| Pre-deployment approval gates | Required (Clause 8). Gated stages with defined criteria. | Recommended (Govern, Map). Risk-tiered review workflows. | Recommended. Adds GenAI-specific validation criteria. | Required (Articles 6, 9, 10). Conformity assessment for high-risk systems. |
| Runtime monitoring | Required (Clause 9). Continuous performance evaluation. | Required (Measure, Manage). Adaptive controls, feedback loops. | Recommended. Content logging, metadata, provenance tracking. | Required (Articles 12, 26). Automatic event logging, deployer monitoring. |
| Case-level decision evidence | Not prescribed. Left to implementation. | Not prescribed. Assumed as part of integrated lifecycle. | Not prescribed. Recommends logging but no case-level schema. | Partially addressed. System-level logs required; case-level linking not explicit. |
| Retention requirements | Not specified. Follows organisational policy. | Not specified. Follows organisational policy. | Not specified. Recommends logging where feasible. | Minimum 6 months for high-risk system logs. GDPR tension unresolved. |
| Human oversight requirements | Required (Annex A controls). Roles and review cadence. | Recommended (Govern). Risk committee, accountability structures. | Recommended. Roles, responsibilities, communication lines. | Required (Article 14). Meaningful human oversight for high-risk systems. |
| Incident response | Required (Clauses 6 and 10). Corrective action process. | Required (Manage). Risk response plans, incident recovery. | Required. Incident disclosure as core consideration. | Required (Article 26). Immediate reporting for serious incidents. |
| Enforcement mechanism | Third-party certification audit. | Voluntary. No enforcement. | Voluntary. No enforcement. | Legal. Fines up to 3% annual turnover or EUR 15M. |
| Scope of applicability | Any organisation using AI. Global. | Any organisation. US-focused, globally adopted. | Organisations using generative AI. | Providers and deployers in EU market. Extraterritorial reach. |
The Gap All Four Share: No Bridge from Approval Record to Case Evidence
Every framework requires pre-deployment approval and post-deployment monitoring. None is naive about this. ISO 42001 Clause 9 requires continuous evaluation. NIST AI RMF builds a four-function feedback loop. NIST 600-1 adds generative-AI-specific logging guidance. The EU AI Act mandates tamper-resistant event logs.
None prescribes how to connect the approval record to the runtime decision at the case level.
The approval record lives in the governance system (a GRC platform, a review board, a release management tool). It says: this system was approved, under this policy, by this person, with this evidence, on this date.
The runtime log lives in the operations system (an observability platform, a model registry, an application database). It says: the system received this input, produced this output, at this time.
The case record often lives in a third system. It holds the operator review, the approval or block, the external call, the captured result, and the ticket resolution. Or it lives in no system at all, reconstructed from screenshots, chat messages, and memory.
The bridge between these three layers is an implementation problem, not a standards problem. The frameworks say you need approval, monitoring, and evidence. They do not say how to keep them connected on a single operational record.
That is the architectural gap this series addresses.
Where Latch Fits: The Case Record Bridges the Gap
Latch sits at the case level, where the approved AI capability meets a live decision.
The governance-first layer (the AI registry, the policy framework, the approval workflows, the release packets) sits above Latch. Latch does not replace that layer. It does not manage model inventories. It does not run conformity assessments. It does not store risk treatment documents.
Latch acts as the runtime layer. AI-assisted intake, policy-bounded recommendations, approval-gated execution, plugin action capture, and immutable audit trails converge on a single case record. The record preserves the recommendation, the available and blocked actions, the approval or override, the execution, the downstream system's response, and the outcome after closure.
The thesis: compliance with any of these four frameworks requires both the approval control plane and the runtime control layer. The connection between them must be deliberate, not assumed.
Continue Reading
- Governance-First Approval Systems for AI: What They Prove, What They Miss, and Where Runtime Evidence Fills the Gap
- The Runtime Control Layer: What This Category of Software Is and Why It Exists
- What to Log When Operators Trigger External Actions
- What Auditors Need to See in a High-Risk Approval Workflow