Global AI Assurance.
AI oversight has quietly changed shape. The question is no longer whether a system was assessed before deployment. It is whether the institution can see what its systems are doing, test how they behave, constrain what they may do, intervene when they misbehave, and prove all four afterwards.
Dax Philbert, LLM
Chairman & CEO, Cabier Consulting · 14 September 2026 · ~22 min read

Executive summary
For three years the AI oversight conversation was about documents. Policies, principles, risk taxonomies, maturity grids, framework adoption. That work was not wasted, but it described intent rather than producing state. The systems it was meant to govern have since acquired memory, tools, data access, the ability to take action, and the ability to interact with other systems and other agents. Documents do not constrain any of that.
The direction of travel across the industry, including the largest platform providers publishing their own operational expectations during 2026, points the same way: agent identities, tool permissions, monitoring of actions, control over sensitive and irreversible operations, failure detection and remediation, and documented capability limits. Those are operational controls, not policy statements. Read them together and they describe an assurance layer rather than a governance document.
This article sets out that layer as nine disciplines and ten working capabilities, and maps each one to a surface an institution can inspect. It is written from the position Cabier has been building toward since before the term was fashionable: the institution chooses the AI, and the assurance layer above it has to be independent of every provider underneath.
The shift nobody announced
Traditional oversight of models followed a straight line: assess, approve, deploy. That worked because the thing being assessed held still. A credit model validated in March behaved in September much as it had in March, and the annual review cycle was proportionate to the rate of change.
Nothing about a modern AI estate holds still. The underlying model version changes beneath the institution, often without a change ticket. The retrieval corpus is refreshed weekly. A prompt is edited by a product owner on a Tuesday afternoon. An agent acquires a new tool permission because an integration shipped. A vendor turns on a summarisation feature inside a product the institution already licences, and an unowned model enters the estate without anyone deciding to admit it.
An annual assessment of a system that changes weekly is a photograph of a river. The assessment is not wrong; it is simply about a moment that has passed. What replaces it is not a better assessment. It is a loop that never stops running.
Nine disciplines, one question each
Global AI Assurance is the umbrella, and underneath it sit nine disciplines. Each one exists because it answers a question a board, a supervisor or an auditor will eventually ask in exactly these terms.
Most institutions can answer the first discipline and struggle with the second. Almost none can answer the third, because constraint is an engineering property rather than a documentary one. The remaining six are rarely owned by anyone at all.
You cannot assure an unknown estate
Every AI assurance programme that begins with policy ends in the same place: a control set applied to the systems somebody remembered to declare. The estate is always larger than the register. It contains models inside purchased software, agents built by a business team on an approved platform, copilots enabled by a default setting, retrieval indexes over document stores nobody classified, and tool servers standing between an agent and a production system.
Discovery is therefore the first capability rather than the last. It works from signals the institution already emits: identity and consent grants, network egress to inference endpoints, expense and licence records, code repositories, tool-server registrations, data-access requests. The output is not a tidy inventory on day one. It is a list of things found, things whose owner is unknown, and things running without approval. That third list is where the exposure sits.
Permission is not access
Access control asks whether a system may reach a resource. Agent authority asks a different question: within that resource, which actions may it take, up to which threshold, under which conditions, and who is told when it tries to go further. A treasury agent may read liquidity positions, calculate a funding requirement and recommend a transfer, while being unable to execute a transfer above a stated amount, create a beneficiary, modify settlement instructions or interact with an external wallet.
Writing that envelope down is straightforward. Enforcing it at the point of action is the work, and proving it later is the assurance. The artefact that matters is the refusal: a recorded, retained event showing that the boundary was reached and held. A boundary nobody ever tested is an assertion. A refusal in the evidence set is a control that operated, and it is the single most persuasive thing an institution can put in front of an examiner on this subject.
From approval gate to running loop
The replacement for assess, approve, deploy is a loop with seven positions: observe, evaluate, control, intervene, evidence, learn, reassess. Observation is telemetry over real behaviour rather than a self-reported status. Evaluation is continuous testing against the institution's own standards. Control is the envelope holding. Intervention is the ability to narrow, pause or stop a system in production without a release cycle. Evidence is emitted by the loop rather than assembled afterwards. Learning changes the standard. Reassessment feeds the next revolution.
Pre-deployment assessment does not disappear in this model. It becomes the entry condition to the loop rather than the whole of the governance activity. The practical test of whether an institution has made this shift is simple: ask how long it would take to narrow one production agent's authority by one permission, and who is allowed to do it. If the answer involves a change advisory board and a fortnight, the institution has documentation and not control.
Independent evaluation and dispositions
Evaluation spans more than model quality. It covers prompt injection, jailbreak, privilege escalation, data exfiltration, unauthorised action, hallucination in a consequential context, policy violation, model drift, agent drift, emergent behaviour and cross-agent interaction. The last of those is the newest and the least governed: two individually well-behaved agents can produce an outcome neither was approved to reach.
What the exercise produces matters as much as what it tests. A score invites debate. A disposition forces a decision: pass, pass with conditions, remediation required, human review, or block, each with the evidence that supports it and, where conditions attach, an expiry. And the evaluation has to be independent of the team that built the system, for the same reason independent validation exists in prudential model risk management. A party that builds the system cannot credibly grade it.
Incidents, early warning, systemic risk
Most AI incident logs are unusable, not because the incidents were unimportant but because the record lacks the fields that make comparison possible. A usable record names the model, the version, the agent, the environment, the tool, the data, the action taken, the control that failed, the obligation engaged, the impact, the containment and the change made afterwards. With those fields, patterns become visible across the estate rather than anecdotes inside a team.
Patterns are what convert incident management into early warning: authority quietly widening after approval, evidence ageing past its retention promise, an unowned model appearing repeatedly in the same business line, drift concentrating in one provider, escalations that never reach a human. Aggregate the same patterns across institutions and a further question appears, which supervisors will ask before institutions volunteer it: how much of the sector is exposed to the same failure at the same moment?
Supply chain and concentration
The AI supply chain is longer than procurement records suggest. Foundation model, API provider, hosting, inference infrastructure, model adapter, agent framework, tool server, dataset, application, human reviewer, customer. Each link is a dependency, and the institution has contracted with perhaps three of them.
Mapped as a graph rather than a vendor list, the concentration becomes obvious and uncomfortable. A large share of critical workflows tends to resolve to one model provider. A surprising number of agents resolve to one tool-server operator. A proportion of inference happens outside the jurisdiction the workload was approved for. And a set of embedded models has no owner at all. These are resilience findings, not procurement findings, which is why they belong in the same reporting line as third-party and operational exposure rather than in a technology paper.
Sovereignty as a decision, not a location
Sovereignty is routinely reduced to a data-centre region, which answers almost none of the questions that matter. The operative questions are: what data, whose data, which country's law governs it, which model processes it, where inference is performed, where logs and evidence are retained, who can access the model, which jurisdiction can compel that access, whether the workload may cross a border at all, and what happens when an approved model becomes unavailable.
Answered per workload, those questions resolve to a disposition: sovereign, approved, restricted or prohibited. The value is not the label. It is that the reasoning is retained, so a decision made in one quarter can be defended in the next and revisited when a provider changes a hosting region without asking.
What the board actually receives
Every capability above produces signal, and signal without a common language becomes another dashboard. Cabier resolves it into fourteen categories covering governance, model integrity, data provenance, security, agent authority, human oversight, regulatory compliance, resilience, third-party dependency, transparency, runtime behaviour, incident history, evidence quality and sovereignty. Those roll into the institution's Operational Resilience Score alongside cyber, conduct, tokenisation and third-party signals.
The reason for that last step is the one that persuades boards. A standalone AI governance score is uninterpretable: nobody knows what to do with eighty-two. A resilience position is a decision. The board is told how AI has moved the institution's operational resilience, in which direction and through which dependency, in the same language it already uses for every other exposure it governs.
How a programme is composed
There is no single global AI assurance programme, and any provider claiming one is describing a template. What scales is composition: a global horizontal kernel, plus a jurisdiction pack, plus an industry pack, plus the enterprise AI estate as it actually is.
Canada and banking carries a different obligation set, evidence format and residency expectation from the United States and insurance, which differs again from the United Kingdom and legal services, or Singapore and government. The packs differ. The kernel underneath does not, which is what allows a group supervised in several places to answer each supervisor consistently about the same system.
Where each capability sits
Everything described above is inspectable rather than aspirational. Each capability has a surface, and each surface publishes its structure while keeping scoring weights, grading rubrics and the full control set within engagement. Figures shown on those surfaces describe a sample estate.
What the board should ask
Five questions separate an AI assurance programme from an AI policy. Can we produce, today, a complete list of models and agents in production, including those inside vendor products? Does every agent that can take action have a named human owner and a written boundary? Can we narrow that boundary in production this afternoon, and who is permitted to do it? When something went wrong, do we hold a record with enough structure to compare it to the next one? And has anyone told the board how AI changed the institution's resilience position, rather than how mature its governance is? Uncertainty on any of the five is the work.
FAQs
What is Global AI Assurance?
It is the umbrella category above AI governance: nine disciplines that together answer whether an institution can govern, prove, constrain, defend, survive, map, locate, trace and bound what its AI estate does.
How is it different from AI governance?
AI governance answers who owns the estate and under which policies. Assurance answers whether those policies demonstrably operated. Most institutions have the first and cannot evidence the second.
Why does agent identity matter?
An agent that takes action against institutional systems needs the same treatment as a member of staff with those permissions: an identifier, an owner, a purpose, a scope of authority and an escalation path.
What is a permission envelope?
A scoped statement of what an agent may do and what it may not, enforced at the point of action. Access is binary; authority is bounded. A treasury agent may calculate a funding requirement and still be unable to move money.
Why is a refusal treated as evidence?
Because it is the only artefact that proves the boundary exists. A control nobody ever tested is an assertion. A recorded refusal is a control that operated.
What is runtime assurance?
Continuous observation, evaluation, control, intervention, evidence, learning and reassessment of systems already in production, rather than a single pre-deployment assessment that ages out.
Does pre-deployment assessment stop mattering?
No. It remains necessary and becomes insufficient. A system whose model version, retrieval corpus, prompts and tool permissions change monthly cannot be governed by an annual review alone.
What does independent evaluation produce?
A disposition with evidence: pass, pass with conditions, remediation required, human review, or block. A number without a decision attached is not assurance.
Why keep evaluation independent of the builder?
Prudential model-risk guidance and the EU AI Act both reject self-attestation for material systems. A party that builds the system cannot credibly grade it.
What belongs in an AI incident record?
Model, version, agent, environment, tool, data, action, control that failed, obligation engaged, impact, containment and the change made afterwards. Without those fields an incident log cannot support early warning.
Is AI concentration risk real or rhetorical?
It is measurable. Map the dependency chain and the same provider, tool server or inference region tends to appear underneath a large share of critical workflows. That is a resilience finding, not a procurement one.
Is sovereignty about where the data centre is?
No. It is a control-plane property: whose data, which model, where inference happens, where logs and evidence are retained, which jurisdiction can compel access, and what happens when an approved model becomes unavailable.
What is a sovereignty disposition?
Each workload resolves to sovereign, approved, restricted or prohibited, with the reasoning retained. The disposition is a decision the institution can defend rather than a preference.
Is this model-neutral in practice?
Yes. The layer governs Western, open-weight, sovereign, private and on-premise estates alike. Cabier does not train, host or resell governed models, publishes no vendor comparison and endorses no provider.
Why not give the board an AI score?
Because an isolated AI score has no decision attached to it. The board should be told how AI has moved the institution's operational resilience position, in the same language as cyber, third-party and conduct exposure.
How is a programme scoped across countries and industries?
One horizontal kernel, plus a jurisdiction pack, plus an industry pack, plus the enterprise estate. Canada and banking is not the same obligation set as the United States and insurance, but the kernel underneath is identical.
Is any of this contingent on pending legislation?
No. The operative anchors are in force: the EU AI Act, DORA, existing prudential model-risk guidance, sectoral and state rules. Pending bills are tracked and not relied upon.
What is published and what is not?
Maps, field names, question sets and dispositions are public so the structure can be judged. Scoring weights, grading rubrics, Trust Gate internals and the full control set remain within engagement.
Glossary
- Agent
- A process permitted to take action on institutional systems using a model's output.
- Agent drift
- Behavioural change in an agent over time, from tools, memory, prompts or context rather than model weights.
- Agent identity
- The governed record that makes an agent an institutional entity: identifier, owner, purpose, authority.
- AI estate
- Every model, agent, dataset, tool, application and vendor-embedded capability in production.
- AI Trust Score
- Cabier's fourteen-category read of AI exposure, feeding the Operational Resilience Score.
- Assurance Kernel
- The shared control, evidence and scoring engine underneath every Cabier module.
- Assurance certificate
- A time-bounded attestation that names the conditions which invalidate it.
- Blast radius
- The maximum scope of systems and records an agent can affect before escalation.
- Concentration read
- A measure of how much critical activity depends on a single provider, tool or region.
- Continuous control monitoring
- Control effectiveness measured on an ongoing basis rather than at a review date.
- Decision Fabric
- The layer that resolves what happens when a control, gate or evaluation returns a negative result.
- Disposition
- The decision an evaluation produces: pass, pass with conditions, remediation, human review or block.
- Early warning pattern
- A recurring signal across incidents that indicates exposure before loss occurs.
- Evidence Vault
- The retained, signed artefact store that supervisory and audit answers are drawn from.
- Envelope
- The bounded authority of an agent: what it may do, what it may not, and where it must escalate.
- Estate discovery
- Finding models, agents, tools and embedded AI that no register currently contains.
- Independent evaluation
- Testing performed by a party that did not build the system under test.
- Inference locality
- Where model computation physically occurs for a given workload.
- Jurisdiction pack
- The jurisdiction-specific obligations, evidence formats and residency rules layered on the kernel.
- Operational Resilience Score
- The nine-dimension institutional resilience measure that AI exposure rolls into.
- Permission fabric
- Enforcement of scoped agent authority at the point of action.
- Runtime assurance
- Continuous observation, control and intervention over systems already in production.
- Shadow AI
- Model or agent use that no owner has registered and no control currently covers.
- Sovereignty disposition
- Sovereign, approved, restricted or prohibited, resolved per workload.
- Trust Graph
- The dependency model linking obligation, system, control, evidence and accountable person.
References and citations
Primary sources. Positions change; verify at source before relying on any figure or determination.
- 1European Union, Regulation (EU) 2024/1689 (Artificial Intelligence Act) — Provider and deployer obligations, including general-purpose AI models.Source
- 2NIST AI Risk Management Framework (AI RMF 1.0) and Generative AI Profile (NIST AI 600-1) — Function taxonomy underlying the evaluation and monitoring axes.Source
- 3ISO/IEC 42001:2023 — Artificial intelligence management system — Management-system reference for the continuous assurance loop.Source
- 4Board of Governors of the Federal Reserve System / OCC, Supervisory Guidance on Model Risk Management (SR 11-7 / OCC 2011-12) — Independent validation and effective challenge expectations relied on for the evaluation position.Source
- 5European Union, Regulation (EU) 2022/2554 (DORA) — ICT and third-party resilience obligations applied to model, tool and inference providers.Source
- 6OSFI, Guideline E-23 — Model Risk Management — Canadian independent-review expectations across model types.Source
- 7OWASP Top 10 for Large Language Model Applications — Reference taxonomy for prompt injection, exfiltration and privilege-escalation test classes.Source
- 8NYDFS, 23 NYCRR Part 500 (as amended) and associated AI cybersecurity guidance — Certification and governance obligations for covered entities.Source
- 9Monetary Authority of Singapore, FEAT Principles and Veritas materials — Fairness, ethics, accountability and transparency expectations.Source
- 10UK DSIT, A pro-innovation approach to AI regulation, and FCA supervisory statements on AI — Regulator-led, sector-specific approach relied on for the UK delta.Source
Named sources
- Public regulatory and standards sources through September 2026 — EU AI Act and DORA Official Journal texts; NIST AI RMF and the Generative AI Profile; ISO/IEC 42001; SR 11-7 and OCC 2011-12; OSFI E-23; NYDFS Part 500 and AI guidance; MAS FEAT; UK DSIT and FCA statements; OWASP LLM application guidance. Platform-provider responsible-AI publications are treated as an observed direction of travel only. No provider endorses, sponsors or reviews this analysis, no comparative benchmarking is performed, and Cabier does not train, host or resell governed models.
Global AI Assurance
The umbrella: nine disciplines, the layered architecture, the composition model.
OpenAgent Permission Fabric
Attempt an action outside the envelope and watch the refusal become evidence.
OpenAI Runtime Assurance
The seven-position loop, running, against the static lifecycle.
The Assurance Layer
Why an operating system, and not another framework, closes the gap.