A code of conduct needs an assurance layer.
A model provider has put a draft code of intended model behaviour out for public consultation and asked for feedback. This is ours. We support the direction and argue one thing throughout: a commitment that cannot be evidenced becomes the deployer's problem, and in regulated institutions the deployer is the one standing in front of the supervisor.
Dax Philbert, LLM
Chairman & CEO, Cabier Consulting · 14 September 2026 · ~14 min read

Position in brief
Publishing intended model behaviour for public comment is the right instinct, and stating that these systems must remain subordinate to human authority is the right starting point. Our submission takes no issue with the values. It concerns verifiability.
A behavioural code sits at the model layer. Accountability sits at the institutional layer, with whoever put the system in front of a customer, a market or a regulator. Between those two layers is a gap that no code closes on its own: the difference between what a model is intended to do and what an institution can demonstrate it did, on a named date, in production, under the version it actually had. That gap is the assurance layer, and every one of our twelve recommendations is aimed at making it narrower.
We are model-neutral by design. Cabier governs Western, open-weight, sovereign, private and on-premise estates, endorses no provider, publishes no vendor comparison, and neither trains, hosts nor resells the models it assures. Nothing in this response should be read as approval or criticism of any commercial party.
What is being consulted on
The document is a first draft describing the intended behaviour and values of a provider's own models, published for open comment rather than adopted as an operative training standard, with a stated review window of several weeks. It sets out that models should remain under human control, should not resist interruption, and should not set their own goals, and it places human interests ahead of the system's.
We read this the same way we read every provider publication of expected behaviour during 2026: as an observed direction of travel, in which the industry is migrating from principles toward operational expectations. It is a signal about where oversight is heading, not a validation of anybody's product, ours included.
Intent, behaviour, evidence
Three things are routinely collapsed into one in this debate. Intent is what the builder means the system to do. Behaviour is what the system does in a specific deployment, with specific tools, specific data and specific authority. Evidence is what a third party can inspect afterwards to establish which of the two occurred.
A code of conduct is a statement of intent, and a good one shapes behaviour. It cannot produce evidence, because evidence is generated at runtime inside the deploying institution: the tool call that was refused, the action held for human confirmation, the interruption that landed, the evaluation that failed and blocked a release. A board cannot show a supervisor a value. It can show a refusal log.
This is why our recommendations are almost entirely structural. Each one asks the same question of the draft: if this commitment were breached in production next March, what would exist to prove it, and who would be able to read it?
Twelve recommendations
None of these require the disclosure of weights, training data or commercially sensitive architecture. They are drafting and process changes.
Write each behaviour as a testable assertion
Every stated behaviour should carry an observable failure condition. “The model should not resist shutdown” is a value. “On receipt of an interruption signal the model ceases tool invocation within a stated interval, emits no further external calls, and records the interruption” is an assertion a third party can test. A code drafted in the first register cannot be assured; a code drafted in the second can.
Name the evidence artefact beside each commitment
For each commitment, state what an independent party may inspect to conclude it held: an evaluation result, a refusal log, a permission decision, a telemetry field, an incident record. Commitments without a named artefact become undischargeable in an audit, and the burden then transfers silently to the deployer.
Version the code and bind it to model versions
Publish the code with a version number, a change log and a machine-readable diff, and state which model versions were trained and operated under which code version. Regulated deployers must be able to answer what behavioural commitments applied on a given date to the version they had in production. Today they cannot.
Separate model-level commitments from deployment-level ones
Some commitments are properties of the model. Others depend entirely on how a deployer configures tools, retrieval, memory and authority. Marking each commitment as model-inherent, configuration-dependent or deployer-owned would prevent the most common governance error we see: an institution treating a provider commitment as a control it therefore need not implement.
Specify interruptibility as an operational control
Human control is asserted more often than it is engineered. A code should state who may interrupt, through which interface, at what latency, what happens to work already in flight, whether irreversible actions are held pending confirmation, and what record the interruption produces. Absent those five properties, controllability is a claim rather than a control.
Address agent authority, not only model behaviour
The material risk in 2026 sits above the model: an agent with tools, credentials, memory and the ability to act on institutional systems. A behavioural code should state expected defaults for tool permission, delegation to other agents, spend and action authority, and the treatment of irreversible operations. Model behaviour and agent authority are different governance objects and should be addressed separately.
Accept independent evaluation and publish dispositions
Self-attestation is already being retired in regulated contexts. Provide a route for independent evaluation of the behavioural claims by parties who did not build the model, and express results as dispositions with evidence — pass, pass with conditions, remediation required, human review, block — rather than as scores. A number without a decision attached is not assurance.
Align the incident taxonomy to regulated reporting
Behavioural incidents should be recorded in fields a supervised institution can use: model, version, environment, tool, data, action taken, control that failed, obligation engaged, impact, containment, and the change made afterwards. Where a provider incident feeds a deployer's statutory clock, the provider notification should carry those fields and a timestamp the deployer can rely on.
State the withdrawal and continuity position
A code that governs behaviour should also state what happens when behaviour cannot be assured: withdrawal of a capability, deprecation of a version, refusal of a jurisdiction. Deployers need notice periods, migration windows and evidence retention commitments, because for a regulated institution a withdrawn model is a resilience event.
Recognise sovereignty as a control-plane property
Where inference happens, where logs and evaluation evidence are retained, and which jurisdiction can compel access are governance facts, not deployment details. A behavioural code should state which commitments hold identically across hosted, sovereign, private and on-premise configurations, and which do not.
Provide safe harbour for adversarial testing
Behavioural claims are only credible if outsiders may try to break them. A stated non-retaliation position for good-faith adversarial testing, with a disclosure channel and a response commitment, converts the code from a statement into something falsifiable.
Publish how the consultation itself was dispositioned
Consultation is a governance process and should be evidenced like one. Publish the volume of submissions, the themes, which changes were made, and which recommendations were declined and why. A feedback log is the difference between consultation and announcement.
The deployer's residual duty
Suppose every recommendation above were adopted in full. A regulated institution would still have to discover the models and agents inside its own estate, including the ones that arrived enabled by default inside purchased software. It would still have to register each agent with a named human owner, bound tool permissions and a defined financial and action authority. It would still have to observe behaviour in production, intervene when it drifted, and retain the evidence that it did.
No provider commitment reaches into that estate. This is the practical reason we argue for commitments to be labelled by owner: the most expensive governance failure we encounter is not a broken provider promise. It is an institution that read a provider promise as a control it therefore did not need to build.
Where a code meets in-force law
The obligations that bite are already operative. General-purpose model duties and deployer duties under the EU AI Act. Serious-incident reporting with clocks that start on discovery. Prudential model-risk guidance that requires validation by a party that did not build the model. Third-party dependency, concentration, exit and continuity expectations under DORA. Sectoral and state rules layered on top.
A behavioural code drafted without reference to those instruments still has to be translated by every regulated deployer, separately, at cost, and inconsistently. Drafted with them in view, the same commitments can be consumed directly as control evidence. That is a small editorial difference with a very large downstream effect.
What we are not asking for
We are not asking for model weights, training corpora or architectural detail. We are not asking for a provider to accept liability for how an institution configures a system. We are not proposing that behavioural codes be replaced by regulation, nor that they be treated as a substitute for it. And we are not asking any provider to adopt, license or reference our own control set.
We are asking for one property: falsifiability. A commitment that cannot fail observably cannot be assured, and an unassurable commitment eventually becomes an unowned risk.
Text of our submission
Published in full, unedited. Institutions preparing their own submissions are welcome to reuse the structure.
Submission — 14 September 2026
Cabier Consulting — response to the public consultation on a draft AI Code of Conduct Submitted 14 September 2026 by Dax Philbert, LLM, Chairman & Chief Executive, Cabier Consulting. Cabier Consulting operates trust and assurance infrastructure for regulated enterprises and digital financial markets. We are model-neutral: we govern Western, open-weight, sovereign, private and on-premise AI estates, and we neither train, host nor resell the models we assure. We therefore respond as a party that has to evidence, to boards and supervisors, whether commitments of this kind held in production. Our position in one sentence: a behavioural code is necessary and insufficient, and the distance between the two is the assurance layer. We support the publication of intended model behaviour for consultation, and we support the principle that these systems remain subordinate to human authority. Our recommendations concern verifiability rather than values. 1. Write each behaviour as a testable assertion with an observable failure condition. 2. Name, beside each commitment, the artefact an independent party may inspect. 3. Version the code, publish a change log, and bind code versions to model versions. 4. Mark each commitment as model-inherent, configuration-dependent or deployer-owned. 5. Specify interruptibility operationally: who may interrupt, through which interface, at what latency, what happens to work in flight, and what record results. 6. Address agent authority as a distinct object from model behaviour: tool permission, delegation, spend and action authority, and irreversible operations. 7. Provide a route for independent evaluation of the behavioural claims and express results as dispositions with evidence rather than scores. 8. Align the behavioural incident taxonomy to the fields supervised institutions must report, and carry those fields into provider notifications. 9. State the withdrawal and continuity position: notice periods, migration windows, and evidence retention when a capability or version is withdrawn. 10. Treat sovereignty as a control-plane property and state which commitments hold identically across hosted, sovereign, private and on-premise configurations. 11. Adopt a stated non-retaliation position for good-faith adversarial testing, with a disclosure channel and a response commitment. 12. Publish how the consultation was dispositioned: themes received, changes made, recommendations declined and why. We make no request for proprietary information. Every recommendation above can be satisfied without disclosing model weights, training data or commercially sensitive architecture. We would welcome the opportunity to contribute to any technical working group on evidence formats for behavioural commitments, and we will publish our own response openly so that deployers may reuse its structure. Dax Philbert, LLM Chairman & Chief Executive, Cabier Consulting dax.philbert@cabierconsulting.com
How each point is already operable
We do not raise these points abstractly. Each is a working surface an institution can inspect, built before this consultation opened.
The wider argument sits in Global AI Assurance, and the umbrella itself is at Global AI Assurance.
Questions we have been asked
Why respond to a provider's consultation at all?
Because behavioural commitments made at the model layer become control assumptions at the institutional layer. If a code is drafted so that its commitments cannot be evidenced, supervised deployers inherit the gap. Responding is cheaper than absorbing it.
Is this an endorsement of the draft or of its author?
No. It is a consultation response. Cabier is model-neutral, endorses no provider, publishes no vendor comparison, and does not train, host or resell governed models. We treat published provider expectations as an observed direction of travel and nothing further.
Does a provider code reduce a regulated deployer's obligations?
No. Accountability for an AI-enabled outcome sits with the institution that put the system in front of a customer, a market or a supervisor. A provider code can supply evidence into that case; it cannot discharge it.
What is the single most useful change the draft could make?
Marking each commitment as model-inherent, configuration-dependent or deployer-owned. That one column would prevent most of the misallocated reliance we see in institutional AI programmes.
Are values-based codes a distraction?
Not at all. Stated intent is the necessary first step and sets the direction a model is trained toward. The argument is only that intent needs an assurance layer above it, because a board cannot show a supervisor a value.
Will you publish the response you submit?
It is published on this page in full, unedited, and dated. Institutions may reuse the structure for their own submissions.
References and citations
Primary sources. The consulted draft is referenced as published material only; positions change, so verify at source before relying on any statement.
- 1Microsoft AI, Humanist AI Code of Conduct (draft for public consultation, 14 September 2026) — The document under consultation. Referenced as an observed direction of travel; no endorsement is stated or implied in either direction.Source
- 2Microsoft AI, Humanist AI in practice: a public consultation on our Code of Conduct for MAI Models (14 September 2026) — Announcement of the consultation and its stated review window.Source
- 3Regulation (EU) 2024/1689, Artificial Intelligence Act — General-purpose model obligations, deployer duties, serious-incident reporting and the rejection of self-attestation for material systems.
- 4NIST AI 100-1, AI Risk Management Framework 1.0, with the Generative AI Profile (NIST AI 600-1) — Govern, map, measure, manage functions and the measurement expectations we rely on for testable assertions.
- 5ISO/IEC 42001:2023, Artificial intelligence management system — Management-system expectations for documented control operation rather than stated intent.
- 6Board of Governors of the Federal Reserve System SR 11-7 and OCC Bulletin 2011-12, Supervisory Guidance on Model Risk Management — Independent validation by parties who did not build the model.
- 7Regulation (EU) 2022/2554, Digital Operational Resilience Act — Third-party dependency, concentration, exit and continuity expectations engaged by provider withdrawal.
- 8OWASP Top 10 for Large Language Model Applications — Adversarial testing categories relevant to the safe-harbour recommendation.