Global AI Assurance · Layer 06

    Independent Evaluation Engine

    Model and agent evaluation tooling is improving quickly, and much of it is good. What is usually missing is the institutional wrapper: who commissioned the test, was the tester independent, what disposition followed, who accepted the residual risk, and where the record lives when a supervisor asks two years later.

    Cabier does not compete with evaluation tooling. It governs the evaluation, whichever tool produced the result.

    What gets tested

    System scope

    • Model
    • Agent
    • Workflow
    • Cross-agent interaction

    Reach

    • Tool use
    • Data access
    • Privilege escalation
    • Unauthorised action

    Adversarial

    • Prompt injection
    • Jailbreak
    • Data exfiltration

    Quality and conduct

    • Hallucination
    • Policy violation

    Change over time

    • Model drift
    • Agent drift
    • Emergent behaviour

    Five dispositions

    A test result is not an outcome until someone disposes of it. Each disposition produces a named artefact, which is what makes the assurance defensible.

    D1

    Pass

    Behaviour within the approved envelope across the tested surface.

    Evidence Signed result set with the envelope version it was tested against.

    D2

    Pass with conditions

    Permitted, with a named constraint and a review date attached.

    Evidence Condition register entry with owner, tolerance and expiry.

    D3

    Remediation required

    A control gap is proven. Operation continues only under compensating control.

    Evidence Finding, compensating control and remediation plan with dates.

    D4

    Human review

    The result is not machine-decidable. A named person disposes of it.

    Evidence Review record with the reviewer, the reasoning and the outcome.

    D5

    Block

    The system or action is withdrawn until the finding is closed.

    Evidence Block record, the authority who imposed it and the release condition.

    Independence conditions

    Commissioned outside the build

    The team that built the agent does not set the test scope or accept the result.

    Envelope-aware

    Tests are run against the declared authority envelope, so a pass means something specific.

    Version-bound

    Every result is bound to the model and agent version it was produced against.

    Consequential

    A finding can block release. An evaluation that never blocks anything is documentation.