Trust in AI is behavioural evidence, not code review

Share
Trust in AI is behavioural evidence, not code review

TL;DR: You can't trust an AI system by reading its documentation. Trust in AI comes from observed behaviour, not policy papers. The only artifact that proves your governance model works is a refusal log. Most regulated firms don't have one.

  • Agentic AI takes actions in production. The failure mode is behaviour that's already happened.
  • Regulators are now asking for demonstrated behaviour, not documentation.
  • A refusal log is the only evidence that policy is enforced, not just written.
  • Most governance failures trace to three missing properties: independence, enforceability, and evidence.
  • A CIO can diagnose this before Friday with five specific checks.

You can't trust an AI system by reading its documentation. You can only trust it by watching what it does when the instruction is wrong.

Most boards don't know this yet. Most CIOs do. The space between the two is where the next generation of regulatory findings will land.

Agentic AI systems in 2026 don't just say things. They take actions. They draft transfer instructions, update contract terms, file compliance reports, and touch client records. The failure mode is no longer an output somebody reviews before it goes anywhere. The failure mode is behaviour that's already happened by the time anyone looks.

That single shift, from generative to agentic, breaks the trust model most regulated businesses inherited from the last software cycle. You used to trust a system because you read the code, tested it in a sandbox, and shipped it once the specification matched the output. That trust model assumes the system does one thing and does it in one place.

An agentic system does thousands of things across dozens of contexts, and the specification is a policy document written in English. There is no code review that certifies English against behaviour. There is only observation.

The trust asymmetry regulators have already noticed

Trust in AI is built the same way trust in a person is built: through repeated, observable behaviour that demonstrates adherence to principles under pressure. A single violation of a stated principle destroys trust faster than a thousand successful interactions build it.

Regulators know this. APRA, ASIC, the SEC, and the FCA have all issued guidance in the last 18 months that shifts the burden of proof from documentation to demonstrated behaviour. The vocabulary varies. The mechanism is identical: show that the system does what the policy says, over time, under pressure, in production.

The organisations that treat governance as infrastructure, something you build once and operate continuously, are the ones that move fast in regulated markets. The ones that treat it as documentation theatre are stuck in pilot purgatory, watching competitors deploy while their own steering committee schedules another review.

Grant Thornton's 2026 governance survey put numbers on the gap. 62% of financial services firms reported an AI governance policy in place. 58% could not produce a refusal log for that same policy. 52% had no independent enforcement layer between the model and the policy at all.

Three numbers, three angles, one failure. Policy exists on paper. Behaviour is not measured. Enforcement is not architecturally independent. In every case, the trust model is documentation. In every case, it will fail the first regulator who asks the right question.

Key point: Regulators have moved past policy documents. They want evidence of behaviour under pressure. If your governance model is documentation-first, it's already behind.

What refusal actually proves

A refusal is the only observable evidence your governance model works.

Not the policy document. That proves the policy exists on paper. Not the vendor's compliance deck. That proves the vendor has a marketing department. Not the training completion rate. That proves people clicked through a slide.

The refusal.

The moment the system was instructed to do something the policy forbids and the system said no. That's the artifact. It's the only artifact that survives contact with a regulator's opening question, which is always some version of: show me a case where the system stopped itself.

If the log is empty, three possibilities exist. Nobody has ever attempted an out-of-policy action, which isn't credible in any organisation with more than 50 employees. The system attempted the action and completed it, which means the policy is decoration. Or the system refused but nobody logged it, which means the audit trail doesn't exist.

Two of those three findings end the same conversation with a regulator.

Key point: A refusal log is the minimum viable proof of governance. Without it, you're not enforcing policy. You're hoping the system behaves.

The architecture that makes refusal possible

Not every AI deployment can refuse. Most can't. The reason is architectural, not philosophical.

For a system to refuse an instruction, three properties have to hold at the same time.

Independence. The policy layer must be separate from the model. If the model and the policy live in the same context window, the model can rewrite the policy on the fly. Any user with a moderately clever prompt can dissolve the guardrail. The policy has to sit in a layer the model can't reach, evaluated by a component the model can't influence.

Enforceability. The policy has to be expressed in a form a machine can evaluate deterministically, not just interpret. A policy that says "the system should be helpful but not harmful" is not enforceable. A policy that says "the system will not draft a transfer instruction over $50,000 without a named human approver in the loop" is. Governance in production is a translation problem before it is a control problem.

Evidence. Every refusal has to write an entry to a log the operator can't modify after the fact. Hash-chained records, immutable storage, timestamped, attributable to the specific instruction that triggered them. Without this, the refusal happened in private, and private refusals are indistinguishable from no refusals at all.

Three properties. Independence, enforceability, evidence. Missing any one and the system is not enforcing policy. It's performing it.

This is the difference between a CRO who can answer the regulator's question in 30 seconds and one who needs a week. The CRO with the architecture pulls the log. The CRO without it hires a firm to write a report.

Key point: The three properties aren't aspirational design goals. They're the minimum structural requirements for a governance layer that holds under scrutiny.

What a CIO can check this week

None of this is theoretical. Every claim above resolves into a specific action a CIO can take before Friday.

  1. Pull the refusal log for the last 90 days. If the log doesn't exist, that's the finding. Stop the review, go build the log, come back when it's running.
  2. Count refusals per 1,000 interactions. Under 0.1% and the policy layer is inert. Between 0.1% and 1% is the plausible operating range for a well-tuned enforcement layer. Above 5% and the policy is too broad, refusing benign work.
  3. Read the top 10 refusal reasons. If they cluster on format errors and rate limits, the policy layer is checking syntax, not intent. That's input validation dressed up in governance vocabulary.
  4. Find one refusal that cost the business money. A blocked transaction, a delayed contract, a client escalation that resolved to "the system was right to stop us." If none exist, the policy has never held under pressure. That means the trust in it is faith.
  5. Ask who authored the policy the AI enforced. If the answer is the vendor, the vendor governs your business. That's a procurement finding, not a technology finding, and it's the single most under-discussed governance risk in regulated AI deployment in 2026.

Five checks. None require a consultant. All fit inside a Thursday afternoon.

Key point: Governance is observable. If you can't answer these five questions from existing data, the architecture isn't there yet.

The board decision this frames

Governance in a mature regulated business is a monthly board number, not a quarterly policy review. Four numbers deserve the seat: production AI systems touching material things, refusals logged, vendor overrides identified, and decisions reviewable end to end.

Miss any one and the board is not governing AI. It's watching a slideshow about AI.

The organisations that will win the next five years in regulated markets are not the ones with the most sophisticated models. Those are commoditising. The winners are the ones who built the substrate underneath the models: the policy layer, the enforcement layer, the evidence layer, the review layer. That substrate is boring. It is expensive to build. It is not a marketing story.

It is the only thing a regulator, a client, or a board will trust when the story stops being enough.

Look at your AI system's refusal log this week. If it's empty, you don't have governance. You have hope.

Trust is behavioural evidence, not code review.

Key point: The board metrics that matter aren't model performance scores. They're the four numbers that prove the enforcement substrate is running.


Frequently asked questions

What is a refusal log in AI governance?
A refusal log is a tamper-evident record of every instance where an AI system declined to complete an action because it conflicted with the governing policy. It's the primary evidence a regulator will ask for when assessing whether policy is enforced in production, not just written on paper.

Why is a policy document not enough to prove AI governance?
A policy document proves the policy was written. It doesn't prove the system follows it. Regulators need demonstrated behaviour under pressure, which means logged refusals, independent enforcement layers, and immutable audit trails.

What's the difference between a generative and an agentic AI system from a governance perspective?
A generative system produces outputs a human reviews before acting. An agentic system takes actions directly, drafting contracts, filing reports, touching client records. The failure mode shifts from a reviewable output to a completed action. That shift breaks the governance models most regulated firms inherited from the last software cycle.

What is an independent policy layer in AI architecture?
An independent policy layer sits outside the model's context window. It evaluates instructions against policy before the model acts, using a component the model can't influence or rewrite. Without independence, a sufficiently clever prompt can dissolve any guardrail the model holds internally.

What does a healthy refusal rate look like?
Between 0.1% and 1% of interactions per 1,000 is the plausible range for a well-tuned enforcement layer. Below 0.1% suggests the policy layer is inert. Above 5% suggests the policy is miscalibrated and blocking legitimate work.

Who should author the AI policy a system enforces?
The organisation deploying the system, not the vendor. If the vendor authored the policy, the vendor governs your business. That's a procurement and governance risk, not a technology one. It's also one of the most under-discussed failure modes in regulated AI deployment.

How do regulators like APRA, ASIC, the SEC, and the FCA assess AI governance?
All four have shifted the burden of proof from documentation to demonstrated behaviour in guidance issued over the last 18 months. The mechanism is consistent: show the system does what the policy says, in production, under pressure, with an audit trail that can't be modified after the fact.

What are the three architectural properties required for enforceable AI governance?
Independence (policy layer separate from the model), enforceability (policy expressed in deterministic, machine-evaluable form), and evidence (immutable, hash-chained refusal logs). Missing any one means the system is performing policy, not enforcing it.


Key takeaways

  • Agentic AI takes actions in production. The governance model that applies to generative systems doesn't hold.
  • A refusal log is the only artifact that proves policy is enforced, not just written.
  • Three architectural properties are required: independence, enforceability, and evidence. All three, not two.
  • 62% of financial services firms have a governance policy. 58% can't produce a refusal log for it.
  • If the vendor authored your AI policy, the vendor governs your business.
  • The board metric is four numbers: systems in production, refusals logged, vendor overrides identified, and decisions reviewable end to end.
  • The organisations that win in regulated markets over the next five years are the ones who built the enforcement substrate, not the ones with the most capable models.

Read more