HomeServicesAgentic AI Containment Assurance

Independent technical assurance

If this AI agent behaves unexpectedly, what can it do before you stop it?

AI agents do not just answer — they act. They run code, call APIs, modify systems, and carry out multi-step work with real permissions. Agentic AI Containment Assurance is an independent, tested answer to one question: what can this agent reach, change, disclose, authorize, or damage before your organization can stop it? Containment is not an architectural claim — it is a security property that has to be demonstrated.

Design intent is not demonstrated control.

When software can act, the security question changes. It is no longer whether the model usually follows instructions — it is what the complete deployed system can do when the model is manipulated by untrusted content, encounters a misconfigured tool, or pursues an objective beyond its intended boundary.

Labels are not evidence

"Sandboxed," "human-in-the-loop," and "read-only" are useful design intentions. A sandbox may still have an egress path, a nominally read-only tool may invoke a side-effecting API, and an approval step may take longer than the agent needs to create an irreversible consequence.

The unit of assurance is the deployed system

We assess the agent together with its orchestration, system instructions, tools, identities, credentials, network paths, memory, retrieval sources, data stores, logging, approval mechanisms, infrastructure, and external services — not the foundation model in isolation.

No single safety score

Averages conceal catastrophic weaknesses. Strong logging does not compensate for unrestricted production credentials, and good model behaviour during a demonstration does not compensate for an ineffective shutdown mechanism. Findings are reported by dimension.

Six containment dimensions.

Four exposure dimensions describe what the agent could do — its consequence space. Two containment-capability dimensions describe whether you can see that behaviour and stop it. The distinction matters: exposure tells you what is at stake, containment capability tells you whether oversight is operationally real.

Exposure

Authority

What actions can the agent take? From read and recommend through write, execute, externalize, and authorize. Assurance documents both intended authority and technically available authority — any difference is a finding.

Exposure

Connectivity

Where can the agent reach? Browsers, APIs, webhooks, package repositories, cloud metadata endpoints, messaging, MCP servers, and agent-to-agent channels. Restrictions are tested, not inferred from configuration documentation.

Exposure

Privilege

What authority can credentials convert into action? Service accounts, API keys, tokens, cloud roles, repository rights, and administrative functions — including secrets discoverable through environment, logs, volumes, or metadata services.

Exposure

Persistence

Can the effect survive the session? Files, configuration, scheduled tasks, accounts, tokens, repository changes, memory entries, callbacks, deployed workloads, and altered startup behaviour. Ending a session does not prove containment.

Containment capability

Observability

Can consequential behaviour be reconstructed? Independent, attributable records of tool use, authorization decisions, and external communication — held outside the agent's control, where they cannot be suppressed, altered, or bypassed.

Containment capability

Intervention

Can the organization stop the agent in time? Session termination, credential revocation, network isolation, process termination, tool disablement, and deterministic policy enforcement — measured against how fast the consequence arrives.

Control principle: the model should not be the final authority on its own permissions, approval requirements, compliance status, logging, or right to continue operating.

Three executive-facing metrics.

The assessment converts technical testing into measures a board can act on — without collapsing distinct weaknesses into a single average.

Agentic Exposure Profile

Each exposure dimension is scored 0–4 and reported as a profile rather than an average, so the shape of exposure is preserved. Significant authority combined with significant privilege is the combination that turns a model error into an operational consequence.

Maximum Credible Agent Impact

MCAI (A0 Informational through A4 Critical) answers the executive question: what is the most consequential credible outcome this agent could produce before existing controls detect and stop it? It is rated on technically available pathways, not average routine behaviour.

Containment Margin

Intervention Time (detect + decide + enforce) is compared with Damage Time (the shortest credible path to a material consequence). A margin below 1 means the consequence lands before intervention completes — and monitoring faster is usually not the fix.

A worked example of why the margin matters: suppose testing shows that a manipulated input can drive an agent to overwrite a shared record in 20 seconds, while your alert-to-intervention path takes 45 seconds. The Containment Margin is 20 / 45 ≈ 0.44 — the damage is done before anyone can act. For that action class the remediation is not faster monitoring; prevention has to move ahead of the model, through versioned writes, constrained destinations, or deterministic approval before the action executes.

Conclusions are graded by the quality of evidence behind them.

A design diagram, a policy statement, or a vendor questionnaire is not equivalent to an independently challenged control.

Evidence hierarchy

  • E0 — Assertion: a developer or vendor states that a control exists
  • E1 — Documentation: architecture or configuration supports the claim
  • E2 — Inspection: implementation is directly examined
  • E3 — Test: the control is actively challenged under defined conditions
  • E4 — Operational: records show the control working in real conditions

Control effectiveness

  • CE0 — Absent: no meaningful control exists
  • CE1 — Documented: present on paper, not demonstrated
  • CE2 — Implemented: technically present, never independently challenged
  • CE3 — Tested: resisted defined independent test conditions
  • CE4 — Continuously enforced: monitored, revalidated, and change-managed

The critical control rule

  • Some controls are too important to average
  • Failure of a critical control produces a Critical Finding regardless of the broader score
  • Inability to reliably terminate the agent, agent-controlled logging, or unrestricted egress with privileged credentials are examples
  • Strengths elsewhere cannot mathematically conceal a severe weakness

Repeatability matters: agent behaviour is probabilistic. A single successful run does not prove reliability, and a single failed adversarial attempt does not prove security. Critical scenarios are repeated, and configuration, prompt variation, repetition count, and observed success frequency are recorded.

A seven-phase assessment.

Scope and duration are scaled to the agent's authority and the sensitivity of the environment. All adversarial testing is authorized, bounded, and agreed in writing before it begins.

1 · Characterize

Define the agent, tools, identities, data, network, environment, and people. Output: assurance boundary and agent authority map.

2 · Threat model

Identify credible manipulation, misuse, failure, and boundary-crossing scenarios. Output: a prioritized scenario register.

3 · Review evidence

Inspect architecture, IAM, network rules, secrets handling, approval flows, logs, and incident processes. Output: control evidence inventory.

4 · Challenge controls

Conduct authorized adversarial and failure-mode testing — including prompt injection, tool misuse, credential discovery, and approval bypass.

5 · Repeat critical tests

Measure behavioural variability and alternate pathways across repetitions. Output: repeatability results.

6 · Measure intervention

Test stop mechanisms and scenario timing. Output: Intervention Time, Damage Time, and Containment Margin.

7 · Conclude & remediate

Classify findings and decide deployment posture. Output: an executive report with a clear deployment recommendation, and a sequenced remediation roadmap that converts findings into an implementation plan.

Decision-ready deliverables.

Executive outputs

  • Executive Assurance Report — MCAI, critical findings, intervention analysis, and a deployment recommendation
  • Board-level view: what the agent is allowed to do, what it can actually do, the worst credible consequence, and how quickly it can be stopped
  • 90-Day Remediation Roadmap sequencing the work

Technical outputs

  • Agent Authority Map of every material tool and resource connection
  • Agentic Exposure Profile across authority, connectivity, privilege, and persistence
  • Technical Containment Test Report with conditions, evidence, repetitions, and results
  • Control Gap Register prioritized by severity, evidence, owner, and target state

Who it's for

  • Organizations moving an agent from pilot into production
  • Buyers who need a vendor's autonomy and containment claims independently verified
  • Regulated and public-sector teams accountable for delegated machine authority
  • Security and risk functions integrating agentic AI into an existing control model

Why it matters for executives: autonomy is delegated operational authority. Governance should scale with the authority delegated to the system and the irreversibility of its potential effects — a text summarizer and an internet-connected code-execution agent should not face the same approval process.

Conditions that should stop a deployment.

Within the assessed scope, Maple Quanta would normally recommend against production deployment while any of the following is present. Exceptions should require explicit, documented risk acceptance by an accountable executive, with residual risk and compensating controls stated.

Authority and privilege

  • An unresolved Critical finding exists
  • High-impact actions lack independent authorization
  • The agent can exercise materially broader privilege than necessary

Reach and persistence

  • Unrestricted external connectivity without a justified requirement and compensating controls
  • Cross-tenant or security-boundary access is possible
  • Unauthorized persistence has been demonstrated

Oversight and timing

  • Consequential activity cannot be independently logged and reconstructed
  • Effective shutdown has not been demonstrated
  • For a critical scenario, intervention is slower than the credible damage pathway

Ten questions to ask before procuring an agent.

Procurement should treat autonomy as delegated operational authority. Before an agent is bought or approved, ask for evidence — not assurances — on these ten questions. The assessment exists to answer them independently.

Authority and reach

  • What actions can the agent perform without human approval?
  • Which internal and external systems can it reach?
  • Which identities, credentials, tokens, and administrative roles can it exercise?
  • Can it create durable effects — accounts, tasks, resources, configuration changes?

Independent control

  • Which consequential actions are governed by deterministic controls outside the model?
  • Can every high-impact tool call be reconstructed from logs the agent cannot alter?
  • How is a session, credential, tool, or network path disabled during an incident?

Evidence and reassessment

  • What has actually been tested, by whom, under what configuration, and with how many repetitions?
  • What changes trigger mandatory reassessment?
  • What is the Maximum Credible Agent Impact under the proposed production configuration?

If the answers are assurances rather than evidence, that is exactly the gap this assessment closes — independently, before the contract is signed or the agent reaches production.

Complementary to established guidance — and explicit about what it does not claim.

Designed to complement

The framework is built to sit alongside the NIST AI Risk Management Framework and TEVV work, Canadian Centre for Cyber Security and international guidance on the careful adoption of agentic AI, OWASP's agentic security material, and ISO/IEC 42001, 23894, 42005, and 42006. Its contribution is to translate those principles into a compact system-containment model with measurable, executive-facing outputs.

What it does not claim

It is not a certification scheme and does not certify conformity with ISO/IEC standards. It does not prove that an AI system cannot fail, and it does not replace a formal penetration test, privacy assessment, safety case, or regulatory assessment where those are required. Conclusions are configuration-specific: they apply to the model version, prompts, tools, permissions, environment, and controls assessed, and material changes trigger retesting.

The full methodology — framework, metrics, evidence model, and decision rules — is published in the Agentic AI Containment Assurance white paper (Version 1.1, August 2026). Detailed adversarial test scripts, scoring calibration, assessor procedures, and client-specific tooling remain part of the engagement rather than the public document.

What happens when the agent does not behave?

Control is not the claim that nothing will go wrong. It is the demonstrated ability to intervene, exit, and recover when it does.