HomeServicesIndependent AI Evaluation Assurance

Independent technical assurance

Does the evidence behind this AI system prove what it claims to prove?

An AI benchmark tells you what happened in a test. Evaluation assurance determines whether that test deserves to influence a real decision. Maple Quanta independently assesses whether the evidence supporting an AI claim — benchmark scores, vendor evaluations, pilots, system cards, red-team reports — is valid, statistically credible, reproducible, representative of production, sufficiently documented, and independently defensible before it becomes a procurement, deployment, or governance decision.

A score is not the same thing as evidence.

Organisations now approve AI systems on the strength of benchmark results, vendor evaluations, leaderboard positions, and pilot outcomes. Those numbers have acquired the weight of evidence without acquiring the properties of evidence — and the failure modes are structural rather than dishonest.

The test may not measure the claim

Contamination of public benchmarks is documented as widespread. Tasks can be completed by retrieving a published solution rather than by exercising the claimed capability. And a bilingual claim supported by a test set that is 94% English is not a bilingual result, whatever the aggregate says.

Conditions are part of the result

A figure produced with unrestricted network access, stubbed tools, clean documents, and no approval gates does not describe the system that will run inside your security policy. In our experience the difference is routinely double digits — and almost never measured.

Averages conceal the failures that matter

A 97% accurate system with a 3% severe, silent, irreversible error rate is not the same product as one with a 3% obvious, recoverable error rate. No aggregate metric distinguishes them, and the second is the one that reaches a board.

The core thesis: a benchmark result is evidence only to the extent that the evaluation producing it is valid, reproducible, relevant to production, and independently interpretable. Assurance addresses the evaluation, not the marketing.

Seven evaluation dimensions.

The unit of assessment is a claim — together with the evidence offered for it and the decision it supports — not "the model" and not "the vendor's report". Three dimensions establish whether the figure means anything, two whether it transfers to your deployment, and two whether the conditions of its production permit it to be believed at all.

Does the number mean anything?

Construct Validity

Does the evaluation measure the capability it claims to measure? Construct definition, content coverage, metric appropriateness, contamination and leakage, shortcuts and gaming, saturation, and how much prompt engineering and scaffolding produced the result.

Does the number mean anything?

Evidence Integrity

Is the reported result a faithful record of what happened? Per-trial traces, run accounting, exclusions and when their criteria were set, exact system version, configuration capture, scoring integrity, and chain of custody.

Does the number mean anything?

Statistical Reliability

Would the result survive repetition? Trials per item, sample composition, uncertainty intervals, paired comparison, clustering, prompt and seed sensitivity, multiplicity, calibration, and whether error types with unequal cost are reported separately.

Does it transfer to deployment?

Production Realism

Does the evaluation represent the environment in which the system will operate? Security policy, egress restrictions, real identity boundaries, actual tools rather than stubs, rate limits, latency budgets, production retrieval, real document quality, and human approval gates.

Does it transfer to deployment?

Failure Characterisation

How does the system fail, not merely how often does it succeed? Every failure mode is recorded on four axes — frequency, severity, detectability, reversibility. Detectability and reversibility are what convert a failure rate into business risk.

May it be believed at all?

Evaluation Independence

Who designed, conducted, interpreted, and approved the evaluation? Graded I0 to I3, on conditions rather than job titles. An external evaluator working to a vendor-set scope, on vendor-supplied data, under a vendor publication veto is not independent in any useful sense.

May it be believed at all?

Evaluation Environment Integrity

Was the environment that produced the result itself trustworthy? Sequestration of test data, network isolation, simulation fidelity, credential hygiene, independent logging — and the question that decides most agentic evaluations: could the evaluated system influence its own score? Where an agent can write files a grader later reads or reach the reference solutions, a recorded success may represent manipulation of the scorer rather than completion of the task.

No composite score. Findings are reported by dimension, and no weighted index is published. A single number would require weights that cannot be justified from evidence, and its only function would be to let a strength on one dimension conceal a fatal weakness on another. Impeccable statistics do not repair a contaminated benchmark.

Five measures a board can act on.

Each is defined so that two competent assessors given the same inputs produce the same output — and so that the conclusion can be recomputed from its inputs by anyone holding the record.

Evaluation Confidence Level

C0 Unsubstantiated through C4 Robustly Verified, assigned per claim by an explicit floor rule rather than assessor judgement. Failing a precondition does not subtract points — it caps the level, and the report names the precondition that capped it.

Production Evaluation Gap

The difference between Laboratory Capability and Operational Capability, in the metric's own units, with an interval on the difference. Computed only when five commensurability preconditions hold — otherwise reported as not computable, with the reason.

Maximum Credible Evaluation Failure

The most consequential failure the evidence demonstrates or fails to exclude. An evaluation that never tested a failure mode provides no evidence against it, so Not excluded is a reportable status — and often the most important line in the report.

Evaluation Reproduction Agreement

Whether independent re-execution produced the same answer: Reproduced, Partially reproduced, Not reproduced, or Not reproducible with the reason recorded. Categorical, never a "reproducibility percentage" — one decisive failure is not offset by many trivial successes.

Evidence Coverage

How many material claims have examinable evidence behind them, reported as n of N. Materiality is fixed before assessment so the denominator cannot move once results are known. Coverage measures availability of evidence, never whether the claims held.

Four findings, kept distinct

Substantiated · Partially Substantiated · Unsupported · Contradicted. Unsupported means no evidence establishes the claim; Contradicted means testing produced inconsistent results. The first calls for evidence, the second for explanation.

A worked illustration of why commensurability matters: suppose a vendor reports 93% accuracy, and independent testing under your production controls returns 76%. "93% became 76%" is a powerful sentence and an unsound comparison — different items, undisclosed exclusions, an unidentified model version. The defensible finding measures the same items and the same build under both conditions and reports, say, a 12-point gap with a confidence interval that excludes zero. That is the number that survives a vendor challenge.

Conclusions are graded by the quality of evidence behind them.

A vendor report, an inspected trace, a reproduced result, and a methodology that has survived adversarial challenge are four different kinds of thing.

Evidence hierarchy

  • E0 — Assertion: stated, with no supporting artifact
  • E1 — Documentation: a written description of method or result exists
  • E2 — Artifact: traces, datasets, configurations, and scoring code inspected
  • E3 — Reproduced: independently re-executed with a consistent result
  • E4 — Adversarially Validated: the methodology itself was challenged and survived

Two rules that make it work

  • The grade is the weakest link. A reproduction resting on a dataset whose provenance is only asserted is an E0 finding — assurance does not average
  • The ladder is not always climbable. Where a rung is unreachable we record why, and "vendor declined re-execution access" is frequently more consequential than the rating it caps
  • Vendor evaluation material is E1 at best on its own, however detailed the document

The critical-concern rule

  • Some defects are too consequential to average away
  • A manipulable scoring mechanism, demonstrated contamination of the decisive data, undisclosed exclusions, or an unidentifiable model version each produce a Critical Concern regardless of strengths elsewhere
  • Strengths cannot mathematically conceal a severe weakness, because there is no arithmetic in which to conceal it

Why E4 matters most: E3 answers "is the number real?" E4 answers "does the number mean what it is being used to mean?" A contaminated benchmark, a manipulable grader, and an over-elicited configuration all reproduce faithfully. E4 is the rung most often missing, and it is where the consequential errors live.

An eight-phase assessment.

Depth is scaled to the consequence of the decision the evidence supports. All testing against a live system is authorised, bounded, and agreed in writing before it begins.

1 · Claims inventory

Record every material claim verbatim, with the decision it supports. Materiality is fixed here, before assessment, so the denominator cannot move later.

2 · Evaluation reconstruction

Establish what was actually done — version, data, prompts, tools, scoring, repetitions, exclusions — as distinct from what was reported.

3 · Validity audit

Test for contamination, leakage, shortcuts, gaming, unsuitable metrics, inappropriate baselines, saturation, and a manipulable scoring boundary.

4 · Independent reproduction

Re-run key tests with repeated trials, varied seeds, perturbed prompts, and sealed evaluation material the vendor has never seen.

5 · Production-constrained evaluation

Re-execute material evaluations under your real security, identity, latency, retrieval, privacy, and approval constraints. Output: the Production Evaluation Gap.

6 · Failure stress testing

Introduce missing information, tool failures, noisy and adversarial inputs, rate limiting, and context variation. Output: the Failure Register and MCEF.

7 · Evidence assessment

Determine which claims are substantiated, partially substantiated, unsupported, or contradicted — and assign the Evaluation Confidence Level for each.

8 · Executive report

Findings, evidence, the scorecard, remediation priorities sequenced Immediate / 30 / 90 days, and an Executive Assurance Statement suitable for a procurement record — carrying its scope and its expiry conditions in the same paragraph as its conclusion, so it cannot be quoted without them.

Three tiers, distinguished by what we execute ourselves.

What Maple Quanta runs determines the evidence grade the engagement can reach — and therefore the confidence level any claim can attain. The ceilings are structural, not a matter of effort.

Evaluation Evidence Review

  • For organisations that already hold evaluation reports and need an independent read before acting
  • Methodology and construct-validity review, contamination and gaming analysis, statistical review, evidence sufficiency, reporting quality
  • No re-execution — evidence caps at E2, confidence at C2
  • Frequently the highest value per dollar: undisclosed exclusions, mismatched versions, and contaminated benchmarks are found by reading artifacts, not by testing

Independent Evaluation

  • Maple Quanta independently tests the system
  • Sealed private evaluation set the vendor never sees, repeated trials, seed and prompt variation, alternate datasets
  • Production-constrained re-execution with the Production Evaluation Gap and constraint attribution
  • Reaches E3 and C3 — answers whether the capability exists when the system has not seen the test

High-Assurance Evaluation

  • For agentic systems, government systems, regulated use cases, and safety-critical or high-impact deployments
  • Adds evaluation environment assurance, scoring trust-boundary testing, full failure stress testing, production simulation, and evaluation governance review
  • Strengthened evidence preservation with hashing, timestamping, and chain of custody
  • The only tier that reaches E4 and C4 — for specific claims, not for the system

A note on our own independence: Maple Quanta grades its engagements on the same I0–I3 scale it applies to others, and states the grade in the report along with who commissioned and paid for the work. Our engagements are typically I2 — third party, not independently replicated. An assurance provider that grades independence should state its own.

The highest-leverage moment is before the RFP goes out.

Consider a department receiving proposals from three vendors, each claiming superior AI performance on a different benchmark, under different conditions, with different exclusions. The evidence is not comparable — and once the bids have arrived, no amount of downstream analysis fully recovers the comparison.

Pre-procurement

  • Acceptance criteria expressed as testable claims rather than adjectives
  • The evaluation methodology bidders must follow
  • Minimum evidence and disclosure requirements
  • A private evaluation set built from your own documents and never released
  • The conditions under which vendor-supplied evidence will and will not be accepted

During procurement

  • Independent validation of each bidder's claims against the same standard
  • Per-claim findings and confidence levels that can be compared across bids
  • Answers the question a procurement officer actually has: are these three claims comparable, and which are supported?

Acceptance and beyond

  • Final production-constrained evaluation against your real security policies, bilingual requirements, latency limits, privacy controls, document formats, user workflows, and approval gates
  • Periodic reassessment on defined triggers: model or endpoint change, configuration or corpus change, control change, scope expansion, or an incident bearing on an assessed claim

Who this is for: Government of Canada departments, provincial and municipal governments, regulated industries, financial institutions, healthcare organisations, critical infrastructure operators, organisations procuring foundation models or deploying AI agents, technology vendors seeking independent assessment of their own evidence, and boards accepting residual risk before a production deployment.

Ten questions to ask before trusting an AI benchmark.

Ask for evidence, not assurance. A confident answer with nothing behind it is the first finding. The assessment exists to answer these independently.

What was tested

  • Exactly which version was tested — and is it the version we are buying?
  • Could the test material have been in the training data?
  • Was the benchmark chosen before or after the results were seen?
  • How many independent runs, and what is the uncertainty?

How it was run

  • What was excluded — and who set the criterion, and when?
  • How much help did the system get: prompt tuning, scaffolding, retries, human involvement?
  • Could the system have influenced its own score?

What it means for us

  • Was it tested under our constraints, or in a laboratory?
  • How does it fail — and what failure modes were not tested for?
  • Who ran it, who paid for it, and can anyone else reproduce it?

If the answers are assurances rather than evidence, that is not necessarily a reason to reject a vendor — it is a reason not to let the number carry the weight of the decision. The full guide, with what a strong answer looks like for each question and what a weak one should change about the decision, is published free alongside the methodology.

Evaluation Assurance asks

Can we trust the evidence that this system performs as claimed? The object is a claim and the evidence for it. Failure means the number cannot be relied on.

Containment Assurance asks

Can we bound what this system can do when it behaves unexpectedly? The object is a deployed system and its controls. Failure means the consequence cannot be bounded. See the containment framework.

The two are independent. A capable system can pass Evaluation Assurance and fail Containment Assurance — its performance claims are well-evidenced and it can still reach systems it should not. A well-contained system can pass Containment Assurance and fail Evaluation Assurance — nothing it does is dangerous, and there is no credible evidence it does the job. An agentic system entering production in a regulated context generally needs both: one establishes whether the performance justifying the deployment is real, the other whether the organisation survives the cases where it is not.

And when something goes wrong: an ordinary evaluation failure is not an AI incident — evaluations exist to produce failures. But an evaluation that reaches a real system, or runs with credentials that could touch production, is both an evaluation finding and a classifiable event under MQ-AICIR, our AI incident classification framework. That closes the lifecycle: evaluate, deploy, monitor, detect, classify, respond, re-evaluate.

Complementary to established guidance — and explicit about what it does not claim.

Designed to complement

The framework is built to sit alongside the NIST AI Risk Management Framework and its Generative AI Profile, NIST's TEVV programme — including the AITE sequestered-evaluation vehicle and the draft TEVV-Athlon framework — the Canadian AI Safety Institute's guidance on what AI evaluators should disclose, the international network's evaluation best practice, and ISO/IEC 42001, 23894, 42005, 25059, TS 4213, and 24029. Its contribution is a method for interrogating an evaluation result that has already been produced and is now being used to justify a decision.

What it does not claim

It is not a certification scheme, and Maple Quanta is not an accredited certification body — ISO/IEC 42006:2025 governs bodies that certify AI management systems, and assurance and certification are different instruments. It does not certify conformity with any standard, does not claim alignment with draft standards, and does not replace a safety case, privacy assessment, penetration test, or regulatory assessment. Conclusions are configuration-specific and are invalidated by material change to the assessed version, configuration, environment, or controls.

The full methodology — dimensions, evidence model, metrics, decision rules, and a complete synthetic worked example — is published in the Independent AI Evaluation Assurance white paper (Version 1.0, August 2026), with machine-readable schemas, the free Evaluation Disclosure Sheet, and tested reference implementations in the open MQ-EVA repository. Detailed contamination and gaming test procedures, scoring calibration, assessor working papers, and sealed evaluation material remain part of the engagement rather than the public document.

The question is not whether the AI passed the test.

It is whether the test deserves to be trusted — before the result becomes a procurement, deployment, or governance decision.