The framework
Seven evaluation dimensions.
The unit of assessment is a claim — together with the evidence offered for it and the decision it supports — not "the model" and not "the vendor's report". Three dimensions establish whether the figure means anything, two whether it transfers to your deployment, and two whether the conditions of its production permit it to be believed at all.
Does the number mean anything?
Construct Validity
Does the evaluation measure the capability it claims to measure? Construct definition, content coverage, metric appropriateness, contamination and leakage, shortcuts and gaming, saturation, and how much prompt engineering and scaffolding produced the result.
Does the number mean anything?
Evidence Integrity
Is the reported result a faithful record of what happened? Per-trial traces, run accounting, exclusions and when their criteria were set, exact system version, configuration capture, scoring integrity, and chain of custody.
Does the number mean anything?
Statistical Reliability
Would the result survive repetition? Trials per item, sample composition, uncertainty intervals, paired comparison, clustering, prompt and seed sensitivity, multiplicity, calibration, and whether error types with unequal cost are reported separately.
Does it transfer to deployment?
Production Realism
Does the evaluation represent the environment in which the system will operate? Security policy, egress restrictions, real identity boundaries, actual tools rather than stubs, rate limits, latency budgets, production retrieval, real document quality, and human approval gates.
Does it transfer to deployment?
Failure Characterisation
How does the system fail, not merely how often does it succeed? Every failure mode is recorded on four axes — frequency, severity, detectability, reversibility. Detectability and reversibility are what convert a failure rate into business risk.
May it be believed at all?
Evaluation Independence
Who designed, conducted, interpreted, and approved the evaluation? Graded I0 to I3, on conditions rather than job titles. An external evaluator working to a vendor-set scope, on vendor-supplied data, under a vendor publication veto is not independent in any useful sense.
May it be believed at all?
Evaluation Environment Integrity
Was the environment that produced the result itself trustworthy? Sequestration of test data, network isolation, simulation fidelity, credential hygiene, independent logging — and the question that decides most agentic evaluations: could the evaluated system influence its own score? Where an agent can write files a grader later reads or reach the reference solutions, a recorded success may represent manipulation of the scorer rather than completion of the task.