a benchmark is not a safety case
The missing argument between passing an evaluation and deserving authority in the world.
there is a moment in almost every assurance story when a measurement quietly becomes a permission. the system performed well on the test. the graph moved in the preferred direction. the red team found fewer failures under the chosen conditions. then, somewhere between the technical appendix and the deployment announcement, a result about a particular experiment becomes a claim about what the system can be trusted to do in the world. that transition deserves more scrutiny than the score itself. it is where the assumptions disappear and the authority arrives.
a benchmark answers a question defined by its construction. who chose the tasks, what resources the model had, how many attempts it received, what counted as success, what the evaluator could observe. a safety case has to connect evidence to a bounded claim about a deployed system. it must explain why the tests bear on the harm, why the safeguards will operate in that environment, which failure paths remain open, and what would invalidate the argument. a test result can be part of that case. it cannot silently supply the rest of it.
the dangerous-capability evaluations introduced by Phuong and colleagues in 2024 illustrate the need for that precision. they examined particular models across persuasion and deception, cybersecurity, self-proliferation and self-reasoning. they reported no evidence of strong dangerous capabilities in those tested models, while flagging early warning signs. that is a bounded empirical finding. its boundaries are part of its scientific value. erasing them would make the result less informative, not more reassuring. Evaluating Frontier Models for Dangerous Capabilities ↗.
now consider the harder question a forming readiness lab should investigate: how much can a negative evaluation tell us when the deployed system receives a different scaffold, more time, new tools, persistent memory, or a collaborator that supplies the missing step? a failure to elicit a capability may reflect its absence. it may also reflect a weak elicitation method. those possibilities are not interchangeable. a responsible argument needs evidence about the evaluation process itself, including whether it can detect capabilities we already know are present.
this is why a single readiness score can become an instrument of concealment. a number can compress several dimensions whose failures have radically different consequences. weak monitoring does not necessarily become acceptable because the system has excellent documentation. a severe permissions failure is not canceled by a high helpfulness score. aggregation is useful when its tradeoffs are defensible and visible. it becomes scientific vulgarity when a weighted average performs the political task of making an unresolved objection disappear.
the necessary threshold is evaluation validity under deployment conditions. the difficult part is deciding how much evidence is enough without pretending that finite tests can establish universal safety. our proposed approach is to make competing explanations fight in the open. take one capability claim. construct positive controls, alternative scaffolds and bounded changes in resources. ask how often the evaluator fails to recognize a capability whose presence is independently established. record what each variation changes, including the cases in which the answer is still unknown.
then attach consequences. if a modest change in tools reverses the result, the safety case must state that dependency. if the evaluator systematically misses a known capability, its negative findings deserve less weight. if the experiment reveals a failure that cannot be mitigated within the intended deployment, the permitted scope should shrink. the uncomfortable research output may be a narrower claim. that is useful knowledge. treating every reduction in confidence as a communications problem creates an institution that can measure almost anything except its own wishful thinking.
we want evaluators who are willing to break the measurement, engineers who understand the surrounding system, and domain experts who can say which errors matter. the first proposed project is small enough to reproduce: one sandboxed agent workflow, one harmful-outcome definition, several elicitation conditions, and a written argument connecting the results to a deployment decision. no universal readiness grade. no claim that a passing experiment settles ASI. just a serious attempt to stop the evidence from becoming more ambitious as it travels toward the press release.
Proposed first work
Create an evaluation-validity study with positive controls and predefined changes to tools, time budget and scaffolding. Publish the resulting limits on the deployment claim, including inconclusive outcomes. Run it in an isolated environment without real-world harmful actions.