Skip to content
← all essays
05 / institutional design · 4 min read

the right to stop must survive the cost of stopping

A proposed institutional test for safety authority: can the evidence change the decision when doing so becomes expensive?

a safety policy is easiest to admire before it has been asked to cancel anything. before a release date moves, before a financing story weakens, before a customer walks, before a rival gets ahead. at that point almost everyone can agree that evidence should govern deployment. the disagreement begins when evidence demands a sacrifice. that is the moment a readiness framework stops being a statement of values and reveals the actual distribution of power inside the institution.

our proposed standard is simple to state and difficult to establish: a credible stop decision must remain possible when the organization strongly prefers to continue. that means someone has the authority to make it, access to the evidence needed to justify it, protection against retaliation for using it, and a procedure that turns the decision into an operational change. if any link is absent, a carefully written policy can become a ceremonial layer above a system that was never designed to accept the answer no.

there are useful foundations. NIST’s AI Risk Management Framework is a voluntary framework for incorporating risk management into the design, development, use and evaluation of AI systems. it supplies a vocabulary and structure for organizational work. it does not, by existing, confer enforceable veto power on an evaluator or resolve a conflict between a deployment team and its commercial leadership. those arrangements have to be built and assessed in the institution that claims to use the framework. NIST AI Risk Management Framework.

for a lab in formation, this is an obligation before it is a criticism of anyone else. we cannot promise independence merely by choosing the word independent. the proposed funding structure, publication rights, conflict-of-interest rules and decision authority must make that independence plausible. a funder can support the work without owning its conclusions. a partner can receive a fair chance to correct factual errors without receiving an indefinite right to suppress them. these distinctions need to appear in actual arrangements, not remain dependent on the personal courage of whoever writes the report.

the threshold feels impossible because institutions are asked to constrain the very activities that sustain them. an evaluator may rely on access from the organization it evaluates. a researcher may depend on a grant renewal. an operator may be judged on uptime while being asked to interrupt a service. the problem is not solved by announcing that everyone should have better incentives. we need arrangements that reduce a specific conflict, make residual dependence visible, and still function when the people involved are tired, pressured or replaced.

one proposed ASI Readiness project is an institutional stress test tied to a real technical finding. begin with a simulated discovery that invalidates a deployment assumption. follow the evidence through reporting, review, disagreement, escalation, suspension and eventual restart. introduce the ordinary pressures that clean process diagrams omit: an unavailable executive, ambiguous responsibility, a disputed result, a major customer's deadline, uncertainty about whether rollback will damage another service. record where the decision stalls and who can override it. the exercise should expose the organization’s actual constraints.

the experiment also has to take mistakes seriously. a stop mechanism with no review, no appeal and no proportionate restart process can become arbitrary power. readiness requires restraint on the people exercising oversight as well as on the system being overseen. the proposed design should distinguish a reversible precaution from a final prohibition, specify what evidence changes the decision, and give affected people a meaningful route to challenge it. authority earns legitimacy through accountable use, not simply through association with the word safety.

the lab should publish the resulting authority map and unresolved conflicts with the participating organization’s consent and appropriate protection for sensitive operational details. the useful outcome may be the discovery that a proposed safeguard cannot actually be invoked. then the institution has a concrete choice: change the arrangement, narrow the deployment, or admit that its assurance story is incomplete. what it should not do is retain the original claim and move the failed mechanism into a future-work section nobody is expected to read.

this is work for security operators, governance researchers, organizational designers and people who understand how decisions happen under pressure. a technically sound warning that cannot travel into action is an unfinished safety system. if we want the public to take the word readiness seriously, the evidence must be able to do more than inform the people with power. under specified conditions, it must be able to stop them.

Proposed first work

Run a consent-based tabletop exercise with one partner, one simulated technical failure and an explicit stop/restart decision. Document authority, escalation, overrides, appeal, recovery and the conflicts that remain. This is an institutional research proposal, not a claim that enforceable arrangements are already in place.