Skip to content
← home
ASI READINESS / FORMATION BRIEF

a lab is forming at the edge of what we know how to govern.

the proposed agenda, the terms of formation, and the people this work needs.

ASI Readiness is forming around a demanding proposition: the conditions required for a survivable transition to more capable AI must become a research agenda, even where satisfying them would force our institutions to change.

we want the alignment researcher whose result needs an operational setting, the security engineer who can turn a threat model into a test, the evaluator who can explain why a reassuring metric fails, the institutional designer who can make a stop decision enforceable, and the critic who can show where the entire proposal is mistaken. the lab should make those contributions meet around shared experiments and consequential decisions.

this is a formation-stage proposal. the research agenda below describes work to undertake. it does not imply an established legal entity, secured funding, appointed staff, institutional affiliations or completed experiments.

the research agenda

01 / evidence that can bear the claim

The necessary threshold: deployment claims whose scope and strength are justified by inspectable evidence.

Why it feels out of reach: evaluations cover finite conditions; deployed agents encounter different tools, incentives, time horizons and collaborators. independent reviewers also face limits on access, time and expertise.

First proposed artifact: a claim-to-evidence record for one deployment, with positive controls, reproducible evaluations, an independent objection and explicit conditions that invalidate the claim.

People this needs: evaluation researchers, statisticians, domain experts, mathematicians, formal-methods researchers and research engineers.

02 / control that survives a stronger adversary

The necessary threshold: operational constraints that remain effective under the capabilities and adversarial strategies assumed in the threat model.

Why it feels out of reach: a monitor may be weaker than the monitored system, share its blind spots, or see too little to intervene. a model’s fluent account of its own behavior is insufficient evidence of what caused that behavior.

First proposed artifact: a reproducible sandbox comparing a bounded control protocol with independent attacks, measuring failures, false alarms, usefulness and human-review costs.

People this needs: alignment scientists, interpretability researchers, security engineers, red-team specialists and systems operators.

03 / institutions that can act on bad news

The necessary threshold: a credible process for narrowing, interrupting and restarting deployment when evidence warrants it.

Why it feels out of reach: the people finding the problem can depend on the people paying for the deployment; authority can be ambiguous precisely when a decision becomes expensive.

First proposed artifact: a consent-based institutional stress test, with an authority map, an escalation path, protected challenge, proportionate appeal and a recovery procedure.

People this needs: governance researchers, operational leaders, organizational designers and specialists in accountable decision-making.

04 / cooperation under pressure

The necessary threshold: agreements that remain useful under disagreement, imperfect information and unequal incentives.

Why it feels out of reach: verification can conflict with confidentiality; participation can impose unequal costs; actors can gain by withholding the evidence everyone needs.

First proposed artifact: a simulated incident-sharing arrangement with explicit reporting incentives, access limits, verification and dispute resolution, followed by a bounded pilot only if justified.

People this needs: cooperative-AI researchers, mechanism designers, political scientists, security practitioners and negotiators.

05 / human agency that survives dependence

The necessary threshold: people affected by AI decisions retain meaningful ways to understand, challenge and change those decisions.

Why it feels out of reach: participation is unequal; affected people may lack technical resources; systems can become embedded before a workable appeal or fallback exists. capability does not settle disputes about authority or values.

First proposed artifact: map one consequential AI-mediated workflow from the affected person’s perspective; test whether an objection reaches someone able to act and whether a usable fallback exists.

People this needs: human-computer interaction researchers, social scientists, civic practitioners, accessibility specialists and people directly affected by the chosen workflow.

proposed terms of formation

Publish the claim and its limits together. Define the conditions of the experiment and the boundary of the conclusion. Label proposals, simulations, replications and results distinctly.

Make criticism operational. Every project should have an objection strong enough to change its next step. Independent review should be able to narrow a claim or halt a proposed pilot.

Protect attribution and permission. Credit contributions through contributor-approved records. Agree how private material may be used before using it. Keep publication rights and authorship expectations explicit.

Give negative results a home. A failed safeguard, an inconclusive test and a rejected hypothesis can prevent expensive mistakes. They should not disappear because they complicate the lab’s story.

Make independence concrete. Proposed funding and partnership terms should disclose relevant conflicts and protect inconvenient findings. These principles become institutional commitments only when reflected in actual governance and agreements.

Use proportionate disclosure. Publish methods and findings where appropriate while protecting sensitive material and avoiding unnecessary release of dangerous capabilities. Explain any restriction and how qualified independent scrutiny can still occur.

the first formation milestones

  1. Select one shared problem with a defined deployment setting and a falsifiable claim.
  2. Identify a project lead, independent reviewer and operational counterpart; document conflicts and contribution expectations.
  3. Establish the budget, access, governance and publication arrangements needed for that work. Announce roles and compensation only once those arrangements exist.
  4. Publish the proposed method and invite criticism before running the study.
  5. Produce a reproducible artifact and report what failed, what held and what remains unknown.
  6. Decide whether to expand, revise or stop that line of work on the evidence.

These are milestones, not dated promises or claims of current completion.

bring the problem that will not fit inside the reassuring story

the useful introduction is a piece of work: a paper, a replication, an evaluation, a threat model, an operational failure, or a carefully argued objection. describe the threshold it bears on, what is already known, the experiment that should come next, and what result would change your mind. include the skills or resources you can contribute and the constraints on using your material.

we want people willing to connect technical findings to institutional consequences and institutional proposals to tests that can fail. affiliation is less useful than intellectual honesty, careful work and the ability to make a disagreement productive. the point of forming a lab is to give that work a place to accumulate.