Skip to content
← all essays
03 / scientific integrity · 4 min read

compute is not a theory of truth

A scientific claim owes us an exact statement, inspectable assumptions, reproducible evidence and visible human credit, whatever produced it.

you can spend a fortune searching for a proof and still owe the world a precise statement of what you proved. you can coordinate thousands of agents and still owe the mathematicians whose ideas defined the search a complete account of their contribution. you can formalize an argument inside a proof assistant and still owe the reader an explanation of why the formal statement corresponds to the advertised problem. compute changes what can be attempted. it does not abolish any of these obligations. the size of the machine is not a substitute for the scope of the theorem.

that standard has to cut through the criticism as well as the publicity. take Navier–Stokes. Charles Fefferman’s official Clay problem description gives four alternatives. the existence-and-smoothness alternatives use zero forcing; the breakdown alternatives allow smooth forcing subject to specified decay or periodicity conditions. the document also distinguishes Euler, with zero viscosity, from the Navier–Stokes prize problem. dismissing a result merely because it is forced would therefore be mathematically careless. the actual question is whether its equations, regularity, domain, forcing and conclusion meet one of the stated alternatives. Official problem description, pages 1–2.

this is a much more demanding criticism than calling every automated result fake. identify the theorem. identify the definitions. identify the gap between the checked statement and the public claim, if there is one. an argument can be formally valid and irrelevant to the proposition being advertised; it can also be computationally expensive and scientifically important. neither enthusiasm nor contempt gets to decide that in advance. the claim must survive contact with the object it claims to describe.

formal verification, discovery and explanation are related but different achievements. a checked proof can remove a class of logical errors. a newly found argument can extend what mathematics knows. an explanation can make the structure intelligible enough to support further conjectures. a single project may accomplish several of these things, or only one. demanding that the authors distinguish them is not hostility to automation. it is the minimum courtesy owed to readers who are being asked to update their beliefs about what a system can do.

the same discipline applies to credit. a result has a history: conjectures, preliminary lemmas, failed constructions, datasets, code, conversations, review and formalization. our proposed lab standard is to document that history through contributor-approved records and clear provenance, respecting confidentiality rather than treating access as permission to disclose. private research does not become ownerless because a model helped process it. and a machine-generated continuation does not make the intellectual work that made the continuation possible vanish into an infrastructure footnote.

for ASI Readiness, the difficult threshold is independent verification at a speed useful for consequential decisions. exhaustive reconstruction of every result is impossible. surrendering judgment to the producer is unacceptable. between those positions lies a research program: structured claim records, reproducible environments, explicit assumptions, inspectable dependencies, independent attempts to falsify the result, and clear distinctions between what was checked and what remains inferred. the artifact should carry enough context that an outside team can locate the most consequential uncertainty without reverse-engineering the marketing.

our first proposed exercise is a claim-to-evidence audit of a public AI-assisted research result. with the relevant permissions, map each headline claim to the actual artifact that supports it. separate search, formalization, novelty and interpretation. have an independent reviewer identify the weakest bridge in the argument. publish the limits of the audit itself. the aim is a method that other teams can reuse, including on our own eventual work. if the method cannot embarrass its creators, it is not an audit; it is brand management wearing a laboratory coat.

the future may contain machines capable of producing mathematics that few humans can follow. that possibility makes provenance and verification more valuable, not quaint. we will need institutions able to say exactly which parts of an argument they have checked, why those checks justify a decision, and when they do not. a civilization that mistakes computational abundance for epistemic authority could automate the production of certainty faster than it can discover its errors. a readiness lab should work on preventing that failure before it becomes the operating model of science.

Proposed first work

Audit one public AI-assisted research claim against its available proof, code or experimental artifact. Publish a claim-to-evidence map, attribution record, independent objections and unresolved gaps. This essay makes no allegation about a specific organization’s private conduct.