Overview
Accepted at AI for Science @ NeurIPS 2026 Paper Slides
Every claim in a paper comes from somewhere.
A declarative, continuous, machine-checkable provenance layer for research claims, throughout the whole research lifecycle.
- Recording: every result is tied to the run that produced it, at the time it’s produced.
- Verifying: every number against its results file, every quotation against its source, every analysis against its plan.
- Agents: reproducible by construction. The record is made as the agent works.
pip install reproducible-science/plugin marketplace add elliottower/reproducible-science/plugin install reproducible-science@reproducible-science
Like git status, for research provenance. In Claude Code, the state of a project’s records
sits above the prompt, with a warning when something no longer matches.
| tool | what it does |
|---|---|
prereg |
freezes a plan before running, records what changed after |
results |
seals inputs, records outputs, binds claims to runs |
citations |
checks that quotations resolve in the sources they cite |
repro |
verifies a paper’s plans, numbers and quotations in one report |
Each is an independent distribution with its own public API, so installing citation verification never drags in a preregistration tool. They live in one repository because a change that crosses two of them should be one commit rather than a release sequence.
The problem
Section titled “The problem”AI tooling makes it quick to generate hypotheses, run analyses and draft manuscripts. Verification and provenance have not kept pace:
- Verifying that a finished paper agrees with the artifacts behind it takes manual or agentic checking, which is expensive and difficult to audit.
- Verification usually reports no denominator for the sources or artifacts checked.
- A verification snapshot goes stale quickly, and a second run is not guaranteed to find the same issues.
- More experiments mean more researcher degrees of freedom and post-hoc analysis, even when unintentional, or done by an agent.
- Verification is typically done after the fact, and does not cover every step of the research lifecycle.
As agents take over more of the research lifecycle, the artifacts they leave behind are worth only as much as the claims inside them can be verified.
The chain
Section titled “The chain”A number in a manuscript names a claim. The claim names a run. The run names its outputs, hashed when they were recorded. The inputs were hashed before the run started.
prereg freeze # lock the planresults seal PREREG.md analysis.py # hash the inputsresults run output.json --run-id exp_001results claim "ICC = 0.42" --run-id exp_001 --location "Table 2"repro verify # check the whole chainLimitations
Section titled “Limitations”It does not decide that a paper is reproducible. It checks relations:
- that a claim addresses an artifact;
- that the artifact is the one that was pinned;
- that the addressed value is what the manuscript prints;
- that a confirmatory run started after the plan it names was registered.
Everything it cannot establish is reported as unestablished rather than assumed.