When a biomedical finding appears first as a preprint, the natural question is whether it will hold up once peer review begins. A large claim-level analysis of bioRxiv now measures exactly that. Among papers that later reached peer-reviewed journals, the central abstract claim usually stayed put. What the study does not show is whether those claims were true.
What the researchers measured
Hao Yin, Ruslan Rust, and colleagues matched every bioRxiv preprint posted from 2018 through 2025 that they could link by DOI to a peer-reviewed research article. The matched set contained 72,644 preprint–publication pairs. For each pair they compared the first preprint version with the published abstract, not the full manuscript.
A large language model (Claude Sonnet 4.6) extracted one primary claim and up to two secondary claims from each abstract, then scored content change as unchanged, minor, or major, and scored hedging as more cautious, more confident, or unchanged. Major change covered direction flips, large shifts in magnitude, changes of population or setting, switches in the kind of effect claimed, or wholesale replacement of the claim. On a 550-pair validation set labelled independently by four raters, model–rater agreement approached the agreement among the raters themselves.
The study itself is a preprint, revised on 29 July 2026. Its headline numbers have held across that revision.
What held, and what shifted
Primary claims were unchanged in 39.8 percent of pairs, revised only in wording or light hedging in 50.0 percent, and substantially revised in 10.2 percent. Nearly nine in ten central claims were intact or only lightly reworked by the time the paper appeared in a journal.
When language about certainty moved, it more often softened than hardened. Claims became more cautious in 8.4 percent of pairs and more confident in 4.2 percent. The type of claim — mechanistic, associative, descriptive, methodological, and so on — was preserved in the great majority of cases; when it changed, the shift usually stayed nearby on the spectrum rather than collapsing into a null result.
Revision was not evenly distributed. In the detailed analysis, major primary-claim change ranged from 7.2 percent in bioinformatics to 17.5 percent in microbiology among large fields. Method claims were revised less often than descriptive, associative, or mechanistic ones. Major revision rose with the interval from preprint to publication, from 7.0 percent in the fastest third of pairs to 14.1 percent in the slowest. The headline mix of unchanged, minor, and major change remained similar across journal-impact tiers, including higher-cited journals.
What the numbers do not license
The authors state the limits plainly. The corpus is, by design, limited to preprints that were later published. Earlier literature puts the share of bioRxiv preprints that never reach a journal at roughly 30 to 35 percent; those manuscripts are outside this comparison. Stability among survivors is not evidence about the full preprint stream.
The analysis also reads abstracts, not full texts. Changes buried in methods, figures, or supplemental tables would not appear. The transition labeled "peer review" here is really the combined effect of author revision, reviewer pressure, and journal production. A minority of pairs were posted close to publication — about 11 percent within 90 days — which can mean the preprint already reflected late-stage review; removing those pairs shifted the major-revision rate only slightly, from 10.2 percent to 10.7 percent.
Most importantly, claim continuity is not scientific correctness. An unchanged wrong claim is still wrong. Nature's coverage of the work noted that some researchers read the result more cautiously for exactly that reason: peer review can leave a central message standing while still improving methods, caveats, or evidence quality that abstracts do not capture.
An exploratory comparison of retraction rates found never-preprinted papers retracted more often than preprinted ones in a restricted journal set. The authors treat that contrast as exploratory; the events are few, time at risk is hard to equalize, and preprinting is not randomly assigned.
How to use the finding
The practical takeaway is narrower than a verdict on preprint reliability. For bioRxiv work that eventually publishes, readers can expect the main abstract claim to look familiar after peer review most of the time, with wording more likely to grow cautious than bold. That is practical guidance for scanning early literature: treat the central claim as a provisional signal worth reading, not as a finished verdict, and watch for the minority of cases where scope, direction, or magnitude actually move.
The study also reframes what peer review appears to do at abstract scale. It more often polishes and hedges than overturns. Whether that is enough quality control for public communication is a separate institutional question. Claim stability answers only the first half of the worry about preprints — whether the message survives — not the second half, whether the message deserved to.
Sources
- Tracking claim changes from preprint to publication across 72,644 biomedical studies using large language models (Yin, Ahn, Forster & Rust, bioRxiv v2) · accessed Jul 30, 2026 · primary
- Think preprints are unreliable? Analysis of 70,000 studies might change your mind (Basu, Nature) · accessed Jul 30, 2026 · independent
- Preprint to Publication companion site and dataset browser · accessed Jul 30, 2026 · primary
- Full-text methods and results of the preprint analysis (v1 full text) · accessed Jul 30, 2026 · primary