Google DeepMind’s Co-Scientist, formalised in a Nature paper in May 2026, compresses the early stage of scientific discovery, generating and refining a research hypothesis, from weeks to days. It works as a multi-agent system built on Gemini: separate AI agents continuously generate hypotheses, critique each other’s reasoning, and refine what survives, in a loop that used to depend on a human researcher’s slower cycle of reading, thinking and drafting. It isn’t a demo. Now rolling out experimentally to individual researchers through Gemini for Science, it has already produced hypotheses that survived contact with a wet lab: a novel drug-repurposing candidate for acute myeloid leukaemia, and a new epigenetic target for liver fibrosis, both independently confirmed rather than merely plausible-sounding. This is a real capability, working, now.
In the same literature this capability is meant to accelerate, something else is happening. Fabricated references in biomedical papers have risen twelvefold in two years. Maxim Topaz, an associate professor at Columbia University’s School of Nursing and Data Science Institute, led an audit of 2.5 million papers — published as a letter in The Lancet in May 2026 — that found roughly one in every 277 PubMed-indexed papers now contains some fabricated references, up from one in 2,828 three years earlier. Review articles, the papers other researchers lean on most heavily to orient themselves in a field, showed a fabrication rate 57% higher than other paper types. It’s already sitting in the permanent scientific record, not a future risk to be managed.
One documented case, precisely
Most of that twelvefold rise is inference from pattern, not confession. But at least one instance is on the record in detail. A 2026 retraction from the Journal of Medical Ethics was explicitly attributed to an author who had used generative AI to identify and understand referenced sources, without independently verifying them before submission. The citations looked right. Nobody checked whether they were.
That’s the mechanism in miniature: a tool that can write a fluent, well-formatted reference list faster than any human, attached to a verification step that still assumes a human wrote it carefully. Multiply that gap across an entire literature moving at AI speed, and the twelvefold rise stops being surprising. The higher fabrication rate in review articles matters more than the raw number suggests, because a review article is exactly the paper other researchers reach for first when they’re trying to orient themselves in an unfamiliar field — it’s the summary everyone trusts precisely because checking every citation in it individually is what the review was supposed to save them from doing.
The system nobody rebuilt
Peer review and citation-checking were designed for a slower, lower-volume era of publishing — one submission at a time, one careful human reader at a time. AI can now generate a polished hypothesis, a literature review, and a citation list faster than that process was ever built to check — and the researchers behind the audit say plainly that the checking side hasn’t caught up to the volume the generating side now produces. There is still no accepted, mandatory standard for flagging an AI-assisted finding as provisional before it enters the permanent scientific record.
My Opinion
I don’t read this as a publishing-etiquette problem, the kind a stricter house style or a sterner editor’s note could fix. Verification is infrastructure, in exactly the same sense that the wet lab confirming Co-Scientist’s drug-repurposing candidate is infrastructure. Nobody would fund a hypothesis-generation system this fast and then leave the checking to whoever happens to notice. Yet that’s precisely what’s happening on the citation side, because the money and the attention are on the generation half of the pipeline, not the verification half. Until that changes, the speed is the risk, not the safeguard.
That gap has consequences beyond the literature itself. A funding committee weighing a grant application, a systematic review synthesising a field, or a clinical trial recruiter screening candidates has no reliable way to know whether the citation or hypothesis in front of them survived independent checking, like Co-Scientist’s confirmed drug-repurposing candidate, or only survived generation, like the reference list behind the Journal of Medical Ethics retraction.
Even Google is candid about this. The company acknowledges Co-Scientist’s own limitations in literature review and factual accuracy — meaning the most rigorously validated AI-assisted discovery tool currently available still depends on external checking that the surrounding system isn’t built to provide at the speed the tool now works at. The researchers behind the audit have recommended that citation-checking move earlier, built into the submission workflow before peer review begins rather than caught, if at all, after publication — a change that depends on publishers adopting it, which none has yet done at scale.
On the page, a hypothesis nobody has independently checked looks identical to one that’s been rigorously verified. Right now, nothing in the publishing system forces that distinction to be visible before the finding starts shaping what gets funded and which patients get recruited into which trials. How many findings currently doing that work have never actually been checked — and would anyone know?
You’re reading The Next Evolution by Neil Catton, articles that explore the human world and the intersection of technology, they try and ask difficult questions - not to scare - but to inform. If someone forwarded this to you, you can subscribe free at neilcatton.substack.com.
Neil Catton is the author of The Next Evolution, The Cognitive Crucible and The Shadow System - available on Amazon, and writes at the intersection of technology, ethics, and human purpose.


