top of page

The Confident Error: Why AI Makes Human Editorial Judgement More Necessary In Academic Copyediting

  • Writer: Julie Pinborough
    Julie Pinborough
  • Jul 17
  • 6 min read

One in every 277 biomedical papers published in early 2026 contained a fabricated reference.

I recently discovered that one in every 277 biomedical papers indexed on PubMed during the first seven weeks of 2026 contained at least one fabricated reference [1]. Three years earlier, in 2023, the figure was one in every 2,828. At the time the audit was conducted, over 98% of those papers had seen no publisher action [1]. These false references have stayed exactly where they were published, quietly available to anyone who chose to trust them.


It’s easy to make fabricated citations the villain here. An LLM (Large Language Model) invents an article that sounds plausible, formats it perfectly, supplies convincing author names and page numbers, and the references survive long enough to make it into print. It’s a satisfying story because it has a clear culprit.


But the harder question we should be asking is why these are getting so difficult to spot.


The answer isn’t really about the invention; it’s about confidence.

Why Fluency Isn’t Evidence


Academic writing has always run on a grammar of uncertainty – one of the reasons it sounds so cautious to non-academics. Generally, researchers don’t claim proof. Instead, evidence supports a conclusion or is consistent with one under specific and stated conditions. These qualifiers aren’t hedging for the sake of politeness but rather tell a reader exactly how much weight a claim can carry.

Additionally, LLMs are optimised for fluency, not for preserving that weighting.


They’re very good at reproducing the statistical shape of academic prose, and in doing so, they tend to iron out the very features that were signalling uncertainty. A once cautious observation becomes a stronger conclusion, and the preliminary evidence becomes a demonstration. None of this requires an outright falsehood – it’s subtle enough to read perfectly well – but it silently changes the relationship between the evidence and what’s being claimed about it.


Citation fabrication works on a similar principle. Models don’t usually invent references from nothing; they’re not entirely ‘hallucinating’ as we have now come to term it. Instead, they blend fragments of real scholarly metadata into something that reads as authentic. A familiar author name, a plausible journal, correct formatting… every component looks real even though the whole doesn’t truly exist.


Not only that, but more often than not, the reader’s confidence is earned by presentation, not verification. This matters more with citations than almost anywhere else in a manuscript, because a citation isn’t just supporting evidence, it’s an address. It lets someone else trace the claim back to its source and judge for themselves whether the interpretation stands up – or not. When the address doesn’t exist, that check simply isn’t available to anyone.


Departures Are More Interesting


The same averaging effect turns up in methods sections. Researchers often deviate from standard procedure precisely because they’re doing something worth doing, like adapting a technique or working around an odd dataset. Those departures are frequently the most interesting part of the paper. A model has no way of recognising methodological originality; it generates methods from the pattern of a million other methods sections, and unusual approaches drift toward the average description of what a methods section normally says. Nothing’s been falsified, but something has just been quietly lost.


And the better the prose reads, the less likely anyone is to notice. Editors and reviewers have always used friction as a signal because awkward phrasing slows us down; it invites a second look. Fluent text doesn’t ask for that scrutiny because fluency reads as reliability, whether it’s deserved or not. Psychologists have a name for this (processing ease), but we don’t really need the label to recognise it in our own reading habits.


This ‘processing ease’ changes what’s being asked of authors. It’s no longer ‘can AI produce acceptable academic prose’ – obviously, it can – it’s whether the author is still close enough to their own argument to notice when a linguistic improvement has moved the meaning. That’s not something we can hand over to the software and ask it to check.


It’s also not something that can be handed entirely to the copyeditor anymore, and this is the part that concerns me more than the researcher-side problem, because it’s closer to home.


Nobody’s Watching The Middle



Academic copyediting has traditionally meant improving clarity while protecting meaning – grammar, ambiguity, terminology, consistency – on the assumption that scientific accuracy is the author’s job and communication is the editor’s. AI disrupts that division of labour, because the manuscript an editor receives may already have passed through a model once, and the editor’s own tools may pass it through a second one before a human has properly checked whether the underlying claim still says what it originally said. Two independent rounds of smoothing, and yet there is nobody in the process whose job it was to catch the drift between them. Errors don’t get caught by repeated editing in this scenario, but they do sometimes get more persuasive.


AI In Academic Copyediting


‘I asked AI’ isn’t a defence, it’s an admission that verification never happened.

So the editorial question shifts from ‘is this sentence correct’ to ‘does the claim still rest on the same evidence it did before every improvement got applied to it’. That’s a considerably heavier lift than a typical copyediting style pass.


It gets more awkward in fields where terminology itself is the site of disagreement: where two research groups use slightly different vocabulary because they genuinely dispute the underlying mechanism. An LLM defaults to the statistically dominant phrasing because that’s what it’s seen most. But inconsistency here isn’t always sloppiness; sometimes it’s the debate, visibly happening in the language. Standardise it, and we’ve made the manuscript tidier and less honest about where the field actually stands.


There’s a live argument in research-ethics circles about whether hallucinated citations count as misconduct, given that citations often function as evidence in the argument itself for establishing priority, justifying method, and supporting an interpretation. I won’t explore that here (that’s a different question for a different blog). What does seem indefensible is the idea that authorship carries less responsibility because AI generated the reference. It doesn’t. ‘I asked AI’ isn’t a defence, it’s an admission that verification never happened.


That’s really the crux of it. Knowledge doesn’t enter the record because it’s written well. It enters because someone puts their name to it and accepts what follows if it’s wrong. Careers and reputations get attached to claims, which is precisely why errors can be corrected – someone owns them. An LLM owns nothing. It can’t issue a correction, defend a methodological choice, or answer a peer reviewer. That’s not a gap that better engineering closes because it isn’t actually a capability gap. Accountability isn’t a feature we train into a system; it belongs to people because they bear the consequences of being wrong.


None of which is an argument against using the tools, for what it’s worth. AI is already catching things that human reviewers routinely miss, such as duplicate images, statistical anomalies, and manuscript overlap, at a scale no editorial office could manage by hand. Used well, it’s a net gain for research integrity rather than a threat to it.


AI Relocates The Risk


The trouble is that AI relocates the risk rather than removing it. Most editorial workflows are still built to catch the errors humans have always made, such as typos, inconsistent formatting, awkward subheadings, or an incomplete reference list. AI’s failure modes are a different shape entirely: for example, confidence where there should be caution or statistical averages where there should be specificity. The process hasn’t caught up to where the risk actually sits now, which might be the real lesson in that PubMed number.


The danger has never simply been that AI gets things wrong sometimes. It’s that polished writing increasingly looks like trustworthy scholarship, whether or not the foundation underneath it is actually substantiated. The manuscript or citation shouldn’t earn trust because it reads convincingly; it should earn it because a confident sentence can still be questioned, a citation can still be followed to its source, and somebody is still willing to answer for both.



Julie Pinborough

About Me


I’m a professional copyeditor and copywriter specialising in academic and non-fiction writing. With backgrounds in law, history, and English literature, and more than a decade as a lecturer in higher education, I help authors communicate complex ideas with clarity and precision.


I believe good editing is less about changing a writer’s voice and more about helping readers hear it.

 
 
bottom of page