AI model flags 261,245 cancer papers for template writing

A language model trained on 2,202 retracted paper-mill articles flagged 261,245 of 2.6 million cancer studies (1999–2024) for writing patterns matching suspected fabrications.

A team led by Queensland University of Technology biostatistician Adrian Barnett published a paper on Aug. 4, 2026 reporting that a BERT-based language model screened 2.6 million cancer studies published from 1999 through 2024 and flagged 261,245 papers, or 9.87%, for writing patterns associated with suspected fabrication.

The model was trained on 2,202 retracted paper-mill articles drawn from a retraction database and then validated against independent expert datasets. In validation tests the classifier reached 91% accuracy in identifying papers that matched the retracted template style used as the training signal.

The share of flagged papers rose over time, from about 1% of annual cancer publications in the early 2000s to more than 16% by 2022. Flagging rates varied by cancer type: gastric cancer papers were flagged at about 22%, bone cancer at about 21% and liver cancer at about 20%.

Adrian Barnett cautioned against treating the percentages as exact counts. He noted the model detects a particular template of writing linked to known retractions and that other templates or more sophisticated fabrication methods could evade detection, so the flagged total may be an underestimate.

The authors describe the tool as a screening filter intended to identify papers with template-like writing patterns for further scrutiny. They say flagged articles should be examined by journals and experts before any editorial action, rather than being automatically retracted.

The paper reports that three scientific journals are piloting the screening technology in editorial workflows to provide automated triage before human review. The authors recommend regular updates to detection tools to keep pace with changes in how paper mills generate manuscripts.

The study notes potential effects if fabricated or low-quality work enters the published record, including possible influence on clinical trial design, research priorities and patient care decisions. The authors present the model and its results as a pattern-based map of template-like writing across the cancer literature, and they call for wider use of automated screening alongside expert follow-up.

The material on GNcrypto is intended solely for informational use and must not be regarded as financial advice. We make every effort to keep the content accurate and current, but we cannot warrant its precision, completeness, or reliability. GNcrypto does not take responsibility for any mistakes, omissions, or financial losses resulting from reliance on this information. Any actions you take based on this content are done at your own risk. Always conduct independent research and seek guidance from a qualified specialist. For further details, please review our Terms, Privacy Policy and Disclaimers.

Articles by this author