Paper mills produce fabricated text. Clinical trial site fraud produces fabricated patients. The two problems are sometimes discussed in the same breath, but they operate at different layers of the scientific process and create different hazards for the published record. A paper mill can be caught by text analysis, citation audits, or pattern recognition in abstracts. Fraudulent trial data, assembled by a site that enrolled people without the disease, copied values from one participant to another, or simply invented case report form entries, looks like real data. It has the right distribution. It has a plausible-looking response curve. It passes statistical review because the numbers were constructed to pass it.
Clinical researchers who publish multi-site trial results carry responsibility for data they did not personally collect. The principal investigator at the coordinating center typically has no independent access to raw patient records from every site. They see what the data management system received. If a site entered fraudulent observations, those observations travel up the pipeline and into the analysis dataset that eventually drives a publication. The PI did not fabricate anything. They published a result that was contaminated before they saw it.
That is not a hypothetical scenario. In January 2026, Science published an investigation into allegations against five South Florida clinical trial sites that had participated in a Phase 2 study of an experimental Alzheimer drug called T3D-959, developed by the company T3D Therapeutics. The allegations, described in a legal complaint filed by T3D Therapeutics in July 2025, included some placebo-group participants appearing to improve at rates inconsistent with Alzheimer's disease progression, individuals enrolled who apparently did not have the disease, blood samples reused across patients, and drug absent from the blood of patients documented as having received it. The results were described in contemporaneous company communications as medically impossible.
Key Distinction
Paper mills produce fake text around real or nonexistent studies. Clinical trial site fraud produces fake data within otherwise real studies. The two problems require different detection methods, and trial site fraud is considerably harder to catch from the published paper alone.
The T3D case did not result in any retracted publications at the time this post was written, because the Phase 2 results were apparently recognized as unreliable before the company published them. That outcome is unusually fortunate. In many documented fraud cases, the contaminated data completes the analysis, drives a publication, passes peer review, and sits in the literature until someone with access to raw site-level data or regulatory inspection records raises a flag years later. By that point, the published result has been cited, and the citations have been cited.
What Clinical Trial Site Fraud Actually Looks Like
The term covers a wide range of practices, and not all of them involve the same degree of deliberate misconduct. At one end is outright fabrication: a site invents participants who never existed, manufactures case report form values, creates plausible-looking biomarker results, and submits the data to the central database. At the other end is the softer failure of a site that enrolls borderline patients who do not strictly meet inclusion criteria because the coordinator wants to hit enrollment targets and earn the per-patient payment.
Between those poles sits a range of practices that researchers and regulators have documented in inspection records over the decades. These include enrolling the same individual under multiple identities, recording visits that did not occur, copying laboratory values from one time point or one patient to another, altering dates in source documents, and administering study drug differently from the protocol while recording compliant administration. The FDA maintains records of sites that have received Warning Letters or been debarred, but the public database represents only the fraction of fraud that was investigated, which is itself a fraction of what occurred.
What makes site fraud particularly problematic for the published record is that it is not randomly distributed across the trial dataset. Fraudulent sites tend to show unusually good results, unusually complete follow-up, or unusually low adverse event rates, because these patterns are what a coordinator fabricating data might logically choose to simulate good performance. Those unusual results are not outliers that analysts remove as implausible. They are often the most favorable entries in the dataset, and they shift the aggregate result in the direction of efficacy.
Common forms of trial site fraud
- 1.Enrolling participants who do not meet inclusion criteria, including people without the target disease.
- 2.Fabricating visits, assessments, or procedures that did not take place.
- 3.Copying laboratory values or clinical scores from other timepoints or other participants.
- 4.Enrolling the same individual under different identities to multiply per-patient payments.
- 5.Altering source documents after the fact to match expectations or resolve protocol deviations.
- 6.Recording drug administration that did not occur or occurred at the wrong dose.
The Professional Patient Problem
Alongside site-level fraud, a separate but related problem involves participants rather than investigators. Researchers studying fraudulent enrollment in clinical trials have found that roughly four percent of trial volunteers at some local clinics are what the field calls professional patients: individuals who enroll repeatedly, sometimes under different identities, for the per-participant payments. They may not have the target condition. They may have been enrolled in multiple competing trials simultaneously. They may report symptom improvements that reflect social desirability rather than pharmacological effect.
In neurodegenerative disease trials, professional patients create a specific statistical signature. Alzheimer's disease follows a fairly predictable cognitive trajectory. Patients who do not have the disease do not follow that trajectory. They may score unexpectedly well at baseline or show improvements on cognitive measures that genuine patients do not show. Those improvements dilute the treatment effect, complicate the placebo-arm trajectory, and in the worst cases produce results that suggest the experimental drug caused harm relative to a placebo group that was quietly improving because many of its members were cognitively healthy.
The professional patient problem is not new. It has been documented in pain, psychiatric, and dermatological trials over many years. What changed is the scale at which it can operate when sites are paid per enrollment and oversight is light. Researchers have described the problem as particularly acute in markets where per-patient fees are high relative to local incomes and where the infrastructure for verifying participant identity or disease status is less robust than in academic medical center settings.
How COVID Expanded the Problem
Decentralized clinical trials accelerated sharply during the COVID-19 pandemic as traditional in-person site visits became impossible. Remote enrollment, telehealth assessments, and at-home outcome measures offered real practical benefits, and many sponsors adopted them permanently after the emergency lifted. But decentralization also removed several of the natural checks on fraud that in-person site visits provide.
When a research coordinator meets a participant face to face, some basic verification happens by default. The person exists. They appear to be the age they claimed. Their physical presentation is at least broadly consistent with their stated health status. In a fully remote trial, those ambient checks are absent. Fraudulent participants can enroll using AI-generated identity documents. Multiple individuals can share an enrollment profile. Research teams at online clinical trial centers have documented cases where staff recognized that two separately enrolled participants appeared to be the same person dressed differently in a video call, or where data entry timestamps were inconsistent with the visit times recorded.
Sponsors and contract research organizations began publishing fraud-prevention checklists for decentralized trials from around 2024 onward. A 2025 paper in a leading implementation science journal described the experience of a Boston research center that detected ten fraudulent participants in an online trial before implementing stronger screening, then identified thirty-seven additional suspicious individuals at the screening stage after adding new verification steps. Those numbers suggest that in an unprotected online enrollment pathway, fraudulent participation rates can be high enough to meaningfully affect small to medium-sized trials.
Decentralized trial enrollment: warning signs
- Unusually rapid enrollment at a single remote site relative to others in the same geography.
- Participants who consistently perform at or near inclusion thresholds without variation.
- Similar response curves across many participants at the same site, suggesting copied or templated data.
- Baseline demographic data that clusters tightly around the enrollment criteria rather than the realistic population distribution.
- Missing or implausible timestamps on assessments, particularly for home-based or app-based measures.
- Unusually low dropout and protocol deviation rates at a specific site.
Why Fraudulent Trial Data Is So Difficult to Catch Before Publication
The most important reason fraudulent site data evades detection is that it was constructed to look plausible. A coordinator fabricating cognitive scores for a patient who does not have Alzheimer's will typically enter scores in the plausible range for a mildly affected patient. The values will not look like outliers. They will not trigger data management queries. They will not flag as implausible in range checks. They will, at worst, contribute an unusually favorable response to the treatment or placebo arm, which may actually be welcomed by the study team as evidence of measurement sensitivity.
Peer review does not access site-level raw data. Reviewers see a manuscript, tables of aggregate results, and possibly a supplement with detailed statistical output. They cannot see that three participants at Site 14 have cognitive trajectories that are inconsistent with a progressive disease. They cannot see that blood samples from Site 7 were drawn on dates that do not appear in the pharmacy administration log. Those inconsistencies live in the source documents and study databases that regulators can subpoena but journals cannot.
Post-publication integrity review faces the same structural barrier. When a reader raises concerns about a multi-site trial, the journal can ask authors to provide data, but the authors themselves may have limited access to site-level raw records. The sponsor may hold those records. The contract research organization may hold portions. If the trial was commercially sponsored, the sponsor may be reluctant to produce documentation that supports a fraud allegation against a site they continue to work with. Journals that send forensic biostatisticians to review data often find that the data they receive is already an aggregated form.
Statistical and AI Detection Methods Being Developed
Researchers have been building statistical tools to detect anomalous site-level patterns since at least the early 2000s, and those tools have become considerably more sophisticated in recent years. Central statistical monitoring uses distribution-based tests to compare each site's data patterns against the expected distribution from the rest of the trial. Sites with unusually low variance in continuous measures, unusually high protocol compliance, or unusual response patterns relative to other sites get flagged for closer human review. These approaches have been validated against historical fraud cases and can detect anomalies that would not be visible to an analyst reviewing aggregate results.
More recently, machine learning approaches have been applied specifically to Alzheimer trial data. A 2026 paper in a clinical pharmacology journal described a paradoxical patient analysis system that identified site-level fraud by looking for participants whose longitudinal profiles were inconsistent with expected disease progression. The approach does not require access to source documents. It applies to data that has already been aggregated into the study database, and it can flag anomalous sites before unblinding or analysis.
These tools are promising, but they are not yet standard practice across the industry. They require computational infrastructure and statistical expertise that many smaller sponsors and academic trial teams do not have in-house. They are also most effective when applied prospectively during data collection, not retrospectively after a study has closed. A detection system that runs during the trial can prompt site audits and potentially exclude contaminated data before analysis. A detection system applied to a completed dataset may identify fraudulent contributions, but the question of what to do with that finding, including whether and how to publish, becomes considerably more complicated.
What Gets Published and What Gets Retracted
Published papers from multi-site trials that included fraudulent sites are a known problem in the literature. Retractions in this category tend to come after regulatory action rather than peer review failure. The sequence typically runs as follows: a drug fails to replicate in a second trial, a regulatory agency conducts an inspection of sites from the first trial, the inspection reveals fabricated or altered records at one or more sites, the agency issues a Warning Letter or initiates debarment proceedings, the sponsor eventually notifies the journal, and the journal retracts.
That sequence takes years. The published paper accumulates citations throughout. When the retraction eventually appears, the citation accumulation does not stop immediately, because retracted papers continue to be cited, sometimes for a decade or more. A 2025 analysis of retraction patterns across five decades of medical publications found that retracted papers in clinical pharmacology and cardiology were among the most heavily cited categories, presumably because they reported efficacy results that subsequent researchers built on before the underlying fraud was identified.
Retraction Watch's managing editor Kate Travis, testifying before a House Science, Space and Technology subcommittee in April 2026, noted that the current retraction rate of approximately 0.2 percent of published papers likely understates the true problem by a factor of ten or more. Scientific misconduct, including fabrication and falsification, accounted for roughly 60 percent of retractions in their database at that time. Those figures cover the full range of research integrity failures, not only clinical trial fraud, but they give some sense of the scale of contamination that may exist in the published record.
The Regulatory Picture in 2026
The FDA's response to clinical trial fraud operates on a different timeline than journals'. The agency conducts bioresearch monitoring inspections of trial sites, both routine and for-cause, and it has authority to issue Warning Letters, require additional studies, or in serious cases move to debar investigators from participating in future regulated research. The FDA's public database of debarred individuals is accessible and names specific investigators and sites. Researchers designing trials should check it before selecting sites, though the list reflects only completed debarment actions and not the longer list of Warning Letters or ongoing investigations.
An internal FDA analysis circulated in 2025 found that roughly 30 percent of clinical studies required to submit results to ClinicalTrials.gov under the agency's reporting rules had not done so within the required timeframe. That non-compliance figure does not indicate fraud, but it suggests that the compliance infrastructure around clinical trial conduct remains weaker than the requirements imply. Sponsors and investigators who do not report timely results often face no immediate consequence, which reflects the limited enforcement resources available relative to the volume of trials being conducted.
Contract research organizations, which manage large proportions of industry-sponsored trials, have increasingly added their own fraud detection protocols to their quality management systems. The STM Integrity Hub, a collaborative project between major publishers and research integrity tool developers, is expanding its scope from post-submission screening to include earlier-stage verification tools. Whether these efforts will be sufficient to intercept site-level fraud before it enters publications is an open question, but the direction is toward earlier detection rather than post-publication correction.
What Clinical Authors and Principal Investigators Should Do
The uncomfortable implication of clinical trial site fraud for authors is that a publication can be compromised by conduct the PI did not oversee, did not witness, and may not have been able to prevent under standard trial management practices. That does not mean there is nothing to do. Several practices reduce exposure meaningfully.
Before selecting trial sites, review the FDA's debarment database and search Warning Letter records by investigator name. Sites that have received fraud-related regulatory action in the past are not automatically disqualified from future work, but their history should inform the level of monitoring they receive. For commercially sponsored trials run through contract research organizations, ask the CRO what central statistical monitoring approach they use, when it is applied, and what the threshold is for triggering a site audit. Central monitoring applied only at study closeout is considerably less useful than monitoring applied during data collection.
When analyzing data from multi-site trials before publication, run site-level consistency checks as part of the standard analysis plan. Examine whether any single site has anomalously low variance, an unusually favorable outcome distribution, or an implausibly low dropout or protocol deviation rate. Those checks do not require specialized software. They can be done in any statistical package with site as a grouping variable. If a site's data looks too clean, take that as an invitation to review the monitoring records before finalizing the analysis rather than after.
Pre-publication site integrity checklist for multi-site trials
- 1.Review the FDA debarment database and Warning Letter records for all enrolled sites and named investigators before finalizing authorship.
- 2.Run site-level variance and distribution checks on all primary and secondary outcomes, flagging any site with statistically implausible homogeneity.
- 3.Compare enrollment-to-completion rates across sites. Sites with unusually low dropout should receive the same scrutiny as sites with unusually high rates.
- 4.Confirm that biomarker or laboratory results from problematic sites are consistent with expected ranges for the enrolled population.
- 5.Document your monitoring approach in the Methods section. Reviewers and editors increasingly regard central statistical monitoring as part of trial conduct, not a bonus.
- 6.If site-level anomalies are found during analysis, consult with your data safety monitoring board or trial steering committee before submitting, not after a reviewer raises the question.
One practical issue concerns the Methods section. Most published clinical trial methods describe enrollment, randomization, interventions, and outcome measures. They rarely describe the monitoring approach used to detect data anomalies during the trial. As publishers and journals increase their scrutiny of multi-site trial submissions, describing your central monitoring approach in the methods, even briefly, demonstrates that the question was addressed before the paper was written rather than discovered after it was submitted.
What Journals Can and Cannot Do About This Problem
Publishers have been steadily expanding their pre-publication integrity screening over the past several years. Springer Nature now has more than 75 full-time research integrity staff. Major publishers use image analysis tools, statistical screening, and authorship pattern checks as part of routine desk review. But those tools are well-matched to text fabrication, image manipulation, and statistical anomalies within a single paper's aggregate results. They are not designed to detect that the underlying patient data was fabricated at a site the journal has never heard of.
Some journals have started requiring authors of multi-site trials to confirm in their submission checklist that central statistical monitoring was conducted, or to provide a summary of site-level audit findings on request. This is not yet universal, but it represents a realistic lever for journals: they cannot inspect trial sites, but they can require authors to attest that someone did.
Post-publication, journals depend heavily on whistleblower reports, regulatory disclosures, and the analysis of independent researchers who access published data and identify inconsistencies. That dependency means the timeline from fraud to retraction is almost always measured in years. Authors who discover post-publication that one of their trial sites has been the subject of a regulatory action face a real decision about whether and how quickly to notify the journal. COPE guidance is clear on this point: authors have an obligation to notify journals when they become aware of potential errors or misconduct that could affect the integrity of their published work. Waiting for the regulatory process to conclude before telling the journal is not a position that COPE or the major publishers support.
A Practical Note for Academic Trial Investigators
The clinical trial fraud problem gets the most attention in the context of industry-sponsored trials with large commercial stakes, and that attention is warranted. But the problem is not confined to commercially sponsored research. Academic multi-site trials, particularly in therapeutic areas where per-patient payments are significant relative to local salaries, are not immune. Collaborative networks in neurology, oncology, and psychiatry that rely on community clinical sites rather than academic medical centers face the same monitoring gaps that the industry does.
The difference for academic investigators is that the institutional incentives to detect and address fraud, rather than absorb it quietly, are generally stronger. An academic principal investigator who discovers anomalies at a site does not have the same commercial interest in suppressing the finding that a sponsor might have in protecting a relationship with a high-enrolling site. That structural advantage should translate into earlier reporting and faster remediation when problems are found.
The practical takeaway for any clinical author is straightforward, even if it is uncomfortable to state plainly. The data that drives a multi-site publication was collected by people you may never have met, at locations you may never have visited, under supervision that varied in rigor from site to site. That is a real source of uncertainty about what your paper actually reports. Statistical monitoring and site audits reduce that uncertainty. Describing them in your methods section discloses it honestly. And checking before you submit, rather than after a reader raises a question, keeps the problem within the range of things you can manage.
Further Reading
Paper Mills in 2026
The related but distinct problem of fabricated text factories and how to protect your citations from contaminated literature.
ClinicalTrials.gov Reporting Requirements
The FDA's 2026 enforcement push and what trial sponsors need to submit and when.
Scientific Image Integrity in 2026
A different layer of data integrity enforcement that now operates in parallel with site-level monitoring at major publishers.
Publisher Integrity Screening in 2026
How major publishers now screen manuscripts before peer review, and what that means for how you describe your methods.
Written by Dr. Meng Zhao
Physician-Scientist · Founder, LabCat AI
MD · Former Neurosurgeon · Medical AI Researcher
Dr. Meng Zhao is a former neurosurgeon turned medical-AI researcher. After years in the operating room, he moved into applied AI for clinical workflows and now leads LabCat AI, a medical-AI company working on decision support and research tooling for clinicians. He built Journal Metrics as a free resource for researchers who need reliable journal metrics without paid database subscriptions.
Related Articles
The 2025 COPE Retraction Guidelines: What Medical Authors Need to Know
COPE released Version 3 of its retraction guidelines in August 2025, adding explicit criteria for undisclosed AI use, paper mills, and batch retractions. Here is what medical authors need to understand about how journals now handle post-publication integrity concerns.
15 min readPublishing EthicsHidden Prompts in Manuscripts: What the AI Peer Review Exploit Means for Medical Authors
Researchers discovered 18 manuscripts on arXiv in July 2025 that contained hidden text instructions designed to manipulate AI-assisted peer review. A JAMA Network Open study found acceptance rates jumping from 0% to nearly 100% with invisible injection. Here is what medical authors need to understand about this emerging threat to research integrity.
16 min readPublishing EthicsThe AI Letter Flood: What the Surge in AI-Generated Correspondence Means for Medical Authors
From 2023 to 2025, roughly 8,000 authors moved from publishing no letters to publishing many, making up 3% of active authors but 22% of all letters in journals including The Lancet and NEJM. Here is what the AI-generated correspondence surge means for researchers whose work gets commented on and for anyone who reads the scientific record.
16 min read