A retraction notice contains a strange little census. Alongside the title, journal, country, subject and reason, it also contains a byline. That byline is not just a list of names. It is a clue.
A single-author retraction often smells different from a twelve-author retraction. A two-author computer-science paper retracted after a peer-review investigation belongs to a different ecosystem than a fourteen-author biomedical paper retracted after image concerns. A 60-author clinical or COVID paper is different again.
Using the uploaded Retraction Watch CSV, I parsed author counts from the semicolon-separated Author field. I treated each semicolon-separated entry as one author entity, so consortium names or study groups listed as one entry count as one entity. This is a practical approximation, not a perfect bibliometric author disambiguation.
The dataset contained:
| Dataset slice | Records |
|---|---|
| Valid publication and notice dates | 70,589 |
| Valid dates plus usable author counts | 70,473 |
| Non-conference records with usable author counts | 56,428 |
For the main analysis, I focused on non-conference records because conference-proceedings batches strongly distort author-count and retraction-speed patterns.
The central result:
Author count appears to predict retraction timing in raw data, but the deeper story is that author count predicts what kind of problem the paper had.
More authors usually means more image/data/biomedical-style retractions. Fewer authors often means more peer-review, paper-mill, plagiarism, or proceedings-cleanup style retractions. The author count is not the disease. It is the footprint of the habitat. š§Ŗš
1. The raw pattern: more authors, slower retraction, until the extreme tail
In non-conference records, the median author count was 4, and the mean was 4.58. Single-author records made up 14.7%, papers with more than 10 authors made up 5.25%, and papers with more than 20 authors were rare, only 0.34%.
The median retraction lag rises from 1.30 years for single-author papers to 2.52 years for papers with 11 to 20 authors. Then the very large-author group drops slightly, mostly because many very large bylines are linked to fast corrections, retract-and-replace events, COVID-era papers, or large collaborative notices.
Non-conference records only. The 11-20 author bucket has the longest median publication-to-notice lag.
Calculated from the uploaded Retraction Watch CSV.
Raw correlation supports the first impression. Across non-conference records, the Spearman correlation between author count and retraction lag was Ļ = 0.151, with a very small p-value. Across all valid records, including conference material, it was Ļ = 0.228.
But this is the tempting trapdoor.
When I fitted a simple adjusted model for log retraction lag, controlling for notice year, society status, subject categories, and broad reason tags, the independent author-count effect almost vanished: p = 0.684. In the same model, reason tags mattered much more. Image concerns, fraud/misconduct, and plagiarism/duplication were associated with longer lag; paper-mill/peer-review issues were associated with shorter lag.
So the correct interpretation is not:
“More authors cause slower retractions.”
It is:
“Many-author papers are more often found in subjects and reason categories that take longer to investigate.”
Authorship is a proxy. It is the smoke, not necessarily the fire.
2. Author count predicts the reason profile very strongly
The strongest pattern in the data is not timing. It is failure mode.
Single-author and low-author records are dominated by peer-review, paper-mill, and process-related tags. As author count rises, image concerns and fraud/misconduct tags rise sharply.
Non-conference records only. Reason categories overlap, so percentages do not sum to 100.
Calculated from the uploaded Retraction Watch CSV.
This is the big result.
| Author bucket | Dominant signal |
|---|---|
| 1 author | Paper-mill/peer-review/process issues, plagiarism/duplication |
| 2-3 authors | Still strongly peer-review/paper-mill dominated |
| 4-5 authors | Mixed zone |
| 6-10 authors | Image and data concerns rise |
| 11-20 authors | Highest image and fraud/misconduct share |
| >20 authors | Rare, heterogeneous, often large collaborations or clinical/public-health papers |
A logistic model confirmed this pattern after controlling for society status, subject and notice year. For every e-fold increase in author count, meaning roughly multiplying the number of authors by 2.7:
| Outcome tag | Direction with higher author count |
|---|---|
| Paper-mill/peer-review/AI tag | Lower odds, OR = 0.637 |
| Image concern | Higher odds, OR = 1.95 |
| Fraud/misconduct | Higher odds, OR = 1.44 |
| Plagiarism/duplication | Slightly lower odds, OR = 0.937 |
That is a useful signature. Low-author retractions often look like publication-process failures. Many-author retractions more often look like data, image, and biomedical-forensic failures.
3. Time trend: the median author count is stable, but the composition changes
The median author count in non-conference retraction records is surprisingly stable, usually around 4. But underneath that calm median, the composition wriggles like a box of eels.
The share of single-author records spikes in years associated with batch corrections and low-author journal clusters. The share of >10-author records rises in some recent years, especially 2022, 2024, 2025 and partial 2026.
Non-conference records only. The year 2026 is partial in the uploaded file.
Calculated from the uploaded Retraction Watch CSV.
The 2023 pattern is especially telling: the year had many retractions, but the >10-author share dropped to 2.98%, while single-author records rose to 18.49%. That fits the earlier observation that 2023 was heavily shaped by publisher-wide and paper-mill/special-issue cleanup.
The 2024 and 2025 pattern is different: fewer single-author records and more >10-author records. That suggests a shift toward biomedical, image/data, clinical, review, or multi-author investigations.
A useful era summary:
| Notice era | Median authors | Single-author share | >10-author share |
|---|---|---|---|
| ≤2009 | 4 | 13.5% | 3.9% |
| 2010-2014 | 4 | 11.8% | 4.8% |
| 2015-2019 | 4 | 17.9% | 5.3% |
| 2020-2026 | 4 | 14.2% | 5.4% |
So the median hardly moves, but the tails do.
4. Subject matters: humanities and social sciences are low-author worlds, biology and medicine are many-author worlds
Author count is strongly shaped by subject.
Biology and medicine have higher author counts, while humanities and social sciences have many single-author records. This is expected from normal field culture, but the retraction patterns mirror it beautifully.
Non-conference records only. Humanities and social sciences have many single-author records; biology and medicine have more large teams.
Calculated from the uploaded Retraction Watch CSV. Subject categories can overlap.
The subject-reason interaction is important:
| Subject | Median authors | Dominant author-count interpretation |
|---|---|---|
| Biology/life sciences | 5 | More image/data/fraud signals, longer investigations |
| Health sciences/medicine | 5 | Multi-author clinical and biomedical work, mixed faster and slower pathways |
| Physical sciences/engineering | 4 | Mixed, includes paper-mill/special-issue and materials-image patterns |
| Social sciences | 2 | Many low-author records, peer-review/process and plagiarism signals |
| Humanities | 1 | Strong single-author culture, plagiarism/duplication and process issues |
| Environmental sciences | 3 | Many low-author batch-like records, but with some multi-author ecological/climate studies |
The adjusted correlations also show this. Author count correlated with lag in biology, medicine, physical sciences and environmental sciences. But it did not meaningfully correlate with lag in social sciences or humanities. In those fields, the author-count range is too compressed, and the retraction machinery is often driven by plagiarism, policy, or peer-review concerns rather than image/data forensics.
5. Country-specific patterns: author count is a collaboration fingerprint
Country analysis needs caution. I used an exploded-country approach: if a record listed China and the United States, it counted once for China and once for the United States. This does not assign responsibility. It maps affiliation presence.
The country patterns are striking.
Top countries by non-conference country-paper occurrences. Multi-country papers are counted once for each country listed.
Calculated from the uploaded Retraction Watch CSV.
Several country signatures emerge.
Russia: the single-author and plagiarism/duplication signature
Russia has a median author count of 2, with 43.4% single-author records. Its plagiarism/duplication/copyright theme is very high, about 72.1% in this non-conference country slice. This is a very different profile from image-heavy biomedical retractions.
Italy, France, Germany, Canada and the United States: many-author long-tail worlds
Italy has the highest >10-author share among the large country groups shown: 23.7%. France, Germany, Canada and the United States also have high >10-author shares.
These countries also have substantial biomedical, clinical, institutional, and long-tail correction profiles. Many-author papers here are often not paper-mill style records. They are more likely to be clinical, biomedical, collaboration-heavy, or image/data investigation records.
Saudi Arabia and Pakistan: many-author, multinational collaboration signatures
Saudi Arabia and Pakistan have high median author counts, both around 6, and high >10-author shares. They also have very high multinational shares in the earlier country analysis. In other words, their author-count pattern is partly a collaboration-network pattern.
China and India: huge counts, moderate author numbers, strong process signals
China has the largest record count by far, median 4 authors, with 14.8% single-author records and only 3.9% >10-author records. Its non-conference records are heavily associated with peer-review/paper-mill/publisher-investigation themes.
India has median 4 authors, low single-author share (4.6%), and modest >10-author share (4.3%). Its profile is also strongly shaped by peer-review and publisher-investigation clusters.
So again, author count is not national character. It is publication ecology.
6. Publishers and journals: the author-count worlds are completely different
The strongest author-count patterns appear when we look at publisher and journal ecosystems.
Some publishers have many low-author records associated with paper-mill or peer-review cleanup. Others have higher-author biomedical, clinical, or image-heavy portfolios.
Non-conference records only. Selected publishers with large record counts are shown.
Calculated from the uploaded Retraction Watch CSV.
Low-author, fast, process-heavy journal clusters
These include many paper-mill or peer-review-heavy journals:
| Journal | Median authors | Single-author share | Median lag | Main signal |
|---|---|---|---|---|
| Arabian Journal of Geosciences | 1 | 52.9% | 0.30 years | Peer-review/process |
| Journal of Environmental and Public Health | 2 | 48.6% | 0.98 years | Peer-review/process |
| Wireless Communications and Mobile Computing | 2 | 38.9% | 1.16 years | Peer-review/process |
| Security and Communication Networks | 2 | 36.0% | 1.38 years | Peer-review/process |
| Computational Intelligence and Neuroscience | 2 | 27.6% | 1.25 years | Peer-review/process |
These are not typical “old lab-data investigation” retractions. They look like special-issue, publisher-audit, paper-mill, or peer-review pipeline failures.
Many-author, slower, biomedical/data-heavy clusters
These include:
| Journal | Median authors | >10-author share | Median lag | Main signal |
|---|---|---|---|---|
| PLoS One | 6 | 16.4% | 4.31 years | Mixed, image/data, peer-review |
| Scientific Reports | 6 | 15.0% | 1.96 years | Image/data and mixed integrity issues |
| Journal of Biological Chemistry | 6 | 8.5% | 7.32 years | Image-heavy, misconduct-heavy |
| Bioscience Reports | 5 | 4.6% | 3.13 years | Image and batch signals |
| Journal of Crohn’s and Colitis | 7 | 17.4% | 11.98 years | Long-lag clinical/review-like correction |
| Cochrane Database of Systematic Reviews | 4 | 1.2% | 8.32 years | Review lifecycle corrections |
This contrast is one of the strongest in the analysis. A two-author paper in a paper-mill-heavy journal and a twelve-author biomedical paper in an image-heavy journal are both “retracted,” but they belong to different weather systems.
7. Society vs non-society journals: society retractions have larger teams
Using the conservative society-linked publisher classification from the previous analysis, society-linked non-conference records had:
| Group | Records | Median authors | Mean authors | Single-author share | >10-author share | Median lag |
|---|---|---|---|---|---|---|
| Society-linked | 4,896 | 5 | 5.86 | 4.7% | 9.7% | 3.00 years |
| Non-society / unclassified | 51,532 | 4 | 4.45 | 15.6% | 4.8% | 1.67 years |
Society-linked records have larger author teams and longer retraction lags. They are also more image-heavy and fraud/misconduct-heavy, while non-society/unclassified records are more peer-review/paper-mill-heavy.
Non-conference records only. Society-linked records are more concentrated in 6-10 and 11-20 author buckets.
Calculated from the uploaded Retraction Watch CSV.
This explains much of the society-journal pattern from the previous post. Society-linked retractions are not necessarily more numerous, but they are more likely to sit in biomedical, biochemical, chemistry, society-proceedings, and higher-team-size journals. That produces a different retraction clock.
Within society-linked records:
| Author bucket | Median lag | Image concerns | Fraud/misconduct |
|---|---|---|---|
| 1 author | 1.97 years | 6.5% | 13.9% |
| 2-3 authors | 2.38 years | 27.2% | 23.4% |
| 4-5 authors | 2.85 years | 38.8% | 27.8% |
| 6-10 authors | 3.51 years | 53.5% | 26.0% |
| 11-20 authors | 3.79 years | 52.6% | 30.4% |
Within non-society/unclassified records, the same direction exists, but the paper-mill/peer-review signal is much stronger in low-author buckets.
So society status modifies the author-count interpretation:
In society journals, more authors usually means more image/data/forensic correction.
In non-society/unclassified journals, low-author retractions are heavily shaped by peer-review and paper-mill correction.
8. The >20-author exception: giant bylines are rare and heterogeneous
The largest author-count records are fascinating exceptions.
The maximum parsed author count in the dataset was 88, from a Science paper on the emergence and spread of the SARS-CoV-2 Omicron variant in Africa. It was retracted very quickly, with a lag of about 0.05 years, and the reason involved contamination/materials and unreliable conclusions.
Other very large bylines include:
| Approx. authors | Type of record | Typical pattern |
|---|---|---|
| 80+ | COVID/genomics/public-health collaborations | Fast correction possible |
| 80+ | Mendelian randomization/large consortium study | Data concerns, retract-and-replace or updated notice |
| 60+ | Clinical trial/COVID ICU papers | Removal, retract-and-replace, date/notice complexity |
| 50+ | Nature/biomedical consortium papers | Institutional/data/manipulation concerns |
This is why the >20-author bucket does not simply continue the lag increase. Very large bylines are a special species. They include consortia, public-health surveillance, clinical collaborations, and multi-country computational studies. Their corrections can be rapid if the problem is centralized, obvious, or administrative.
The author-count curve therefore has a bend:
Retraction lag increases from 1 author to 11-20 authors, but the extreme mega-author bucket is too rare and too heterogeneous to behave like a simple continuation.
9. Hypothesis-by-hypothesis evaluation
| Hypothesis | Result | Evidence |
|---|---|---|
| More authors mean slower retraction | Raw yes, adjusted no | Spearman Ļ = 0.151 in non-conference records, but adjusted model author effect p = 0.684 |
| Author count predicts reason type | Strongly supported | Low-author records are peer/paper-mill-heavy; many-author records are image/data/fraud-heavy |
| Single-author records are mostly in humanities/social sciences | Partly true, but not enough | Humanities and social sciences are single-author-heavy, but many low-author records also come from publisher batch corrections |
| Many-author records are mostly biomedical/clinical | Broadly supported | Biology and medicine have median 5 authors and the largest >10-author shares |
| Country patterns differ by author count | Supported | Russia has a single-author/plagiarism signature; Italy, France, Germany, US, Saudi Arabia and Pakistan have stronger many-author signatures |
| Publisher/journal ecosystems differ | Strongly supported | Hindawi/Springer/IOS low-author process-heavy clusters vs PLoS/BMC/Cell Press/JBC many-author data/image clusters |
| Society journals have different author-count behavior | Supported | Society-linked records have higher median authors, fewer single-author records, more >10-author records and longer lag |
| Mega-author papers behave like ordinary many-author papers | Not supported | >20-author records are rare, mixed and often corrected faster than 11-20 author papers |
10. What the author count really tells us
A byline is not just a list of contributors. In this dataset, it behaves like a weak but useful diagnostic.
| Author-count pattern | Likely retraction ecology |
|---|---|
| 1 author | Plagiarism, peer-review/process, humanities/social sciences, batch cleanup |
| 2-3 authors | Paper-mill/peer-review-heavy, computing/engineering and special-issue clusters |
| 4-5 authors | Transition zone, mixed problems |
| 6-10 authors | Biomedical, image/data, society-journal and lab-science records rise |
| 11-20 authors | Strong image/data/fraud/institutional-investigation signal |
| >20 authors | Consortia, clinical/public-health, large collaborations, heterogeneous and rare |
The most important conclusion is that author count is a context marker. It points toward field, journal, publisher, collaboration structure and reason category. It does not by itself tell us whether a paper is fraudulent, careless, or unlucky.
A one-author retraction may be plagiarism.
A three-author retraction may be a paper-mill node.
A seven-author retraction may be duplicated western blots.
A fifteen-author retraction may be a clinical or biomedical investigation.
An eighty-author retraction may be a fast correction in a consortium study.
The byline is a map legend, not the map.
Data cautions
Several caveats matter:
- Author count was parsed from the Retraction Watch author field, using semicolon-separated entries. Group authors may be counted as one entity.
- Records are not always unique scientific articles, because some entries are updated notices, expressions of concern, corrections, or retract-and-replace events.
- Conference records were excluded from the main analysis, because conference-proceedings batches strongly distort author-count and timing.
- Country fields were exploded, so multinational papers count once for each country listed.
- No publication denominator is available, so this analysis describes retraction records, not retraction rates per published paper.
- Reason categories overlap, so percentages do not sum to 100.
Final thought: the byline is the paper’s seismograph
The number of authors on a retracted paper does not tell us guilt. It tells us terrain.
Small bylines often sit in fast-moving process failures: peer review, paper mills, special issues, plagiarism, metadata cleanup. Medium-to-large bylines sit more often in slow-moving forensic failures: images, data, misconduct investigations, biomedical records, society journals, and clinical ecosystems. Very large bylines are rare exceptions, often shaped by consortia and centralized corrections.
So the authorship pattern is not a morality score. It is a seismograph.
It tells us whether the tremor came from a paper-mill factory floor, a humanities desk, a computational special issue, a biochemical blot archive, a clinical collaboration, or a giant pandemic consortium.
The byline, quiet little row of names that it is, carries the crackle of the whole publishing ecosystem. š¬š