Tuesday, August 4, 2026

Why Tempo and Mode in Evolution Still Matters

Why should modern students read an older work like Tempo and Mode in Evolution?

Because many of its questions are still alive.

How fast does evolution happen? Are evolutionary rates constant or variable? Do some lineages evolve faster than others? How do we connect fossil patterns with genetic mechanisms? What counts as evidence for gradual change, rapid change, stasis, or branching evolution?

These questions remain central in evolutionary biology, even in the genomic era. Today, we can sequence genomes, estimate divergence times, detect selection, study developmental pathways, and model trait evolution statistically. Yet the fossil record still provides something no genome alone can offer: direct evidence of morphology across deep time.

Simpson’s chapter is valuable because it teaches scientific caution. It warns against easy stories. Fossils must be interpreted carefully. Rates must be defined clearly. Traits must be measured thoughtfully. Patterns must not be confused with mechanisms.

The chapter also offers a vision of synthesis. Evolutionary biology becomes strongest when palaeontology, genetics, systematics, ecology, and statistics work together. Fossils show long-term outcomes. Genetics explains inheritance and variation. Ecology explains selective context. Systematics reveals relationships. Statistics help separate signal from noise.

In that sense, Tempo and Mode in Evolution is not merely about old fossils. It is about how to think scientifically across time scales.

Evolution has no single rhythm. Some lineages creep. Some sprint. Some pause. Some branch into wild experiments. Some vanish. Some leave only fragments, a tooth here, a skull there, a clue pressed into stone.

To study tempo and mode is to listen carefully to that deep-time orchestra.

The music is ancient, but the questions still hum. 🦴🧬

Monday, August 3, 2026

The Human Proteome’s Missing Pieces: Why Thousands of Proteins Remained Hidden and How Scientists Planned to Find Them

The human genome gave biology an extraordinary catalogue of approximately 20,000 protein-coding genes. Yet a gene sequence is not the same thing as proof that its corresponding protein actually exists. A gene may be transcribed only in a handful of cells, activated briefly during development or disease, translated into a tiny peptide, embedded deep inside a membrane, or produced at concentrations far below the reach of routine experiments.

This gap between what the genome predicts and what experiments can confidently detect lies at the heart of the article “Accelerating the search for the missing proteins in the human proteome,” published in Nature Communications in 2017.

The article reviews the progress of the Human Proteome Project, explains why many predicted proteins were still considered “missing,” examines difficult examples such as olfactory receptors, prestin and interleukin-9, and proposes a community knowledge platform called MissingProteinPedia. Its central argument is both practical and philosophical:

A protein may be missing from a high-stringency proteomics database without being absent from human biology.

The challenge, therefore, is not merely to produce more mass spectra. It is to connect molecular clues scattered across proteomics, genetics, transcriptomics, physiology, pharmacology, microscopy and decades of published or unpublished research.


From the human genome to the human proteome

Sequencing the human genome revealed the instructions encoded in DNA. The Human Proteome Project, or HPP, was created to determine how those instructions are translated into the functioning molecular machinery of human life.

Launched by the Human Proteome Organization in 2010, the HPP had two broad objectives.

First, it sought to complete a reliable protein parts list for the human body. Ideally, this would include at least one experimentally supported protein product from every protein-coding gene, along with splice variants, post-translational modifications and amino-acid variations.

Second, it aimed to make proteomics a mature partner to genomics in biomedical and clinical research. That required more than detecting proteins. It required standardized repositories, reproducible analytical pipelines, accepted confidence thresholds and community-wide rules for deciding when a protein had truly been observed.

The HPP was organized into two overlapping branches:

  • The Chromosome-centric Human Proteome Project, or C-HPP, divided the search according to chromosomes.
  • The Biology and Disease Human Proteome Project, or B/D-HPP, investigated proteins through tissues, pathways, biological systems and diseases.

These efforts were supported by three major pillars:

  1. Mass spectrometry
  2. Affinity reagents, particularly validated antibodies
  3. Knowledgebases that integrate and curate evidence

Databases and resources such as neXtProt, PeptideAtlas, ProteomeXchange, the Human Protein Atlas and GPMDB became essential parts of this ecosystem. Together, they created the scaffolding needed to turn enormous quantities of experimental data into defensible protein identifications.


What exactly is a “missing protein”?

The word missing can be misleading. It does not necessarily mean that the protein does not exist. Instead, it means that the available evidence did not satisfy the HPP’s accepted criteria for confirmation at the protein level.

The article describes the neXtProt protein-existence system, which placed human proteins into five categories.

PE1: Evidence at the protein level

These proteins had strong experimental evidence, derived from approaches such as mass spectrometry, validated antibodies, Edman sequencing or structural determination.

Under the stringent mass-spectrometry criteria adopted in 2016, a typical PE1 assignment required at least two highly confident, uniquely mapping peptides of nine or more amino acids that were not nested within one another.

PE2: Evidence at the transcript level

The corresponding RNA had been detected, but sufficiently strong evidence for the protein itself was lacking.

PE3: Evidence from homologous species

A related protein had been found in another organism, suggesting that the human protein probably exists, but direct human transcript or protein evidence was insufficient.

PE4: Predicted evidence

The protein was predicted from genomic information, but experimental support remained very limited.

PE5: Dubious or questionable proteins

These entries were considered unlikely to produce functional proteins, often because of disrupted genes, pseudogene-like characteristics or missing transcriptional features.

The article defines PE2, PE3 and PE4 proteins collectively as the missing proteins. PE5 proteins were excluded because they were considered dubious rather than simply undetected.

Using the February 2016 neXtProt release, the authors reported:

Protein-existence categoryNumber of proteins
PE116,518
PE22,290
PE3565
PE494
PE5588
Total20,055

Thus, 2,949 proteins were classified as PE2–PE4 and therefore considered missing under the HPP framework.

The comparison shown in Box 1 on page 2 also demonstrates genuine progress. Between 2013 and 2016, PE1 assignments increased from 15,649 to 16,518. This occurred even while the criteria became stricter and hundreds of previously accepted PE1 proteins were downgraded.


Why high-stringency evidence matters

Proteomic experiments produce immense volumes of data. Mass spectrometry breaks proteins into peptides, measures their mass-to-charge characteristics and compares the resulting fragmentation spectra against candidate sequences.

The difficulty is that a plausible peptide-spectrum match is not automatically correct.

False assignments can arise from:

  • Noisy or incomplete spectra
  • Very short peptide sequences
  • Closely related proteins with shared peptides
  • Sequence variants
  • Leucine/isoleucine ambiguity
  • Incorrect database assumptions
  • Multiple testing across huge search spaces
  • Inadequate false-discovery-rate control

For this reason, the HPP moved toward increasingly rigorous standards. The article discusses recommendations such as a 1% protein-level false discovery rate, accompanied by peptide-level and peptide-spectrum-match-level error estimates. Raw spectra should also be deposited publicly so that claims can be independently re-examined.

The requirement for two unique peptides of at least nine amino acids greatly lowers the probability of a random or ambiguous assignment. Yet the authors emphasize an important caveat: even two qualifying peptides do not make an identification absolutely unquestionable. Confidence improves dramatically, but scientific evidence remains open to reassessment.

This tension became visible after two large “draft human proteome” studies reported many previously undetected proteins using criteria that differed from HPP standards. Some of their assignments, particularly those involving olfactory receptors, were later criticized for marginal spectra, short peptides, ambiguous matching and insufficient false-discovery control.

The resulting dispute was productive. It exposed weaknesses in large heterogeneous datasets and pushed the field toward clearer metrics, transparent workflows and stronger deposition standards.


Why can a real protein remain invisible?

The article’s most important contribution is its explanation of why certain proteins repeatedly escape standard detection. The proteome is not a static warehouse. It is closer to a city at night: some buildings glow continuously, while others illuminate one room for a few seconds under very specific conditions.

1. Extremely low abundance

Some proteins are present at only a few copies per cell. Their peptide signals may be drowned out by abundant structural proteins, serum proteins or housekeeping proteins.

2. Highly restricted tissue expression

A protein may occur only in the inner ear, a small brain nucleus, a particular epithelial layer or a few sensory neurons. Whole-tissue analysis can dilute its signal almost beyond recovery.

3. Narrow temporal expression

Some proteins are produced only:

  • During a particular developmental stage
  • After immune activation
  • Under environmental stress
  • During infection
  • In a specific disease state
  • At a particular point in a cell’s differentiation

A sample taken at the wrong time may contain no detectable protein.

4. Difficult subcellular localization

Proteins confined to cilia, axon terminals, extracellular vesicles, membrane microdomains or other specialized structures may require targeted enrichment.

5. Membrane association and hydrophobicity

Transmembrane proteins are difficult to dissolve, purify and digest. Their trypsin cleavage sites may be buried in the lipid bilayer or shielded by neighbouring molecules.

6. Poor compatibility with trypsin digestion

Bottom-up proteomics commonly uses trypsin to cut proteins into peptides. Some proteins simply do not generate two suitable, uniquely mapping tryptic peptides of the required length.

7. Small proteins and bioactive peptides

Short secreted peptides may be biologically crucial but too small to satisfy rules designed for conventional proteins. The article cites the orexigenic neuropeptide QRFP as an example. A peptide can control appetite or cell signalling and still remain excluded from PE1 because it cannot generate two independent peptides of nine amino acids.

8. Extensive post-translational modification

Glycosylation, phosphorylation, cleavage and other modifications can alter peptide masses and complicate database matching.

9. Insolubility or cross-linking

Proteins that form dense extracellular structures, complexes or cross-linked assemblies may resist conventional extraction.

10. Genuine absence under ordinary conditions

Some predicted gene products may not be translated in normal physiology. Others may be expressed only under rare circumstances, or perhaps not at all.

The authors therefore recommend targeted approaches such as subcellular enrichment, better fractionation, analysis of rare tissues, improved membrane-protein workflows, alternative proteases, higher-sensitivity instruments, validated antibodies and examination of samples under diverse physiological and pathological conditions.


Which protein families contained the largest gaps?

The authors performed bioinformatics analyses of missing proteins according to their families, domains, biological classes, pathways and evolutionary relationships.

The chart on page 5 identifies several especially prominent categories:

  • Olfactory receptors
  • Zinc-finger proteins
  • Other transmembrane proteins
  • Coiled-coil-domain proteins
  • Homeobox proteins
  • Keratin-associated proteins
  • Sperm-related and testis-expressed proteins
  • Solute carriers
  • β-defensins
  • PRAME-family proteins
  • Other G-protein-coupled receptors

The analysis showed encouraging progress for most major groups between 2013 and 2016. The glaring exception was the olfactory-receptor family, whose representation among missing proteins actually increased proportionally.

The comparison of UniProt families on page 5 also revealed that some families were dominated by proteins lacking high-stringency evidence. GPCR type 1 proteins were particularly enriched among PE2–PE4 entries. Apart from a few families such as Kruppel C2H2 zinc-finger proteins and peptidase C19 proteins, between 50% and 95% of the members of many top-ranked missing-protein families remained unconfirmed.

This pattern suggests that missingness is not random. Certain biological architectures create systematic blind spots. Large membrane-receptor families, recently expanded gene families and proteins with highly specialized expression are naturally harder for conventional proteomics to capture.


The olfactory receptor problem

Olfactory receptors became the article’s central case study because they represented the largest family of missing proteins.

These proteins belong to the G-protein-coupled receptor superfamily. GPCRs respond to an astonishing range of signals, including photons, neurotransmitters, hormones, nutrients, metals and volatile chemicals. They are also among the most important classes of pharmaceutical targets.

The article describes five major GPCR branches:

  1. Rhodopsin or class A
  2. Secretin
  3. Adhesion or class B
  4. Glutamate or class C
  5. Frizzled and taste receptor 2

The phylogenetic analysis on page 6 shows missing GPCRs distributed throughout these branches, but the greatest concentration lies within the rhodopsin branch containing olfactory receptors.

At the time of analysis, the genome contained approximately 480 olfactory-receptor genes. Twelve were classed as PE5. The remainder encoded 411 unique proteins, of which only two were listed as PE1 and 409 remained PE2–PE4.

In other words, almost the entire human olfactory-receptor repertoire was missing by HPP standards.


Functional evidence versus proteomic proof

The scarcity of mass-spectrometry evidence did not imply biological inactivity.

Functional screening had identified agonists for numerous olfactory receptors. One study tested receptors against 73 candidate ligands and found agonists for 27 receptors, including 18 that had previously been orphans.

Such experiments show that a receptor can:

  • Be expressed in a heterologous system
  • Reach the cell membrane
  • Bind or respond to an odorant
  • Activate downstream signalling

That is compelling biological evidence. Yet it may not satisfy a formal PE1 mass-spectrometry criterion.

The article highlights an inconsistency in the 2016 classifications. OR1D2 and OR2AG1 were listed as PE1, but the reported supporting evidence did not appear to meet the contemporary HPP mass-spectrometry thresholds. OR1D2 lacked MS or antibody evidence, while OR2AG1 was associated with only a single seven-amino-acid peptide. Meanwhile, another receptor with comparable functional evidence remained PE4.

The lesson is not that functional studies are unreliable. It is that the rules for incorporating non-MS evidence had not been standardized as clearly as the rules for mass spectrometry.


Reanalysing more than 122,000 peptide-spectrum entries

To investigate olfactory receptors more systematically, the authors searched public proteomic repositories including GPMDB, PRIDE, ProteomicsDB, MAXQB and Human ProteinPedia.

They aggregated 122,717 peptide-spectrum entries of at least seven amino acids.

The filtering process was severe:

  1. Removal of non-unique and decoy peptides left 4,751 potentially proteotypic olfactory-receptor peptides.
  2. Only 286 carried a high-confidence score from search engines such as SEQUEST, Mascot or MaxQuant.
  3. Manual examination of spectral quality reduced the set to 64 strong spectra representing 24 peptides.
  4. After merging overlapping peptides, only 23 unique olfactory-receptor peptides remained.

These data provided some mass-spectrometry evidence for 23 of the 409 missing olfactory receptors, approximately 5.6%.

However, none had the two qualifying peptides needed to satisfy the strict PE1 standard. Fourteen proteins were supported by only one seven- or eight-amino-acid peptide, while nine had one peptide longer than nine amino acids.

The authors describe these receptors as proteins “waiting in the wings.” The available evidence pointed researchers toward promising targets, but additional confirmation was required.

The analysis is an excellent demonstration of the difference between a clue and a completed identification. A single credible spectrum may not close the case, but it can reveal where to search next.


Olfactory receptors are not confined to the nose

Another reason olfactory receptors should not be dismissed is their expression outside nasal tissue.

The article notes evidence for olfactory-receptor expression in multiple epithelial tissues, where they may perform broader chemosensory roles. Thus, searching only the olfactory epithelium may be unnecessarily restrictive.

A receptor could be present:

  • In very few sensory neurons
  • On cilia that are difficult to isolate
  • In non-nasal epithelial tissues
  • At low abundance
  • Under specific hormonal, metabolic or disease conditions

The appropriate search strategy is therefore not simply “analyse more nose tissue.” It requires combining functional biology, tissue-expression information and targeted proteomic design.


Chromosome 7 as a miniature map of the problem

The Australian and New Zealand C-HPP teams focused on chromosome 7. The chromosome map on page 8 plots 757 PE1 proteins and 139 PE2–PE4 proteins along its length.

The missing proteins were not restricted to gene-poor regions. In fact:

  • 56% came from regions described as having high gene density
  • 12% came from moderately dense regions
  • 25% came from low-to-moderate-density regions
  • Only 1.5% came from low-density regions

This observation argues against a simplistic explanation that missing proteins merely originate in poorly populated or poorly annotated genomic regions. Their invisibility is more likely related to expression level, tissue specificity, protein chemistry or sampling conditions.

Among the chromosome 7 missing proteins, 27 were GPCRs. These included:

  • 15 olfactory receptors
  • 6 taste-related receptors
  • 4 orphan GPCRs
  • The serotonin receptor 5-HT5A
  • Metabotropic glutamate receptor 8

The article then uses selected receptors to show how rich biological evidence can coexist with absent high-stringency protein detection.


5-HT5A: functional, but elusive

The HTR5A gene encodes the serotonin receptor 5-HT5A.

Evidence discussed in the article includes:

  • Human brain mRNA detected by in situ hybridization and PCR
  • G-protein activation and inhibition of adenylyl cyclase in expression systems
  • Altered behaviour in knockout mice
  • Changed responses to the serotonin-related compound LSD in knockout animals

Yet the authors found no convincing human protein localization by immunohistochemistry or western blot.

The likely explanation is not necessarily non-existence. The receptor may be expressed at extremely low levels in narrowly restricted brain regions, making it difficult to detect in bulk tissue.


mGlu8: low expression and complex biology

GRM8 encodes metabotropic glutamate receptor 8.

The article reports:

  • Functional signalling in heterologous expression systems
  • Low and anatomically restricted mRNA in human brain
  • Expression reported in cancer cells, hippocampal cells and astrocytes
  • Associations with epilepsy and multiple sclerosis tissue
  • Physiological consequences after gene deletion in mice

Its large size, complex gene structure and possible alternative splicing may create multiple protein forms and further complicate detection.

Again, the biology looks active, while the conventional proteomic signal remains faint.


GPR22: a protein hiding in conditional biology

GPR22 provides an even more uncertain case.

Its mRNA had been detected in human heart and brain, but no ligand had been identified. In experimental systems, its unusual AT-rich coding sequence appeared to interfere with efficient expression. Signalling could be restored after modifying the sequence composition.

Knockout mice did not display an obvious baseline phenotype. However, under cardiac stress, animals lacking GPR22 developed heart failure more rapidly, suggesting that its role may emerge only under particular physiological challenges.

This illustrates a deeper point: some proteins may look unimportant under routine laboratory conditions because their function is conditional. Their biological significance may appear only during injury, infection, stress or disease.


Prestin: a well-known protein that remained technically “missing”

Prestin, encoded by SLC26A5, is one of the article’s most striking examples.

It is widely described as the motor protein of cochlear outer hair cells and is central to the mechanics of hearing. The article found abundant indirect and functional evidence:

  • More than 90 peer-reviewed publications
  • Numerous commercially available antibodies
  • Known chloride and bicarbonate relationships
  • Human variants associated with deafness and other phenotypes
  • Multiple transcripts
  • Copy-number variants in clinical databases
  • Experimental genetic tools in model organisms

Yet prestin remained PE2 because high-stringency endogenous MS or accepted antibody evidence was unavailable.

Why?

Prestin occupies a near-perfect hiding place:

  1. It is expressed in cochlear outer hair cells.
  2. These cells are rare.
  3. Human inner-ear tissue is extremely difficult to obtain.
  4. Only a few hundred outer hair cells may be collected by specialized microdissection.
  5. Prestin is a hydrophobic membrane protein.
  6. Membrane localization complicates extraction and tryptic digestion.
  7. Its concentration is far below what routine proteomic workflows prefer.

Synthetic peptides corresponding to prestin could be produced and measured, but synthetic standards do not prove that endogenous prestin peptides have been recovered from human tissue.

Prestin reveals a classification paradox. Biologists may regard the protein as firmly established, yet the HPP’s specific evidence machinery may still classify it as missing. This is not necessarily a flaw in quality control. It shows that biological knowledge and assay-specific proof answer related but different questions.


Interleukin-9: when experimental design determines visibility

Interleukin-9 offers a different kind of challenge.

Small secreted signalling proteins are often:

  • Produced transiently
  • Released only after stimulation
  • Present at low concentrations
  • Heavily modified
  • Surrounded by vastly more abundant extracellular proteins
  • Too short to yield multiple qualifying tryptic peptides

The authors examined the secretome of activated primary T cells. Conventional secretome studies often use serum-free media to avoid contamination, but serum deprivation stresses cells and can produce a flood of stress- and apoptosis-related proteins.

The authors instead cultured cells for several days in the presence of fetal bovine serum. This created another problem: approximately 95% of detected peptides originated from bovine serum proteins.

After excluding bovine proteins and proteins released by resting human T cells, they identified secretory proteins associated specifically with activated cells, including IL-9.

Their MS analysis detected two IL-9 peptides:

  • YPLIFSR, seven amino acids
  • SLLEIFQK, eight amino acids

The fragmentation spectra are shown in Figure 6 on page 10.

Both peptides were predicted to be unique to IL-9, but neither reached the nine-amino-acid threshold. Thus, the experiment provided biologically meaningful and apparently specific evidence while still falling short of formal PE1 requirements.

IL-9 demonstrates that a fixed peptide-length rule can disadvantage small proteins whose sequence simply cannot produce the required set of tryptic peptides.


MissingProteinPedia: a home for clues that do not fit the final verdict

The article’s proposed solution is MissingProteinPedia, a communal database designed to complement rather than replace the high-stringency HPP system.

The distinction is crucial.

The HPP’s official pipelines answer:

Does this protein satisfy the agreed criteria for high-confidence identification?

MissingProteinPedia would answer:

What does the scientific community know about this protein, and what clues might help us find it conclusively?

The architecture illustrated in Box 2 on page 3 connects information from sources such as:

  • neXtProt
  • Human Protein Atlas
  • PeptideAtlas
  • ProteomeXchange
  • PRIDE
  • PASSEL
  • MassIVE
  • GPMDB
  • ProteomicsDB
  • MaxQB
  • PubMed
  • UniProt
  • GeneCards
  • GeneRIFs
  • ProtAnnotator
  • Individual laboratories

It was also intended to accommodate preliminary, unpublished, proprietary or legacy observations through protected collaboration interfaces.

Text-mining tools could gather literature associated with genes, proteins and synonyms. Users could add annotations, while administrators could curate material before public release.

The platform was envisioned as searchable and sortable by criteria such as chromosome, tissue and keyword.


Low-stringency does not mean low value

Calling MissingProteinPedia “low-stringency” could sound as though it were designed to collect unreliable information. That is not the authors’ intention.

Rather, the platform separates evidence gathering from final adjudication.

A single peptide spectrum may be insufficient for PE1, but it can identify:

  • A promising tissue
  • A likely stimulation condition
  • A useful peptide target
  • A possible splice isoform
  • A candidate antibody
  • A suitable subcellular fraction
  • A disease state in which the protein is enriched

Likewise, an old laboratory notebook, unpublished western blot or commercial antibody result might not independently prove a protein’s existence. Yet, when combined with transcript data, animal phenotypes and targeted MS evidence, it may reveal the experimental route required for confirmation.

MissingProteinPedia was therefore conceived as a hypothesis engine. It would collect the breadcrumbs without pretending that every breadcrumb is the loaf.

The authors explicitly state that the platform would not initially judge submitted evidence by the same standards used for official HPP reclassification. Instead, it would expose possibilities that might otherwise remain hidden in unpublished experiments, commercial records or disciplinary silos.


The paper’s proposed roadmap

The article presents five major recommendations for accelerating completion of the human proteome.

1. Maintain rigorous HPP standards

Researchers and journals should follow current HPP data-submission guidelines and high-stringency reanalysis metrics.

Lowering standards would make the proteome appear complete more quickly, but it would fill the catalogue with questionable assignments.

2. Consolidate mass-spectrometry evidence

All relevant MS data should be deposited in accessible systems such as ProteomeXchange, including datasets from repositories not fully integrated into the HPP pipeline.

Claims involving missing proteins should be accompanied by transparent raw data.

3. Develop formal rules for non-MS evidence

The community should agree on how to evaluate evidence from:

  • Antibodies
  • Functional assays
  • Structural biology
  • Protein interactions
  • Imaging
  • Genetics
  • Pharmacology
  • Cell biology
  • Other experimental methods

Without common rules, equivalent evidence may produce inconsistent classifications.

4. Hold annual evidence-review jamborees

The authors propose community meetings resembling the annotation jamborees used during the Human Genome Project.

At these events, experts could review proposed upgrades and downgrades, examine disputed evidence and document the rationale behind classification decisions.

5. Capture all relevant knowledge in MissingProteinPedia

Every credible clue concerning PE2–PE4 proteins should be assembled in a shared resource, creating a bridge from exploratory evidence to targeted high-stringency validation.


What the figures collectively reveal

The article’s visual evidence tells a coherent story.

Box 1, page 2

The PE classification diagram shows both progress and increasing strictness. More proteins reached PE1 between 2013 and 2016, even as the minimum peptide requirements became harder to satisfy.

Box 2, page 3

The MissingProteinPedia workflow places the proposed resource between official HPP databases, public literature, independent repositories and laboratory-generated evidence. It is designed as a connective layer rather than a competing authority.

Figure 1, page 4

The extrapolation plots estimate how quickly different databases were reducing the proportion of missing proteins. The projections varied substantially, suggesting that completion depended heavily on the evidence source and classification pipeline. The authors noted that some trajectories implied much slower completion than optimistic HPP timelines.

Figures 2 and 3, page 5

These show that missing proteins are enriched in particular families rather than uniformly distributed across the proteome. Olfactory receptors form the largest and most stubborn group.

Figure 4, page 6

The GPCR phylogenetic trees reveal dense clusters of missing proteins, especially among olfactory receptors. The figure also overlays functional ligands and partial peptide evidence, visually demonstrating that “missing” receptors may already carry several kinds of incomplete evidence.

Figure 5, page 8

The chromosome 7 map shows that missing proteins are spread along the chromosome and are not predominantly confined to gene-poor regions.

Figure 6, page 10

The IL-9 spectra illustrate the threshold problem directly: two apparently proteotypic peptides are detected, but both are shorter than the required nine residues.


The article’s deeper scientific message

At first glance, the search for missing proteins appears to be a technical problem in analytical chemistry. The article shows that it is actually a problem of measurement, ontology and scientific governance.

Measurement

Some proteins fall outside the practical detection range of standard workflows because of abundance, chemistry, localization or timing.

Ontology

Scientists must decide what “protein existence” means. Does functional activity count? Does a validated antibody count? Is one unique spectrum sufficient? What about a resolved structure, a disease-causing mutation or a knockout phenotype?

Different fields naturally answer these questions differently.

Governance

A global project requires rules that are accepted, transparent and consistently applied. It must also preserve the ability to revise earlier conclusions.

The HPP’s strict criteria protect the reliability of the human protein catalogue. However, strict criteria can become a narrow funnel through which certain legitimate biological objects cannot easily pass. The article does not advocate abandoning the funnel. It advocates building a richer map around it.


Strengths of the article

Several features make the paper especially valuable.

It avoids equating non-detection with non-existence

This is perhaps its most important intellectual contribution.

It supports its argument with diverse examples

Olfactory receptors, prestin, GPCRs and IL-9 fail detection for different reasons. Together, they show that missingness has multiple causes.

It defends rigorous standards

The authors do not suggest promoting proteins to PE1 merely because circumstantial evidence exists.

It recognizes disciplinary fragmentation

Important evidence may sit in genetics, pharmacology or physiology databases that proteomics pipelines do not routinely inspect.

It proposes an actionable infrastructure

MissingProteinPedia is presented not simply as a concept, but as an integrated platform with literature mining, repository links, user annotation and protected collaboration.


Important limitations and cautions

Because the review was published in January 2017, its numerical summaries describe the state of neXtProt and the HPP primarily in 2016. The figures should therefore be read historically, not as present-day counts.

Its completion projections were based on short-term linear extrapolation. Scientific progress rarely proceeds linearly. Instrument improvements, new sample types, changes in classification rules and database reanalysis can all produce sudden jumps or reversals.

The proposed low-stringency collection model also creates challenges:

  • Weak evidence can accumulate rapidly.
  • Duplicate claims may appear convincing through repetition.
  • Commercial antibodies may lack adequate validation.
  • Unpublished observations can be difficult to reproduce.
  • Contradictory evidence requires visible provenance.
  • Community curation can become a substantial workload.

A successful system therefore needs to distinguish clearly between:

  • Submitted observations
  • Curated evidence
  • Independently replicated findings
  • Official PE classification

The article itself respects this distinction, but any implementation must preserve it carefully.


Final perspective: finding the protein means finding its context

The search for missing proteins cannot be completed by repeatedly analysing the same convenient tissues with the same digestion enzymes and the same database filters.

Researchers must ask more biological questions:

  • In which cells is the protein expressed?
  • At what developmental stage?
  • Under which disease or stress condition?
  • In which organelle or membrane compartment?
  • Which protease will generate observable peptides?
  • Is the active product a short, processed peptide?
  • Does an isoform alter the expected sequence?
  • Can functional or genetic data guide targeted MS?
  • Is the protein present only after a stimulus?

The article’s enduring idea is that context is itself an experimental reagent.

A low-abundance receptor in a rare sensory cell, a cytokine secreted only after activation and a membrane motor confined to the inner ear are not missing in the same way. Each requires a different scientific key.

The Human Proteome Project had already built a powerful high-stringency engine for determining when evidence was strong enough. MissingProteinPedia was proposed as the complementary scouting network, collecting clues from the wider scientific landscape and directing that engine toward the places where hidden proteins were most likely to emerge.

Completing the human proteome, in this vision, is not merely an exercise in filling empty database cells. It is an effort to understand where, when, how and why every protein participates in human biology. That turns the proteome from a parts list into something far richer: a dynamic molecular atlas of what it means to be human. 

Why Some Traits Evolve Faster Than Others

Not all traits evolve at the same speed. This is one of the most important messages emerging from Simpson’s discussion of rates.

A lineage is not a block of clay reshaped uniformly. It is a living system made of parts with different functions, developmental constraints, genetic architectures, and ecological roles. Teeth, limbs, skulls, body size, ornamentation, and internal anatomy may each follow different evolutionary rhythms.

Why might one trait evolve faster than another?

First, selection may act more strongly on some traits. Teeth may respond quickly to dietary change. Limb proportions may shift with habitat use. Body size may change in response to climate, predation, or resource pressures.

Second, some traits may be developmentally constrained. A structure deeply integrated with many other body systems may be less free to vary without harmful side effects.

Third, genetic variation may differ among traits. Some traits may have abundant variation available for selection, while others may be more canalised.

Fourth, the fossil record itself may bias our view. Hard parts, such as teeth and bones, preserve better than soft tissues, behaviour, or physiology. We may think teeth evolve dramatically, partly because teeth are what we can measure most easily.

Simpson’s treatment of rate reminds us that evolution is mosaic. One part of an organism may change while another remains stable. This mosaic evolution is crucial for interpreting fossils. A species may look “primitive” in one trait and “advanced” in another. Evolution does not renovate the whole house at once. Sometimes it remodels the kitchen, reinforces the roof, and leaves the attic full of ancestral furniture.

For students, this is a liberating idea. Evolution is not a straight path from old to new. It is a patchwork of changes, constraints, experiments, and histories.

Different traits carry different clocks.

Sunday, August 2, 2026

Microevolution and Macroevolution: Bridging Two Scales

One of the grand tensions in evolutionary biology is the relationship between small-scale and large-scale change.

Microevolution refers to changes within populations: shifts in allele frequencies, variation, selection, drift, mutation, and gene flow. Macroevolution refers to larger patterns: the origin of species, long-term trends, major morphological transitions, adaptive radiations, and extinction.

Simpson’s work is important because it argues that these should not be treated as separate universes. Large-scale evolution must be somehow connected to processes acting within populations. But the connection is not always simple.

A small genetic change can have large morphological effects. A long period of microevolution may produce only modest visible change. A lineage may undergo extensive genetic turnover while appearing morphologically stable in the fossil record. Conversely, major anatomical shifts may occur in relatively short geological intervals.

The chapter’s discussion of evolutionary rates helps bridge these scales. By estimating how fast traits change in fossil lineages, palaeontologists can ask whether observed macroevolutionary patterns are compatible with known biological processes.

For example, if a fossil lineage shows a gradual change in tooth structure over millions of years, this may fit comfortably with cumulative selection. If a lineage appears suddenly transformed, scientists must ask whether the fossil record is incomplete, whether change occurred in a small, isolated population, or whether the trait evolved unusually rapidly.

The key is not to reduce macroevolution to a single population-genetic formula. Nor is it to treat macroevolution as magical. The challenge is to connect mechanisms and history without flattening either.

Microevolution provides the gears. Macroevolution shows the architecture built over deep time.

Simpson’s project was to bring the gears and the cathedral into the same conversation.

Friday, July 31, 2026

The Mode of Evolution: How Change Happens

If “tempo” asks how fast evolution happens, “mode” asks how it happens.

Mode concerns the mechanisms, patterns, and pathways of evolutionary change. Does change occur through gradual transformation within lineages? Through branching and divergence? Through adaptation to new environments? Through differential survival of populations and species? Through shifts in developmental patterns?

Although Chapter 1 focuses strongly on rates, it sits within the larger purpose of Tempo and Mode in Evolution: to connect palaeontology with evolutionary theory. Fossils show patterns. Genetics and population biology suggest mechanisms. Simpson wanted these worlds to speak to each other.

This was especially important because palaeontologists and geneticists historically studied evolution at different scales. Geneticists often examined variation within populations and short-term changes. Palaeontologists examined large-scale transformations across geological time. One group had mechanisms; the other had history. Simpson’s work helped stitch the two together.

Mode matters because the same rate of change can arise through different processes. A lineage may change rapidly because of strong natural selection. Another may appear to change rapidly because a new species migrated into the fossil record while the older form disappeared. A third may show change because of shifts in developmental timing or ecological opportunity.

Tempo without mode is just a speed reading. Mode asks what engine is running beneath the hood.

For students of evolution, this distinction is powerful. When we see a pattern, we should not immediately assume a process. A trend in fossil size does not automatically prove directional selection. A sudden appearance does not automatically prove sudden evolution. A stable form does not mean no genetic change occurred.

The mode of evolution is the hidden machinery behind the visible fossil pattern.

To understand evolution fully, we need both the clock and the mechanism.

Retraction Without Borders: What the Country Column Reveals About Scientific Correction

A retraction is not only a paper-level event. It is also a map event.

Every country listed in a retraction record is a small coordinate in the geography of scientific correction. But that map is tricky. A country tag does not mean “this country caused the problem.” It usually means at least one author affiliation was linked to that country. A China-United States paper, for example, counts as both China-linked and United States-linked. That makes the country column less like a passport stamp and more like a collaboration fingerprint.

Using the uploaded Retraction Watch CSV, I analyzed 70,589 records with valid publication and notice dates. For the country analysis, I focused mainly on 56,472 non-conference records, because conference-proceedings batches strongly distort timing patterns. I excluded records tagged as conference abstracts/papers or with conference-like journal titles.

When a paper listed multiple countries, I used an exploded-country approach: one China-United States paper counts once for China and once for the United States. That produced 73,172 non-conference country-paper occurrences.

The central result:

Country-specific retraction patterns are real, but they are mostly explained by publication ecosystems: subject mix, journal clusters, publisher pipelines, multinational collaboration patterns, and reason categories. Country is the visible flag; the machinery underneath is journal, publisher, subject, and failure mode.


1. The first map: countries differ strongly in retraction timing

Among countries with at least 500 non-conference country-paper occurrences, the median time from publication to retraction notice varies widely.

Japan has the longest median lag, 4.60 years, followed by France, Russia, Italy, the United States, Canada, the United Kingdom and Germany. Countries such as China, Pakistan, Saudi Arabia, South Korea, Turkey, Ethiopia and India have shorter medians.

Median time to retraction by country

Non-conference records only. Countries shown are among the largest by retraction-record count.

0years2years4years6yearsChinaUnited StatesIndiaRussiaSaudi ArabiaIranUnited KingdomJapanPakistanGermanySouth KoreaEgyptItalyFranceCanada

Calculated from the uploaded Retraction Watch CSV. Multi-country papers are counted once for each listed country.

A Kruskal-Wallis test across countries with at least 500 records confirmed that country-associated lag distributions differ strongly: H = 2493.7, p < 1e-300. That is statistically thunderous.

But significance is not explanation. The country label bundles together subject mix, publishers, journals, collaboration patterns and reason types. The real question is: what kind of retraction ecosystem is each country attached to?


2. Multinational versus single-country papers: the raw story is misleading

At first glance, multinational papers seem slower in the full dataset:

DatasetSingle-country medianMultinational medianTest
All dated records1.29 years1.70 yearsMann-Whitney p = 4.0e-146, Cliff’s delta = 0.153
Non-conference records1.71 years1.77 yearsMann-Whitney p = 0.088, Cliff’s delta = 0.011

Once conference records are removed, the difference almost disappears. In journal-like records, multinational status alone is not a strong raw predictor of retraction delay.

Even more interesting: after adjustment for country presence, notice year, society-linked status, broad subject, broad reason tags, and top journal or publisher buckets, multinational papers were associated with shorter, not longer, retraction lag. In the most adjusted journal-bucket model, multinational status was associated with about 12.9% shorter log-lag. In a logistic model for very late retraction, defined as more than 10 years after publication, multinational papers had lower odds: OR = 0.52, p = 2.8e-24.

That does not mean international collaboration protects papers from long-lag problems. It means multinational records in this database are often concentrated in recent, publisher-detected clusters and fast correction pathways. The country count is not the cause; it is a shadow cast by the publication ecosystem.


3. Multinational share has changed over time

The proportion of multinational records among non-conference retractions has not been stable. It was modest through much of the 2000s and 2010s, dipped during some batch-retraction years, then rose sharply in the most recent years of the uploaded database.

Multinational share of non-conference retraction records over time

Share of non-conference records listing more than one country. The year 2026 is partial.

0%9%18%27%36%20002002200420062008201020122014201620182020202220242026

Calculated from the uploaded Retraction Watch CSV.

The 2023 spike in total retractions was not especially multinational: only 15.9% of non-conference records listed more than one country. But 2024, 2025 and partial 2026 show much higher multinational shares, 26.2%, 31.1% and 33.5%.

That suggests a shift in the correction landscape. Recent corrections include more internationally networked papers, or at least more records with multinational affiliation footprints.


4. The country map has two axes: multinational share and retraction lag

Some countries in the dataset are mostly single-country retraction ecosystems. Others are overwhelmingly multinational.

Saudi Arabia, Pakistan, Malaysia and Ethiopia have very high multinational shares, above 80%. China and Russia have low multinational shares, about 13%. Japan also has a relatively low multinational share but a long median lag. France, Canada, Australia and the United Kingdom have high multinational shares and moderate to long lags.

Multinational share versus median retraction lag by country

Each point is a country with at least 500 non-conference country-paper occurrences.

1years2years3years4years5years0255075100

Calculated from the uploaded Retraction Watch CSV.

This plot punctures a simple assumption: multinational does not automatically mean slow. Pakistan, Saudi Arabia, Malaysia and Ethiopia are highly multinational but have short median lags. Japan is less multinational but much slower. France is both highly multinational and slow.

The explanation is not collaboration size alone. It is which collaboration networks are attached to which journals and reasons.


5. Country-specific reason signatures are sharp

Reason tags were grouped into broad themes: paper mill/peer-review/AI, image concerns, fraud/misconduct, plagiarism/duplication/copyright, and data/results/method concerns. Categories overlap, so percentages do not add to 100.

Country-specific retraction reason signatures

Selected countries. Reason categories overlap, so percentages do not sum to 100.

Paper mill / peer-review / AI
Image concerns
Fraud / misconduct
Plagiarism / duplication
0%20%40%60%80%ChinaUnited StatesIndiaRussiaSaudi ArabiaJapanUnitedKingdomFranceItalyPakistan

Calculated from the uploaded Retraction Watch CSV.

The differences are not cosmetic. Chi-square tests with Benjamini-Hochberg correction showed strong reason enrichment patterns.

Examples:

CountryStrong enrichment signalApprox. odds ratio versus rest
JapanFraud/misconductOR ≈ 8.46
RussiaPlagiarism/duplication/copyrightOR ≈ 7.76
EthiopiaPaper mill/peer-review/AIOR ≈ 6.39
ChinaPaper mill/peer-review/AIOR ≈ 5.72
United StatesFraud/misconductOR ≈ 3.60
GermanyFraud/misconductOR ≈ 3.31
United StatesImage concernsOR ≈ 2.15
ItalyPlagiarism/duplication/copyrightOR ≈ 2.18
PakistanPaper mill/peer-review/AIOR ≈ 2.06
IndiaPaper mill/peer-review/AIOR ≈ 1.62

This is the first real explanatory layer. China, India, Pakistan, Saudi Arabia and Ethiopia have large paper-mill or peer-review-process signatures. Japan and the United States show stronger fraud/misconduct and image/data signals. Russia is dominated by plagiarism/duplication/copyright. Italy has a strong plagiarism plus image profile.

Different countries in the retraction database are not merely faster or slower. They fail through different channels.


6. Country pairs: not all collaborations have the same correction clock

The most common country-pair co-occurrence was China + United States, with 902 records, median lag 2.31 years. But other large pairs, such as India + Saudi Arabia, Pakistan + Saudi Arabia, China + Pakistan and China + Saudi Arabia, have much shorter medians, around 1.5 to 1.6 years.

Largest country-pair clusters in retraction records

Pairs are co-occurrences in multi-country non-conference records. A paper with three countries contributes to three pair counts.

02505007501,000China + United St...India + Saudi ArabiaPakistan + Saudi...China + PakistanEgypt + Saudi ArabiaChina + Saudi ArabiaEthiopia + IndiaChina + South KoreaUnited Kingdom +...China + IndiaGermany + United...China + United Ki...India + United St...Italy + United St...Canada + United S...

Calculated from the uploaded Retraction Watch CSV.

Pair-level Mann-Whitney tests compared each country pair with all other multinational records, with FDR correction. Several pairs had significantly shorter lags:

Faster-than-background pairRecordsMedian lagCliff’s delta
Pakistan + United States1050.93 years-0.375
Saudi Arabia + United Kingdom731.24 years-0.293
Jordan + Saudi Arabia1061.39 years-0.252
Pakistan + United Kingdom831.20 years-0.236
Ethiopia + Saudi Arabia1191.42 years-0.235
China + South Korea3231.37 years-0.188

And several pairs had significantly longer lags:

Slower-than-background pairRecordsMedian lagPattern
France + Saudi Arabia1017.17 yearsVery long-lag biomedical cluster
Japan + United States1975.61 yearsImage/fraud-heavy biomedical profile
Italy + United States2363.84 yearsImage/plagiarism-heavy life-science profile
Spain + United States1233.00 yearsLonger biomedical/data profile
South Korea + United States1162.88 yearsMixed but slower
United Kingdom + United States2932.26 yearsMedicine/biomedicine-heavy, long tail
China + United States9022.31 yearsMixed, more image and biology than China-only clusters

The China-United States pair is especially important because it is the largest pair and does not resemble the rapid paper-mill/peer-review clusters that dominate some other China-linked records. It has a higher image-concern share, 38.6%, and a much higher society-linked share, 15.2%, than China’s overall country profile.


7. Country-pair trends have surged recently

Many large multinational pair clusters are recent. The 2020-2026 era dominates for China-Pakistan, China-Saudi Arabia, India-Saudi Arabia, Pakistan-Saudi Arabia and China-India. China-United States was already present earlier, but it also increased sharply after 2020.

Country-pair retraction clusters by era

Selected country pairs, counted as co-occurrences in non-conference multinational records.

0200400600800≤20092010-20142015-20192020-2026

Calculated from the uploaded Retraction Watch CSV. The 2020-2026 era includes partial 2026.

This is one of the strongest time-specific signals in the dataset.

Older multinational retraction clusters often involve the United States, United Kingdom, Canada, Germany, Italy and Japan. Newer multinational clusters increasingly involve China, India, Pakistan, Saudi Arabia, Ethiopia, Egypt and other countries in large publisher-audit or paper-mill-linked networks.

Again, this is not a national guilt map. It is a map of how publication pipelines globalized.


8. Exact country combinations sharpen the story

Pair co-occurrence is generous: a five-country paper contributes ten country pairs. Exact country sets are stricter.

The largest exact multinational country set is China;United States, with 668 records, median lag 2.50 years, image concerns 42.8%, fraud/misconduct 15.7%, and plagiarism/duplication 37.3%.

That is very different from exact China;South Korea, with 237 records, median lag 1.30 years, paper-mill/peer-review/AI tags 86.9%, and image concerns only 4.2%.

Other exact combinations:

Exact country setRecordsMedian lagMain signature
China;United States6682.50 yearsImage/data/fraud, mixed biomedicine
China;South Korea2371.30 yearsPeer-review/paper-mill-heavy
China;Pakistan2031.68 yearsPeer-review/paper-mill-heavy
Ethiopia;India2001.57 yearsVery high peer-review/paper-mill signal
Egypt;Saudi Arabia1892.51 yearsMixed, image and plagiarism
Italy;United States1326.84 yearsSlow image/plagiarism-heavy profile
Japan;United States1246.02 yearsSlow image/fraud profile
Canada;United States1182.89 yearsBiomedical/data long-tail profile

This exact-combination view shows that the same country can participate in different retraction worlds. China + United States behaves unlike China + Pakistan. United States + Japan behaves unlike United States + Pakistan. Pair identity matters because it captures networks, journals, subjects and institutions better than single-country labels.


9. Subject effects: countries do not retract in the same disciplinary universe

Subject mix is a major confounder.

China-linked retractions are spread across biology/life sciences, medicine and physical sciences/engineering, but with a strong paper-mill/peer-review signal. India-linked records are heavily physical sciences/engineering. Russia-linked records are unusually social-science-heavy and plagiarism-heavy. Japan-linked records are medicine and biomedical-heavy and long-lag. France is strongly biology/life-science-heavy and long-lag.

Country-level subject summaries show the pattern:

CountryStrong subject signalMedian lag interpretation
ChinaPhysical/engineering, biomedical, computing-heavy clustersShorter, many publisher-audit and paper-mill/peer-review records
IndiaPhysical/engineering and computing clustersShorter to moderate, process-heavy
RussiaSocial sciences and humanities-heavyPlagiarism/duplication signature, longer median but low after-10-year share
JapanMedicine and biomedical-heavyLong median and long tail
United StatesBiology/life sciences and medicine-heavyImage, fraud and long-tail clusters
FranceBiology/life sciences-heavyLong median, fewer process-heavy records
Saudi Arabia/Pakistan/EthiopiaHighly multinational, many physical/engineering and publisher-audit clustersShort median, low after-10-year share

This is why country comparisons must be subject-aware. A country whose retracted records come from computational special issues will look different from a country whose retracted records come from decades-old cancer biology, anesthesiology, or molecular medicine.


10. Publisher effects: countries travel through different publishing pipelines

The largest country-publisher clusters are deeply asymmetric. China + Hindawi alone has 9,960 records, median lag 1.28 years, and 98.4% paper-mill/peer-review/AI tags. China + IEEE is largely conference-driven and therefore excluded here, but China still dominates several non-conference publisher clusters.

Largest country-publisher clusters

Non-conference country-paper occurrences. Counts are database records, not rates relative to total output.

03K6K9K12KChina | HindawiChina | SpringerChina | ElsevierChina | WileyUnited States | E...China | Springer...China | IOS Press...China | SpandidosChina | SAGEIndia | SpringerIndia | ElsevierIndia | Springer...India | HindawiChina | Taylor &...China | PLoS

Calculated from the uploaded Retraction Watch CSV.

This plot explains a lot of the country pattern.

China looks fast partly because China-linked records are heavily concentrated in fast publisher-audit clusters: Hindawi, Springer, IOS Press/Sage, SAGE, and several journal families with near-total paper-mill/peer-review tagging.

But China is not uniformly fast. China + Spandidos has a median lag of 5.64 years, with high image and plagiarism/duplication signals. China + PLoS has a median lag of 3.92 years and a mixed image/data profile.

Likewise, the United States does not have one pattern. United States + Elsevier has a median lag of 1.33 years, while United States-linked Journal of Biological Chemistry and PLoS One clusters have much longer medians.

So publisher explains country signal, but not completely. The same country changes character when it moves through a different publisher pipeline.


11. Journal effects: country-journal clusters are the real engines

Country-journal clusters are even more revealing than country-publisher clusters.

The largest clusters are dominated by China-linked papers in specific journals with high paper-mill/peer-review tags:

Country-journal clusterRecordsMedian lagPaper-mill/peer-review/AI
China, Computational and Mathematical Methods in Medicine9971.24 years99.6%
China, Journal of Healthcare Engineering9821.62 years99.7%
China, Journal of Intelligent & Fuzzy Systems9741.90 years94.8%
China, Computational Intelligence and Neuroscience8881.22 years99.9%
China, Security and Communication Networks8571.36 years99.4%
China, Arabian Journal of Geosciences7670.30 years100.0%
India, Journal of Intelligent & Fuzzy Systems4291.56 years99.8%
United Kingdom, Cochrane Database of Systematic Reviews3218.90 years0.3%
China, PLoS One6554.01 yearsMixed image/data profile
India, Soft Computing3033.54 years99.0%

This is a crucial finding:

Countries do not retract. Journal-country pipelines retract.

China in Arabian Journal of Geosciences behaves very differently from China in PLoS One. India in Journal of Intelligent & Fuzzy Systems behaves differently from India in Elsevier biomedical journals. The United Kingdom in Cochrane reviews behaves nothing like the United Kingdom in ordinary research-article clusters.


12. Society versus non-society country patterns

Using a conservative high-confidence society-linked publisher classification, society-linked records are unevenly distributed across countries.

Among large country groups:

CountrySociety-linked share of non-conference records
United States27.7%
Japan23.9%
South Korea21.2%
Russia17.6%
Italy16.4%
France16.3%
Canada15.5%
Spain15.3%
China5.1%
India6.1%
Saudi Arabia3.3%
Pakistan4.0%
Ethiopia1.0%
Malaysia1.7%

This partially explains the long-tail geography. The United States and Japan are more represented in society-linked biomedical and life-science journals, where image/data and misconduct investigations often take longer. China, India, Saudi Arabia and Pakistan are more represented in non-society/unclassified publisher-audit clusters, where paper-mill or peer-review problems can be corrected in batches.

But society status is not enough. In adjusted models, country effects persisted even after controlling for society-linked status, subject, reasons, notice year and top journal/publisher buckets. That means society versus non-society is one ingredient, not the whole recipe.


13. Adjusted models: country signal persists, but shrinks into ecosystem signal

I fitted robust OLS models using log(1 + lag years) as the outcome. These models included country-presence indicators for the largest countries, multinational status, notice year, society-linked status, broad subject flags and broad reason flags. I then added publisher buckets and journal buckets.

The model R² increased as publication ecosystem terms were added:

ModelControls added
BaseCountry + year + society + subject + reasons0.229
Publisher modelBase + top publisher buckets0.258
Journal modelBase + top journal buckets0.302

Adding journal effects explains substantially more variation, confirming that journal pipelines are central.

In the journal-adjusted model, several country effects persisted:

Country presenceApprox. adjusted effect on log-lagInterpretation
Japan+51.4%Much longer lag even after controls
France+31.8%Longer lag
Germany+23.6%Longer lag
Russia+22.5%Longer lag
United Kingdom+16.3%Longer lag, but shrinks strongly after journal control
Italy+13.3%Longer lag
United States+10.4%Longer lag
China-19.0%Shorter lag
India-9.2%Shorter lag
Pakistan-10.6%Shorter lag
Saudi Arabia-6.6%Shorter lag

For very late retractions, defined as more than 10 years, a publisher-adjusted logistic model showed:

Country presenceAdjusted odds ratio for >10-year retraction
United Kingdom2.77
Japan2.53
Germany2.47
Italy1.61
China0.16
Russia0.11
Saudi Arabia0.18
Pakistan0.08
India0.38

The United States and France were not statistically strong in this specific logistic model after controls, even though their raw long-tail shares were high. That suggests their long-tail patterns are more heavily explained by journal/publisher/subject/reason mix.

The statistical interpretation:

Country effects remain after adjustment, but much of the country pattern is really journal-publisher-subject-reason structure wearing a country label.


14. Hypothesis testing summary

HypothesisResultEvidence
Countries differ in time to retractionSupportedKruskal-Wallis H = 2493.7, p < 1e-300
Multinational papers are slower to retractNot supported for journal-like recordsNon-conference Mann-Whitney p = 0.088, Cliff’s delta = 0.011
Multinational status predicts late retraction after adjustmentOpposite directionLogistic OR for >10-year retraction = 0.52, p = 2.8e-24
Country-specific reason profiles differStrongly supportedMultiple FDR-corrected chi-square enrichments
Country pairs have specific retraction clocksSupportedSeveral pair-level Mann-Whitney tests significant after FDR
Journal/publisher effects explain country differencesPartly supportedR² rises from 0.229 to 0.302 when journal buckets are added
Society-linked journals explain long-lag country profilesPartly supportedHigher society-linked share in US/Japan/Europe, but adjusted country effects persist
Subject mix explains country patternsPartly supportedJapan/US/France more biomedical, Russia more social science, China/India/Saudi/Pakistan more process-heavy publisher clusters

15. The exceptions are the most informative part

Exception 1: China is fast overall, but not always fast

China’s median lag is 1.50 years, but China + PLoS has a median lag of 3.92 years, and China + Spandidos has 5.64 years. So “China-linked retractions are fast” is only true in the aggregate because the aggregate is dominated by fast publisher-audit clusters.

Exception 2: Multinational does not mean long-lag

Saudi Arabia, Pakistan, Malaysia and Ethiopia are highly multinational in this dataset, but their medians are short and their after-10-year shares are tiny. Their multinational records are often in recent, publisher-audit, peer-review, or paper-mill-related clusters.

Exception 3: Japan has low multinational share but very long lag

Japan has only about 25% multinational records but the longest country median, 4.60 years, and the highest after-10-year share among large countries, 26.4%. This points toward older biomedical, clinical, institutional and fraud/misconduct-heavy corrections.

Exception 4: Russia is slow but not long-tail

Russia has a median lag of 3.19 years, but only 0.83% after 10 years. Its signature is not late biomedical correction. It is plagiarism/duplication-heavy, often in social-science or humanities-like spaces.

Exception 5: China-United States is not like China-Pakistan

China + United States has a median lag of 2.31 years, image concerns 38.6%, fraud/misconduct 14.5%, and society-linked share 15.2%. China + Pakistan has median lag 1.58 years, paper-mill/peer-review tags 65.2%, and no after-10-year records in this dataset. Same China label, different retraction ecosystem.


16. The final map: countries are not causes, they are coordinates

The country column is tempting. It invites rankings. It whispers: faster, slower, better, worse. But the database resists that crude reading.

Country-specific differences are real. Japan, France, Russia, Italy and the United States have longer median lags. China, Pakistan, Saudi Arabia, Ethiopia, South Korea and India have shorter medians. Some countries are image-heavy, some plagiarism-heavy, some paper-mill/peer-review-heavy, some long-tail biomedical.

But the deeper conclusion is not national character. It is publication ecology.

Countries appear in different parts of the publishing machine:

  • China, India, Pakistan, Saudi Arabia, Ethiopia: large recent publisher-audit and paper-mill/peer-review clusters.
  • Japan, United States, Germany, Italy, France, United Kingdom: more biomedical, society-linked, image/data, misconduct and long-tail correction clusters.
  • Russia: a strong social-science and plagiarism/duplication signature.
  • China-United States, Japan-United States, Italy-United States: slower, more biomedical and image-heavy collaboration clusters.
  • China-Pakistan, Pakistan-Saudi Arabia, India-Saudi Arabia, Ethiopia-India: newer, faster, more process-heavy multinational clusters.

The country column is therefore not the verdict. It is the first clue.

The real retraction geography is made of journals, publishers, subjects, collaborations, editorial systems, institutional investigations, paper-mill audits, image forensics and time. It is not a flat political map. It is a weather map with moving storms.

Some storms are old and forensic.
Some are recent and industrial.
Some gather around journals.
Some gather around publishers.
Some cross borders so quickly that the country label becomes less useful than the network itself.

Science corrects itself, but the correction travels through pipes. The country column tells us where the pipes surface. 🔬📍