The human genome gave biology an extraordinary catalogue of approximately 20,000 protein-coding genes. Yet a gene sequence is not the same thing as proof that its corresponding protein actually exists. A gene may be transcribed only in a handful of cells, activated briefly during development or disease, translated into a tiny peptide, embedded deep inside a membrane, or produced at concentrations far below the reach of routine experiments.
This gap between what the genome predicts and what experiments can confidently detect lies at the heart of the article “Accelerating the search for the missing proteins in the human proteome,” published in Nature Communications in 2017.
The article reviews the progress of the Human Proteome Project, explains why many predicted proteins were still considered “missing,” examines difficult examples such as olfactory receptors, prestin and interleukin-9, and proposes a community knowledge platform called MissingProteinPedia. Its central argument is both practical and philosophical:
A protein may be missing from a high-stringency proteomics database without being absent from human biology.
The challenge, therefore, is not merely to produce more mass spectra. It is to connect molecular clues scattered across proteomics, genetics, transcriptomics, physiology, pharmacology, microscopy and decades of published or unpublished research.
From the human genome to the human proteome
Sequencing the human genome revealed the instructions encoded in DNA. The Human Proteome Project, or HPP, was created to determine how those instructions are translated into the functioning molecular machinery of human life.
Launched by the Human Proteome Organization in 2010, the HPP had two broad objectives.
First, it sought to complete a reliable protein parts list for the human body. Ideally, this would include at least one experimentally supported protein product from every protein-coding gene, along with splice variants, post-translational modifications and amino-acid variations.
Second, it aimed to make proteomics a mature partner to genomics in biomedical and clinical research. That required more than detecting proteins. It required standardized repositories, reproducible analytical pipelines, accepted confidence thresholds and community-wide rules for deciding when a protein had truly been observed.
The HPP was organized into two overlapping branches:
-
The Chromosome-centric Human Proteome Project, or C-HPP, divided the search according to chromosomes.
-
The Biology and Disease Human Proteome Project, or B/D-HPP, investigated proteins through tissues, pathways, biological systems and diseases.
These efforts were supported by three major pillars:
-
Mass spectrometry
-
Affinity reagents, particularly validated antibodies
-
Knowledgebases that integrate and curate evidence
Databases and resources such as neXtProt, PeptideAtlas, ProteomeXchange, the Human Protein Atlas and GPMDB became essential parts of this ecosystem. Together, they created the scaffolding needed to turn enormous quantities of experimental data into defensible protein identifications.
What exactly is a “missing protein”?
The word missing can be misleading. It does not necessarily mean that the protein does not exist. Instead, it means that the available evidence did not satisfy the HPP’s accepted criteria for confirmation at the protein level.
The article describes the neXtProt protein-existence system, which placed human proteins into five categories.
PE1: Evidence at the protein level
These proteins had strong experimental evidence, derived from approaches such as mass spectrometry, validated antibodies, Edman sequencing or structural determination.
Under the stringent mass-spectrometry criteria adopted in 2016, a typical PE1 assignment required at least two highly confident, uniquely mapping peptides of nine or more amino acids that were not nested within one another.
PE2: Evidence at the transcript level
The corresponding RNA had been detected, but sufficiently strong evidence for the protein itself was lacking.
PE3: Evidence from homologous species
A related protein had been found in another organism, suggesting that the human protein probably exists, but direct human transcript or protein evidence was insufficient.
PE4: Predicted evidence
The protein was predicted from genomic information, but experimental support remained very limited.
PE5: Dubious or questionable proteins
These entries were considered unlikely to produce functional proteins, often because of disrupted genes, pseudogene-like characteristics or missing transcriptional features.
The article defines PE2, PE3 and PE4 proteins collectively as the missing proteins. PE5 proteins were excluded because they were considered dubious rather than simply undetected.
Using the February 2016 neXtProt release, the authors reported:
| Protein-existence category | Number of proteins |
|---|
| PE1 | 16,518 |
| PE2 | 2,290 |
| PE3 | 565 |
| PE4 | 94 |
| PE5 | 588 |
| Total | 20,055 |
Thus, 2,949 proteins were classified as PE2–PE4 and therefore considered missing under the HPP framework.
The comparison shown in Box 1 on page 2 also demonstrates genuine progress. Between 2013 and 2016, PE1 assignments increased from 15,649 to 16,518. This occurred even while the criteria became stricter and hundreds of previously accepted PE1 proteins were downgraded.
Why high-stringency evidence matters
Proteomic experiments produce immense volumes of data. Mass spectrometry breaks proteins into peptides, measures their mass-to-charge characteristics and compares the resulting fragmentation spectra against candidate sequences.
The difficulty is that a plausible peptide-spectrum match is not automatically correct.
False assignments can arise from:
-
Noisy or incomplete spectra
-
Very short peptide sequences
-
Closely related proteins with shared peptides
-
Sequence variants
-
Leucine/isoleucine ambiguity
-
Incorrect database assumptions
-
Multiple testing across huge search spaces
-
Inadequate false-discovery-rate control
For this reason, the HPP moved toward increasingly rigorous standards. The article discusses recommendations such as a 1% protein-level false discovery rate, accompanied by peptide-level and peptide-spectrum-match-level error estimates. Raw spectra should also be deposited publicly so that claims can be independently re-examined.
The requirement for two unique peptides of at least nine amino acids greatly lowers the probability of a random or ambiguous assignment. Yet the authors emphasize an important caveat: even two qualifying peptides do not make an identification absolutely unquestionable. Confidence improves dramatically, but scientific evidence remains open to reassessment.
This tension became visible after two large “draft human proteome” studies reported many previously undetected proteins using criteria that differed from HPP standards. Some of their assignments, particularly those involving olfactory receptors, were later criticized for marginal spectra, short peptides, ambiguous matching and insufficient false-discovery control.
The resulting dispute was productive. It exposed weaknesses in large heterogeneous datasets and pushed the field toward clearer metrics, transparent workflows and stronger deposition standards.
Why can a real protein remain invisible?
The article’s most important contribution is its explanation of why certain proteins repeatedly escape standard detection. The proteome is not a static warehouse. It is closer to a city at night: some buildings glow continuously, while others illuminate one room for a few seconds under very specific conditions.
1. Extremely low abundance
Some proteins are present at only a few copies per cell. Their peptide signals may be drowned out by abundant structural proteins, serum proteins or housekeeping proteins.
2. Highly restricted tissue expression
A protein may occur only in the inner ear, a small brain nucleus, a particular epithelial layer or a few sensory neurons. Whole-tissue analysis can dilute its signal almost beyond recovery.
3. Narrow temporal expression
Some proteins are produced only:
-
During a particular developmental stage
-
After immune activation
-
Under environmental stress
-
During infection
-
In a specific disease state
-
At a particular point in a cell’s differentiation
A sample taken at the wrong time may contain no detectable protein.
4. Difficult subcellular localization
Proteins confined to cilia, axon terminals, extracellular vesicles, membrane microdomains or other specialized structures may require targeted enrichment.
5. Membrane association and hydrophobicity
Transmembrane proteins are difficult to dissolve, purify and digest. Their trypsin cleavage sites may be buried in the lipid bilayer or shielded by neighbouring molecules.
6. Poor compatibility with trypsin digestion
Bottom-up proteomics commonly uses trypsin to cut proteins into peptides. Some proteins simply do not generate two suitable, uniquely mapping tryptic peptides of the required length.
7. Small proteins and bioactive peptides
Short secreted peptides may be biologically crucial but too small to satisfy rules designed for conventional proteins. The article cites the orexigenic neuropeptide QRFP as an example. A peptide can control appetite or cell signalling and still remain excluded from PE1 because it cannot generate two independent peptides of nine amino acids.
8. Extensive post-translational modification
Glycosylation, phosphorylation, cleavage and other modifications can alter peptide masses and complicate database matching.
9. Insolubility or cross-linking
Proteins that form dense extracellular structures, complexes or cross-linked assemblies may resist conventional extraction.
10. Genuine absence under ordinary conditions
Some predicted gene products may not be translated in normal physiology. Others may be expressed only under rare circumstances, or perhaps not at all.
The authors therefore recommend targeted approaches such as subcellular enrichment, better fractionation, analysis of rare tissues, improved membrane-protein workflows, alternative proteases, higher-sensitivity instruments, validated antibodies and examination of samples under diverse physiological and pathological conditions.
Which protein families contained the largest gaps?
The authors performed bioinformatics analyses of missing proteins according to their families, domains, biological classes, pathways and evolutionary relationships.
The chart on page 5 identifies several especially prominent categories:
-
Olfactory receptors
-
Zinc-finger proteins
-
Other transmembrane proteins
-
Coiled-coil-domain proteins
-
Homeobox proteins
-
Keratin-associated proteins
-
Sperm-related and testis-expressed proteins
-
Solute carriers
-
β-defensins
-
PRAME-family proteins
-
Other G-protein-coupled receptors
The analysis showed encouraging progress for most major groups between 2013 and 2016. The glaring exception was the olfactory-receptor family, whose representation among missing proteins actually increased proportionally.
The comparison of UniProt families on page 5 also revealed that some families were dominated by proteins lacking high-stringency evidence. GPCR type 1 proteins were particularly enriched among PE2–PE4 entries. Apart from a few families such as Kruppel C2H2 zinc-finger proteins and peptidase C19 proteins, between 50% and 95% of the members of many top-ranked missing-protein families remained unconfirmed.
This pattern suggests that missingness is not random. Certain biological architectures create systematic blind spots. Large membrane-receptor families, recently expanded gene families and proteins with highly specialized expression are naturally harder for conventional proteomics to capture.
The olfactory receptor problem
Olfactory receptors became the article’s central case study because they represented the largest family of missing proteins.
These proteins belong to the G-protein-coupled receptor superfamily. GPCRs respond to an astonishing range of signals, including photons, neurotransmitters, hormones, nutrients, metals and volatile chemicals. They are also among the most important classes of pharmaceutical targets.
The article describes five major GPCR branches:
-
Rhodopsin or class A
-
Secretin
-
Adhesion or class B
-
Glutamate or class C
-
Frizzled and taste receptor 2
The phylogenetic analysis on page 6 shows missing GPCRs distributed throughout these branches, but the greatest concentration lies within the rhodopsin branch containing olfactory receptors.
At the time of analysis, the genome contained approximately 480 olfactory-receptor genes. Twelve were classed as PE5. The remainder encoded 411 unique proteins, of which only two were listed as PE1 and 409 remained PE2–PE4.
In other words, almost the entire human olfactory-receptor repertoire was missing by HPP standards.
Functional evidence versus proteomic proof
The scarcity of mass-spectrometry evidence did not imply biological inactivity.
Functional screening had identified agonists for numerous olfactory receptors. One study tested receptors against 73 candidate ligands and found agonists for 27 receptors, including 18 that had previously been orphans.
Such experiments show that a receptor can:
-
Be expressed in a heterologous system
-
Reach the cell membrane
-
Bind or respond to an odorant
-
Activate downstream signalling
That is compelling biological evidence. Yet it may not satisfy a formal PE1 mass-spectrometry criterion.
The article highlights an inconsistency in the 2016 classifications. OR1D2 and OR2AG1 were listed as PE1, but the reported supporting evidence did not appear to meet the contemporary HPP mass-spectrometry thresholds. OR1D2 lacked MS or antibody evidence, while OR2AG1 was associated with only a single seven-amino-acid peptide. Meanwhile, another receptor with comparable functional evidence remained PE4.
The lesson is not that functional studies are unreliable. It is that the rules for incorporating non-MS evidence had not been standardized as clearly as the rules for mass spectrometry.
Reanalysing more than 122,000 peptide-spectrum entries
To investigate olfactory receptors more systematically, the authors searched public proteomic repositories including GPMDB, PRIDE, ProteomicsDB, MAXQB and Human ProteinPedia.
They aggregated 122,717 peptide-spectrum entries of at least seven amino acids.
The filtering process was severe:
-
Removal of non-unique and decoy peptides left 4,751 potentially proteotypic olfactory-receptor peptides.
-
Only 286 carried a high-confidence score from search engines such as SEQUEST, Mascot or MaxQuant.
-
Manual examination of spectral quality reduced the set to 64 strong spectra representing 24 peptides.
-
After merging overlapping peptides, only 23 unique olfactory-receptor peptides remained.
These data provided some mass-spectrometry evidence for 23 of the 409 missing olfactory receptors, approximately 5.6%.
However, none had the two qualifying peptides needed to satisfy the strict PE1 standard. Fourteen proteins were supported by only one seven- or eight-amino-acid peptide, while nine had one peptide longer than nine amino acids.
The authors describe these receptors as proteins “waiting in the wings.” The available evidence pointed researchers toward promising targets, but additional confirmation was required.
The analysis is an excellent demonstration of the difference between a clue and a completed identification. A single credible spectrum may not close the case, but it can reveal where to search next.
Olfactory receptors are not confined to the nose
Another reason olfactory receptors should not be dismissed is their expression outside nasal tissue.
The article notes evidence for olfactory-receptor expression in multiple epithelial tissues, where they may perform broader chemosensory roles. Thus, searching only the olfactory epithelium may be unnecessarily restrictive.
A receptor could be present:
-
In very few sensory neurons
-
On cilia that are difficult to isolate
-
In non-nasal epithelial tissues
-
At low abundance
-
Under specific hormonal, metabolic or disease conditions
The appropriate search strategy is therefore not simply “analyse more nose tissue.” It requires combining functional biology, tissue-expression information and targeted proteomic design.
Chromosome 7 as a miniature map of the problem
The Australian and New Zealand C-HPP teams focused on chromosome 7. The chromosome map on page 8 plots 757 PE1 proteins and 139 PE2–PE4 proteins along its length.
The missing proteins were not restricted to gene-poor regions. In fact:
-
56% came from regions described as having high gene density
-
12% came from moderately dense regions
-
25% came from low-to-moderate-density regions
-
Only 1.5% came from low-density regions
This observation argues against a simplistic explanation that missing proteins merely originate in poorly populated or poorly annotated genomic regions. Their invisibility is more likely related to expression level, tissue specificity, protein chemistry or sampling conditions.
Among the chromosome 7 missing proteins, 27 were GPCRs. These included:
-
15 olfactory receptors
-
6 taste-related receptors
-
4 orphan GPCRs
-
The serotonin receptor 5-HT5A
-
Metabotropic glutamate receptor 8
The article then uses selected receptors to show how rich biological evidence can coexist with absent high-stringency protein detection.
5-HT5A: functional, but elusive
The HTR5A gene encodes the serotonin receptor 5-HT5A.
Evidence discussed in the article includes:
-
Human brain mRNA detected by in situ hybridization and PCR
-
G-protein activation and inhibition of adenylyl cyclase in expression systems
-
Altered behaviour in knockout mice
-
Changed responses to the serotonin-related compound LSD in knockout animals
Yet the authors found no convincing human protein localization by immunohistochemistry or western blot.
The likely explanation is not necessarily non-existence. The receptor may be expressed at extremely low levels in narrowly restricted brain regions, making it difficult to detect in bulk tissue.
mGlu8: low expression and complex biology
GRM8 encodes metabotropic glutamate receptor 8.
The article reports:
-
Functional signalling in heterologous expression systems
-
Low and anatomically restricted mRNA in human brain
-
Expression reported in cancer cells, hippocampal cells and astrocytes
-
Associations with epilepsy and multiple sclerosis tissue
-
Physiological consequences after gene deletion in mice
Its large size, complex gene structure and possible alternative splicing may create multiple protein forms and further complicate detection.
Again, the biology looks active, while the conventional proteomic signal remains faint.
GPR22: a protein hiding in conditional biology
GPR22 provides an even more uncertain case.
Its mRNA had been detected in human heart and brain, but no ligand had been identified. In experimental systems, its unusual AT-rich coding sequence appeared to interfere with efficient expression. Signalling could be restored after modifying the sequence composition.
Knockout mice did not display an obvious baseline phenotype. However, under cardiac stress, animals lacking GPR22 developed heart failure more rapidly, suggesting that its role may emerge only under particular physiological challenges.
This illustrates a deeper point: some proteins may look unimportant under routine laboratory conditions because their function is conditional. Their biological significance may appear only during injury, infection, stress or disease.
Prestin: a well-known protein that remained technically “missing”
Prestin, encoded by SLC26A5, is one of the article’s most striking examples.
It is widely described as the motor protein of cochlear outer hair cells and is central to the mechanics of hearing. The article found abundant indirect and functional evidence:
-
More than 90 peer-reviewed publications
-
Numerous commercially available antibodies
-
Known chloride and bicarbonate relationships
-
Human variants associated with deafness and other phenotypes
-
Multiple transcripts
-
Copy-number variants in clinical databases
-
Experimental genetic tools in model organisms
Yet prestin remained PE2 because high-stringency endogenous MS or accepted antibody evidence was unavailable.
Why?
Prestin occupies a near-perfect hiding place:
-
It is expressed in cochlear outer hair cells.
-
These cells are rare.
-
Human inner-ear tissue is extremely difficult to obtain.
-
Only a few hundred outer hair cells may be collected by specialized microdissection.
-
Prestin is a hydrophobic membrane protein.
-
Membrane localization complicates extraction and tryptic digestion.
-
Its concentration is far below what routine proteomic workflows prefer.
Synthetic peptides corresponding to prestin could be produced and measured, but synthetic standards do not prove that endogenous prestin peptides have been recovered from human tissue.
Prestin reveals a classification paradox. Biologists may regard the protein as firmly established, yet the HPP’s specific evidence machinery may still classify it as missing. This is not necessarily a flaw in quality control. It shows that biological knowledge and assay-specific proof answer related but different questions.
Interleukin-9: when experimental design determines visibility
Interleukin-9 offers a different kind of challenge.
Small secreted signalling proteins are often:
-
Produced transiently
-
Released only after stimulation
-
Present at low concentrations
-
Heavily modified
-
Surrounded by vastly more abundant extracellular proteins
-
Too short to yield multiple qualifying tryptic peptides
The authors examined the secretome of activated primary T cells. Conventional secretome studies often use serum-free media to avoid contamination, but serum deprivation stresses cells and can produce a flood of stress- and apoptosis-related proteins.
The authors instead cultured cells for several days in the presence of fetal bovine serum. This created another problem: approximately 95% of detected peptides originated from bovine serum proteins.
After excluding bovine proteins and proteins released by resting human T cells, they identified secretory proteins associated specifically with activated cells, including IL-9.
Their MS analysis detected two IL-9 peptides:
-
YPLIFSR, seven amino acids
-
SLLEIFQK, eight amino acids
The fragmentation spectra are shown in Figure 6 on page 10.
Both peptides were predicted to be unique to IL-9, but neither reached the nine-amino-acid threshold. Thus, the experiment provided biologically meaningful and apparently specific evidence while still falling short of formal PE1 requirements.
IL-9 demonstrates that a fixed peptide-length rule can disadvantage small proteins whose sequence simply cannot produce the required set of tryptic peptides.
MissingProteinPedia: a home for clues that do not fit the final verdict
The article’s proposed solution is MissingProteinPedia, a communal database designed to complement rather than replace the high-stringency HPP system.
The distinction is crucial.
The HPP’s official pipelines answer:
Does this protein satisfy the agreed criteria for high-confidence identification?
MissingProteinPedia would answer:
What does the scientific community know about this protein, and what clues might help us find it conclusively?
The architecture illustrated in Box 2 on page 3 connects information from sources such as:
-
neXtProt
-
Human Protein Atlas
-
PeptideAtlas
-
ProteomeXchange
-
PRIDE
-
PASSEL
-
MassIVE
-
GPMDB
-
ProteomicsDB
-
MaxQB
-
PubMed
-
UniProt
-
GeneCards
-
GeneRIFs
-
ProtAnnotator
-
Individual laboratories
It was also intended to accommodate preliminary, unpublished, proprietary or legacy observations through protected collaboration interfaces.
Text-mining tools could gather literature associated with genes, proteins and synonyms. Users could add annotations, while administrators could curate material before public release.
The platform was envisioned as searchable and sortable by criteria such as chromosome, tissue and keyword.
Low-stringency does not mean low value
Calling MissingProteinPedia “low-stringency” could sound as though it were designed to collect unreliable information. That is not the authors’ intention.
Rather, the platform separates evidence gathering from final adjudication.
A single peptide spectrum may be insufficient for PE1, but it can identify:
-
A promising tissue
-
A likely stimulation condition
-
A useful peptide target
-
A possible splice isoform
-
A candidate antibody
-
A suitable subcellular fraction
-
A disease state in which the protein is enriched
Likewise, an old laboratory notebook, unpublished western blot or commercial antibody result might not independently prove a protein’s existence. Yet, when combined with transcript data, animal phenotypes and targeted MS evidence, it may reveal the experimental route required for confirmation.
MissingProteinPedia was therefore conceived as a hypothesis engine. It would collect the breadcrumbs without pretending that every breadcrumb is the loaf.
The authors explicitly state that the platform would not initially judge submitted evidence by the same standards used for official HPP reclassification. Instead, it would expose possibilities that might otherwise remain hidden in unpublished experiments, commercial records or disciplinary silos.
The paper’s proposed roadmap
The article presents five major recommendations for accelerating completion of the human proteome.
1. Maintain rigorous HPP standards
Researchers and journals should follow current HPP data-submission guidelines and high-stringency reanalysis metrics.
Lowering standards would make the proteome appear complete more quickly, but it would fill the catalogue with questionable assignments.
2. Consolidate mass-spectrometry evidence
All relevant MS data should be deposited in accessible systems such as ProteomeXchange, including datasets from repositories not fully integrated into the HPP pipeline.
Claims involving missing proteins should be accompanied by transparent raw data.
3. Develop formal rules for non-MS evidence
The community should agree on how to evaluate evidence from:
-
Antibodies
-
Functional assays
-
Structural biology
-
Protein interactions
-
Imaging
-
Genetics
-
Pharmacology
-
Cell biology
-
Other experimental methods
Without common rules, equivalent evidence may produce inconsistent classifications.
4. Hold annual evidence-review jamborees
The authors propose community meetings resembling the annotation jamborees used during the Human Genome Project.
At these events, experts could review proposed upgrades and downgrades, examine disputed evidence and document the rationale behind classification decisions.
5. Capture all relevant knowledge in MissingProteinPedia
Every credible clue concerning PE2–PE4 proteins should be assembled in a shared resource, creating a bridge from exploratory evidence to targeted high-stringency validation.
What the figures collectively reveal
The article’s visual evidence tells a coherent story.
Box 1, page 2
The PE classification diagram shows both progress and increasing strictness. More proteins reached PE1 between 2013 and 2016, even as the minimum peptide requirements became harder to satisfy.
Box 2, page 3
The MissingProteinPedia workflow places the proposed resource between official HPP databases, public literature, independent repositories and laboratory-generated evidence. It is designed as a connective layer rather than a competing authority.
Figure 1, page 4
The extrapolation plots estimate how quickly different databases were reducing the proportion of missing proteins. The projections varied substantially, suggesting that completion depended heavily on the evidence source and classification pipeline. The authors noted that some trajectories implied much slower completion than optimistic HPP timelines.
Figures 2 and 3, page 5
These show that missing proteins are enriched in particular families rather than uniformly distributed across the proteome. Olfactory receptors form the largest and most stubborn group.
Figure 4, page 6
The GPCR phylogenetic trees reveal dense clusters of missing proteins, especially among olfactory receptors. The figure also overlays functional ligands and partial peptide evidence, visually demonstrating that “missing” receptors may already carry several kinds of incomplete evidence.
Figure 5, page 8
The chromosome 7 map shows that missing proteins are spread along the chromosome and are not predominantly confined to gene-poor regions.
Figure 6, page 10
The IL-9 spectra illustrate the threshold problem directly: two apparently proteotypic peptides are detected, but both are shorter than the required nine residues.
The article’s deeper scientific message
At first glance, the search for missing proteins appears to be a technical problem in analytical chemistry. The article shows that it is actually a problem of measurement, ontology and scientific governance.
Measurement
Some proteins fall outside the practical detection range of standard workflows because of abundance, chemistry, localization or timing.
Ontology
Scientists must decide what “protein existence” means. Does functional activity count? Does a validated antibody count? Is one unique spectrum sufficient? What about a resolved structure, a disease-causing mutation or a knockout phenotype?
Different fields naturally answer these questions differently.
Governance
A global project requires rules that are accepted, transparent and consistently applied. It must also preserve the ability to revise earlier conclusions.
The HPP’s strict criteria protect the reliability of the human protein catalogue. However, strict criteria can become a narrow funnel through which certain legitimate biological objects cannot easily pass. The article does not advocate abandoning the funnel. It advocates building a richer map around it.
Strengths of the article
Several features make the paper especially valuable.
It avoids equating non-detection with non-existence
This is perhaps its most important intellectual contribution.
It supports its argument with diverse examples
Olfactory receptors, prestin, GPCRs and IL-9 fail detection for different reasons. Together, they show that missingness has multiple causes.
It defends rigorous standards
The authors do not suggest promoting proteins to PE1 merely because circumstantial evidence exists.
It recognizes disciplinary fragmentation
Important evidence may sit in genetics, pharmacology or physiology databases that proteomics pipelines do not routinely inspect.
It proposes an actionable infrastructure
MissingProteinPedia is presented not simply as a concept, but as an integrated platform with literature mining, repository links, user annotation and protected collaboration.
Important limitations and cautions
Because the review was published in January 2017, its numerical summaries describe the state of neXtProt and the HPP primarily in 2016. The figures should therefore be read historically, not as present-day counts.
Its completion projections were based on short-term linear extrapolation. Scientific progress rarely proceeds linearly. Instrument improvements, new sample types, changes in classification rules and database reanalysis can all produce sudden jumps or reversals.
The proposed low-stringency collection model also creates challenges:
-
Weak evidence can accumulate rapidly.
-
Duplicate claims may appear convincing through repetition.
-
Commercial antibodies may lack adequate validation.
-
Unpublished observations can be difficult to reproduce.
-
Contradictory evidence requires visible provenance.
-
Community curation can become a substantial workload.
A successful system therefore needs to distinguish clearly between:
-
Submitted observations
-
Curated evidence
-
Independently replicated findings
-
Official PE classification
The article itself respects this distinction, but any implementation must preserve it carefully.
Final perspective: finding the protein means finding its context
The search for missing proteins cannot be completed by repeatedly analysing the same convenient tissues with the same digestion enzymes and the same database filters.
Researchers must ask more biological questions:
-
In which cells is the protein expressed?
-
At what developmental stage?
-
Under which disease or stress condition?
-
In which organelle or membrane compartment?
-
Which protease will generate observable peptides?
-
Is the active product a short, processed peptide?
-
Does an isoform alter the expected sequence?
-
Can functional or genetic data guide targeted MS?
-
Is the protein present only after a stimulus?
The article’s enduring idea is that context is itself an experimental reagent.
A low-abundance receptor in a rare sensory cell, a cytokine secreted only after activation and a membrane motor confined to the inner ear are not missing in the same way. Each requires a different scientific key.
The Human Proteome Project had already built a powerful high-stringency engine for determining when evidence was strong enough. MissingProteinPedia was proposed as the complementary scouting network, collecting clues from the wider scientific landscape and directing that engine toward the places where hidden proteins were most likely to emerge.
Completing the human proteome, in this vision, is not merely an exercise in filling empty database cells. It is an effort to understand where, when, how and why every protein participates in human biology. That turns the proteome from a parts list into something far richer: a dynamic molecular atlas of what it means to be human.