Tuesday, September 8, 2026

The Counterfeit Information Factory: How Disinformation Fakes the Messenger, the Meaning and Reality

People often imagine disinformation as a lie wearing a cheap moustache. In practice, the most effective versions are far better dressed.

The photograph may be genuine. The statistic may be accurate. The document may be authentic. Yet the photograph is from another war, the statistic has lost its denominator, and the document has been released at precisely the moment when it can cause maximum damage.

Disinformation works by counterfeiting one or more of three things:

  1. The messenger: Who supposedly created or shared the information?
  2. The meaning: How should genuine facts, images or statistics be interpreted?
  3. Reality itself: Did the reported event, statement or evidence ever exist?

Researchers Claire Wardle and Hossein Derakhshan identify seven major forms of problematic information: satire or parody, false connection, misleading content, false context, impostor content, manipulated content and fabricated content. These categories often overlap, and the same campaign may use several at once.

Read the Information Disorder report

Before Entering the Factory: Three Important Definitions

  • Misinformation is false information shared without an intention to cause harm.
  • Disinformation is false information deliberately created or distributed to deceive or harm.
  • Malinformation is genuine information used strategically to cause harm, such as private information released selectively or at a damaging moment.

The difference is often not the post itself, but the intention and circumstances surrounding it.


Part I: Faking the Messenger

1. Impostor Content: Borrowing Somebody Else’s Credibility

Instead of inventing a new source, a disinformation operator can impersonate one that people already trust. A copied logo, familiar page layout and almost-correct web address can turn an anonymous claim into something that appears to come from a newspaper, government department, university or international organization.

A striking example is the Doppelganger operation documented by EU DisinfoLab. Investigators found websites imitating at least 17 established media organizations, including Bild, 20 Minutes, The Guardian and ANSA. Operators bought domains resembling those of legitimate outlets, copied their visual designs and published fake articles, videos and polls.

Read EU DisinfoLab’s investigation

In 2024, the United States Department of Justice announced the seizure of 32 domains allegedly used in a related Russian government-directed campaign. According to the department, the operation used copied news sites, advertisements, artificial intelligence tools and social media profiles posing as ordinary citizens to direct readers toward counterfeit pages.

Read the US Department of Justice announcement

Why this strategy is effective: Impostor content does not need to build trust from scratch. It steals the accumulated reputation of a familiar institution. Readers may recognize the logo and stop checking.

Why it is difficult to counter: Fact-checking the article’s claims is not enough. Investigators must also inspect domain names, publication histories, account creation dates and technical infrastructure. Counterfeit sites can disappear and reappear under new addresses, making them slippery digital eels.

2. False Personas and Astroturfing: Manufacturing a Crowd

Sometimes the message is less important than the apparent identity of the people spreading it. Operators create fake local activists, concerned parents, veterans, journalists or community organizations. A centrally managed campaign then appears to be an organic public movement.

The 2018 United States indictment of individuals associated with the Internet Research Agency alleged that Russian operators created hundreds of social media accounts posing as Americans. They purchased advertisements, contacted real activists and organized political rallies. The indictment described the use of opposing personas and campaigns, allowing the same operation to inflame multiple sides of a dispute.

The Department of Justice explicitly noted that the charges did not establish that the conduct changed the election result.

Read the Department of Justice indictment summary

Why this strategy is effective: Astroturfing creates the illusion of consensus. People are influenced not only by evidence but also by what appears to be the opinion of their community. A hundred coordinated accounts can make a marginal claim look like a public uprising.

Why it is difficult to counter: An individual post may contain no obviously false statement. The deception becomes visible only when analysts examine synchronized posting, repeated wording, shared infrastructure, unusual account networks and concealed organizational control. Moderators must also distinguish manipulation from genuine collective action.

3. Synthetic Impersonation: Making Someone Appear to Speak

Deepfakes and cloned voices extend impostor content into audiovisual territory. The attacker does not merely imitate a news organization. The attacker manufactures the apparent presence of a real person.

In March 2022, a fabricated video appeared to show Ukrainian President Volodymyr Zelenskyy telling Ukrainians to put down their weapons. The video was technically poor and quickly exposed, but it demonstrated how synthetic media could be inserted into a military crisis, when audiences have little time to verify what they see.

Read Reuters’ report on the Zelenskyy deepfake

Why this strategy is effective: Video and audio create a feeling of direct witnessing. A written claim says, “This person surrendered.” A deepfake appears to say, “Watch the surrender happening.”

Why it is difficult to counter: Detection is a moving target. Compression, reposting and screen recording can erase forensic clues. Even an unsuccessful deepfake can create general uncertainty, allowing people to dismiss authentic recordings as artificial. Verification increasingly requires trusted publication channels, provenance records and independent corroboration.


Part II: Faking the Meaning

4. False Connection: When the Headline, Image and Story Do Not Agree

False connection occurs when a headline, photograph, caption or thumbnail creates an impression that the underlying material does not support.

A revealing experiment came from the satirical website Science Post, which published the headline, “Study: 70% of Facebook users only read the headline of science stories before commenting.” Beneath the opening paragraph was largely meaningless placeholder text. The headline nevertheless circulated widely, illustrating how a claim can travel independently of the content supposedly supporting it.

Read The Washington Post’s discussion of headline-only sharing

In deliberate disinformation, the same structure can be used without the joke. A dramatic photograph may accompany an unrelated article, a cautious scientific paper may receive a sensational headline, or a video thumbnail may depict an event absent from the footage.

Why this strategy is effective: The headline reaches people who never open the article. It can plant an association between a person and an accusation even when the article quietly concedes that the evidence is weak.

Why it is difficult to counter: Platforms often distribute headlines, thumbnails and fragments rather than complete articles. A correction buried in paragraph twelve cannot repair an impression formed in half a second while scrolling.

5. Misleading Framing: Using Facts to Manufacture a False Conclusion

Misleading content frequently contains real numbers, genuine quotations or accurate observations. The deception lies in what is omitted, compared or implied.

During the COVID-19 pandemic, posts argued that vaccines were ineffective because a large number, or sometimes a majority, of recorded COVID-19 deaths occurred among vaccinated people. The raw count could be accurate in highly vaccinated populations, but it ignored the much larger size and older age profile of the vaccinated group. When mortality rates were compared properly, the risk of severe illness and death was lower among vaccinated adults.

Read Reuters’ explanation of the statistical issue

The numbers were not necessarily invented. Their meaning was bent.

Common framing techniques include:

  • Reporting totals while hiding rates or denominators
  • Selecting a convenient starting date
  • Comparing unlike populations
  • Quoting a person accurately but removing the surrounding explanation
  • Presenting correlation as proof of causation
  • Highlighting one study while ignoring the larger body of evidence

Why this strategy is effective: Misleading framing offers plausible deniability. The operator can say, “Every number I used was real.” It is also highly adaptable, because the same dataset can be sliced into many emotionally useful shapes.

Why it is difficult to counter: A simple “true” or “false” label is inadequate. Correcting the claim may require explaining statistics, sampling, confounding variables and uncertainty. The misleading version fits on a meme; the correction arrives carrying luggage.

6. False Context: Genuine Evidence from the Wrong Place or Time

False context attaches an inaccurate date, location, identity or event to authentic material. It is among the cheapest and most effective forms of visual deception because no sophisticated editing is required.

At the beginning of Russia’s 2022 invasion of Ukraine, a dramatic video was shared as footage of a Russian aircraft being shot down. The clip was genuine, but it actually showed a warplane being downed in Libya in 2011. Other old images and videos were similarly recirculated as evidence of current events.

Read the Associated Press fact check

In another case, a 2016 photograph of Ukrainian children saluting troops was shared as though it had been taken during the 2022 invasion.

Read the Associated Press report on the reused photograph

Why this strategy is effective: Genuine media looks convincing because its pixels contain no obvious fabrication. The emotional impact is real even when the caption is false. Old disasters also provide an enormous warehouse of reusable smoke, explosions and frightened crowds.

Why it is difficult to counter: Image-forensic tools may correctly conclude that the photograph is authentic while missing that it has been miscaptioned. Verification requires reverse-image searching, geolocation, weather analysis, landmark comparison and investigation of the earliest known upload.


Part III: Faking Reality

7. Manipulated Content: Editing Genuine Material to Change Its Message

Manipulated content begins with something real and alters it. The change may be crude, such as cropping out relevant details, or subtle, such as changing playback speed, removing a sentence or rearranging several clips.

In 2019, a video of United States House Speaker Nancy Pelosi was slowed to make her speech appear slurred. The underlying footage was genuine, but the altered speed created a false impression about her condition.

Read the Associated Press report on the altered Pelosi video

Manipulation can include:

  • Cropping a photograph to remove explanatory context
  • Splicing sentences from different moments
  • Changing audio speed or pitch
  • Adding or deleting people and objects
  • Altering screenshots
  • Translating speech inaccurately
  • Editing charts by truncating axes or changing labels

Why this strategy is effective: Manipulated material retains a visible connection to reality. Viewers can recognize the person, location or event, which makes the alteration harder to reject than a completely invented story.

Why it is difficult to counter: The question is no longer simply, “Is this file genuine?” Investigators must determine which parts are genuine, what was changed and whether the alteration affects the meaning. Small edits can require frame-by-frame comparison with the original recording.

8. Fabricated Content: Inventing the Event Itself

Fabricated content is the classic all-out falsehood: a nonexistent quotation, invented document, fictional crime, fake scientific discovery or event that never occurred.

One of the most widely discussed examples from the 2016 United States election was a completely false article claiming that Pope Francis had endorsed Donald Trump. The story imitated the structure of political reporting but had no factual basis. It became a standard example in later information-disorder research.

Read the Information Disorder report

Another modern variant is the fully fabricated “news video.” In 2023, a clip imitating BBC News branding falsely claimed that Ukraine had supplied weapons to Hamas. The BBC and Bellingcat confirmed that the video was not their work and that the supposed report did not exist.

Read the Associated Press fact check

Why this strategy is effective: Fabrication gives the operator complete narrative freedom. There is no inconvenient evidence to accommodate, because the quotation, witness and event can all be designed for maximum emotional impact.

Why it is difficult to counter: Individual fabrications can sometimes be disproved quickly, but unlimited new ones can be produced at low cost. The defender must investigate every claim; the fabricator merely presses “publish” again. Corrections also tend to travel less dramatically than the original accusation.


Part IV: The Borderlands

9. Satire Stripped of Its Warning Label

Satire is not inherently disinformation. Its purpose is usually criticism, comedy or exaggeration, and it does not depend on audiences believing it literally.

Trouble begins when satirical material loses its source label or is deliberately presented as genuine.

In 2012, China’s People’s Daily treated an Onion article naming Kim Jong-un its “Sexiest Man Alive” as a real accolade and produced an extensive online photo presentation. The original item was an obvious joke to its intended audience, but that context did not survive its journey.

Read Reuters’ report on the incident

More recently, Reuters documented a fabricated screenshot styled as a Donald Trump social-media post. It originated as satire but was subsequently shared without that context by users who appeared to regard it as authentic.

Read the Reuters fact check

Why this strategy is effective when weaponized: Satire is emotional, memorable and highly shareable. A malicious distributor can also retreat behind “It was only a joke” after the false impression has spread.

Why it is difficult to counter: Satire is legitimate expression. Platforms cannot simply remove every absurd or fictional statement. The task is to identify when parody has been deceptively relabeled, not to appoint an algorithm as the Minister of Humor.

10. Malinformation: Using Truth as a Weapon

Not every harmful information operation depends on false content. Genuine documents, private correspondence or personal information can be selectively released, stripped of context or timed for maximum disruption.

During the 2017 French presidential election, emails from Emmanuel Macron’s campaign were released shortly before the legally mandated media blackout preceding the final vote. The Information Disorder report classifies this as malinformation: substantially genuine private material deployed at a moment designed to maximize harm and minimize the opportunity for examination or response.

Read the Information Disorder report

Why this strategy is effective: Authentic documents are difficult to dismiss. A large leak also creates an information avalanche in which journalists and citizens cannot examine every file before rumors and selective interpretations begin circulating.

Why it is difficult to counter: Suppressing genuine information can conflict with press freedom and the public interest. Journalists must distinguish legitimate revelations from strategically curated dumps, protect private individuals and avoid amplifying unsupported interpretations.


Why Disinformation Campaigns Combine These Strategies

The most sophisticated disinformation does not choose only one drawer from the toolbox.

A campaign might:

  1. Create a fabricated claim.
  2. Publish it on a cloned newspaper website.
  3. Attach an authentic but unrelated photograph.
  4. Promote it through fake local personas.
  5. Use genuine statistics framed misleadingly.
  6. Encourage real users and journalists to debate it.
  7. Claim censorship when platforms intervene.

The objective is often not to make everyone believe one specific story. It may be enough to produce confusion, exhaust investigators, inflame existing divisions or make citizens conclude that no source can be trusted.

This is why disinformation cannot be addressed through fact-checking alone. Source-checking is equally important, because the origin of a message may reveal more than its surface content. Visual misinformation is particularly difficult to trace, and careless debunking can provide additional publicity to a previously obscure claim.


A Practical Checklist for Evaluating Suspicious Information

Who Is Speaking?

  • Is this the authentic account, website or institution?
  • Is the domain name spelled correctly?
  • Can the statement be found on the organization’s official channels?
  • Does the account have a credible posting history?

What Has Been Done to the Meaning?

  • Does the headline accurately reflect the article?
  • Is the statistic a rate or merely a total?
  • What information is missing from the quotation?
  • Is the image really from the stated date and location?

What Has Been Done to Reality?

  • Has the audio been slowed, clipped or synthetically generated?
  • Does an earlier version of the image or video exist?
  • Do independent sources confirm that the event occurred?
  • Can the original document, recording or dataset be located?

Conclusion: There Is No Single Antidote

The central lesson is simple: disinformation is not merely false content. It is the engineering of false belief.

Sometimes it counterfeits the speaker. Sometimes it bends a fact until it points in the opposite direction. Sometimes it manufactures an event from empty air. Sometimes it uses the truth itself as a blade.

Recognizing which component has been falsified is the first step toward choosing the right response. A false statistic needs analysis. A stolen photograph needs provenance. A cloned newspaper needs source verification. A coordinated influence network needs behavioral investigation.

There is no single antidote because there is no single poison. The information factory has many assembly lines, and each leaves a different kind of fingerprint.

Why Do Some Hindu Temples House Both Shiva and Vishnu? A Journey Through History, Geography, and Harmony

Walk into a Hindu temple almost anywhere in India, and you may expect it to be dedicated to a single deity—perhaps Shiva, Vishnu, Devi, Ganesha, or Murugan. But every so often, you encounter something fascinating: a temple where Shiva and Vishnu stand side by side. Sometimes they occupy separate shrines within the same complex. Occasionally, they even share the same sanctum in a combined form such as Harihara.

How common is this? The answer reveals a remarkable story of India's religious traditions—not one of rigid boundaries, but of dialogue, evolution, and coexistence.

More Common Than Many People Realize

Having both Shiva and Vishnu in the same temple is not the norm, but neither is it particularly rare.

Across India, one can find several patterns:

  • A temple primarily dedicated to Shiva with a shrine for Vishnu.
  • A Vishnu temple that includes a shrine for Shiva.
  • Large temple complexes where multiple deities are worshipped independently.
  • Temples dedicated to Harihara—a deity combining aspects of Shiva (Hara) and Vishnu (Hari).

This arrangement reflects a simple reality: for much of Hindu history, devotees often revered multiple deities without seeing them as mutually exclusive.

The Ancient Roots of Shared Worship

In the earliest centuries of the Common Era, the distinction between Shaivism and Vaishnavism was often less pronounced than it later became.

As temple building flourished between the 5th and 10th centuries, many royal dynasties patronized different traditions simultaneously. Kings frequently supported Shiva temples, Vishnu temples, Jain monuments, and Buddhist monasteries within the same kingdom.

This broad patronage encouraged temple complexes that welcomed diverse forms of worship.

Rather than asking, "Which god is greater?", many communities asked a different question: "How can different paths lead to the same divine?"

South India: A Rich Tradition of Coexistence

Perhaps nowhere is this more visible than in South India.

The great temple cities of Tamil Nadu and Karnataka often contain multiple shrines dedicated to different deities. Even when a temple is unmistakably Shaiva or Vaishnava, secondary shrines are common.

In many Shiva temples, visitors encounter:

  • Vishnu
  • Ganesha
  • Murugan (Kartikeya)
  • Devi
  • Navagrahas

Likewise, major Vishnu temples often include shrines dedicated to Shiva.

This reflects centuries of living religious traditions where families and communities participated in festivals across sectarian lines.

Karnataka and the Legacy of Harihara

Karnataka occupies a special place in this story.

The region saw the flourishing of the Harihara tradition, where Shiva and Vishnu are represented as one unified deity.

The Chalukyas and Hoysalas, known for their magnificent temples, frequently patronized both Shaiva and Vaishnava traditions. Rather than emphasizing rivalry, many monuments celebrated theological unity.

Harihara sculptures remain among the most beautiful artistic expressions of this synthesis.

Tamil Nadu: Distinct Traditions, Shared Spaces

Tamil Nadu is famous for both the Shaiva Nayanmars and the Vaishnava Alvars, whose devotional poetry transformed Hindu worship.

Although these traditions developed their own philosophies and temple networks, everyday religious life often remained inclusive.

Many devotees continue to visit both Shiva and Vishnu temples, especially during major festivals.

Even where theological schools differ, shared pilgrimage practices have long been part of Tamil religious culture.

North India: Flexible Temple Traditions

Northern India presents a somewhat different picture.

Many temples evolved over centuries through repeated rebuilding, expansion, and restoration. As new shrines were added, temple complexes often became home to several deities.

It is common to find:

  • Shiva lingams
  • Vishnu or Rama shrines
  • Hanuman temples
  • Durga shrines

all within the same complex.

The emphasis is often less on sectarian identity and more on providing a complete sacred space for worship.

Kerala: Harmony in Temple Practice

Kerala's temples frequently include multiple shrines within a carefully planned sacred enclosure.

A temple may be dedicated to Shiva while also housing Vishnu, Bhagavathy, Ayyappa, Ganapati, and Subrahmanya.

The architectural layout naturally accommodates several deities without diminishing the importance of the principal deity.

Why Do Some Regions Show More Integration?

Several historical factors influenced these patterns.

Royal Patronage

Many dynasties consciously supported multiple religious traditions. Patronage was often a way of strengthening social harmony and demonstrating that the king served all his subjects.

Local Customs

Village traditions rarely fit neatly into philosophical categories. Families often maintained devotional practices centered on several deities across generations.

Temple Expansion

Large temples rarely emerged all at once. As patronage increased over centuries, additional shrines were added, naturally creating multi-deity complexes.

Were There Ever Differences Between Shaivas and Vaishnavas?

Yes—but the picture is more nuanced than it is sometimes portrayed.

Different philosophical schools developed distinct ideas about theology, ritual, and scripture. Scholars debated these questions vigorously, and devotional literature occasionally reflected spirited disagreements.

Yet these intellectual differences coexisted with widespread social interaction. Many families visited both Shiva and Vishnu temples. Kings patronized both traditions. Temple towns celebrated festivals that involved multiple deities.

The everyday practice of Hinduism often proved more inclusive than formal theological debates might suggest.

The Symbolism of Unity

Perhaps the most beautiful expression of this shared heritage is Harihara, where Shiva and Vishnu are represented as a single divine form.

Rather than asking devotees to choose one over the other, Harihara conveys a profound philosophical idea: that different manifestations of the divine can coexist without contradiction.

This idea appears repeatedly in Indian art, literature, and temple architecture.

So, How Common Is It?

The answer depends on what we mean.

If we ask whether every temple houses both Shiva and Vishnu, the answer is clearly no. Most temples have a principal deity and a distinct identity.

But if we ask whether it is unusual to find both within the same temple complex, the answer is also no. Across much of India—especially in large historic temples—this arrangement has been common for centuries.

Regional traditions certainly vary. South India often preserves elaborate temple complexes with multiple shrines. North India frequently reflects centuries of gradual additions. Karnataka is renowned for Harihara traditions, while Kerala naturally integrates several deities into a single sacred precinct.

Despite these differences, one theme emerges consistently.

A Shared Sacred Landscape

India's temples are more than places of worship—they are living records of history.

They reveal that religious traditions evolved not only through philosophy and scripture but also through communities, artisans, kings, pilgrims, and families who worshipped together over generations.

The presence of Shiva and Vishnu in the same temple is not an exception to Hindu tradition. In many places, it is one of its enduring expressions: a reminder that diversity has long existed alongside devotion, and that different paths have often shared the same sacred space.

Monday, September 7, 2026

How Often Do People Actually Fail a PhD?


A global look at PhD pass rates, failed vivas, attrition—and why the numbers are surprisingly difficult to compare

In my previous post, When Geniuses Nearly Failed Their PhDs, I looked at some remarkable stories from the history of science: Werner Heisenberg's disastrous oral examination, David Hestenes's actual failure of his doctoral oral, J. Robert Oppenheimer's intimidating examination, and several other memorable cases.

But those stories raise a natural question:

How often does someone actually fail a PhD?

The answer is surprisingly difficult to give.

One might imagine that universities routinely publish a simple statistic such as:

"92% of PhD students pass, while 8% fail."

In reality, there is no universally comparable international PhD pass rate. And there is a very important reason for this.

"Failing a PhD" can mean several completely different things.

A student can leave a doctoral programme before submitting a thesis. A candidate can fail a qualifying examination. A submitted thesis can be rejected by external examiners. A candidate can fail a viva. A thesis can be returned for major corrections or resubmission. And, in some countries, a doctorate that has already been awarded can subsequently be subjected to another layer of quality control.

These are not equivalent events.

Once this distinction is made, a fascinating pattern emerges:

The proportion of people who fail to complete a PhD can be substantial, while the proportion of candidates who actually reach the final examination and fail there is usually surprisingly small.


Three different meanings of "PhD failure"

Before comparing countries, it is useful to distinguish three statistics.

1. Attrition

This is the proportion of students who begin a PhD but leave without obtaining the doctorate.

This can happen for many reasons: academic difficulties, financial problems, changes in career plans, health or family circumstances, problems with supervision, loss of funding, or simply the realisation that a PhD is no longer what the student wants.

2. Completion rate

This is the proportion of students who eventually obtain the doctoral degree, usually measured within a specified period.

A student who takes nine years to finish may count as a successful completion in one dataset and as a non-completer in another if the study only follows students for six or seven years.

3. Final-examination failure

This is the number of candidates who actually reach the thesis examination or viva and are ultimately denied the doctorate.

This is the statistic most relevant to the terrifying stories in the previous article.

And it is generally much smaller than the attrition rate.


The United Kingdom: about 96% of viva candidates succeed

The United Kingdom provides one of the more useful datasets for understanding what happens at the final examination.

An analysis based on information obtained from 14 UK universities examined the outcomes of 26,076 doctoral candidates who sat their vivas between 2006 and 2017.

Of these candidates, 25,063 ultimately succeeded.

That is just over 96%.

Approximately 4% therefore did not receive an immediate successful outcome at the viva stage.

But even this number requires some caution.

A PhD viva in Britain is not normally a simple "pass/fail" event. Examiners can require corrections, major corrections, or resubmission.

Indeed, the most common outcome is not an immaculate pass.

It is:

"Yes—but please fix these things."

This is important because a candidate who leaves the examination room with several pages of corrections has not necessarily had a disastrous viva.

They may be on the normal path to receiving the degree.

But what about everyone who started the PhD?

This produces a dramatically different number.

The same analysis estimated an overall UK PhD completion rate of about 80.5%, with approximately 16.2% attributed to students leaving their programmes early and about 3.3% to students who failed the viva.

The precise figures should not be treated as an official national UK statistic—the dataset came from 14 universities—but the distinction is extremely useful.

It tells us something important:

Most PhD "failure" occurs before the final viva, not during it.


Sweden: when final rejection becomes extraordinarily rare

Sweden provides an even more striking example.

A study examined doctoral dissertation rejections in the humanities, law and social sciences across Swedish universities between 1984 and 2017.

The researchers identified only 18 rejected dissertations among 15,477 submitted dissertations.

That is approximately 0.12%.

The cases involved dissertations that had actually reached the examining stage and were rejected by the examining committees.

At first glance, a rejection rate of 0.12% seems almost unbelievable.

Does that mean Swedish doctoral students are nearly infallible?

Obviously not.

The explanation is more interesting.

Doctoral education contains multiple layers of quality control before the final examination. Supervisors, departments, research groups and faculties all have opportunities to identify serious problems.

A weak dissertation can therefore be stopped, delayed or revised long before it reaches the final public defence.

The Swedish study itself discusses this tension: the rarity of rejection may reflect effective quality control, but it may also raise the question of whether some dissertations that should have been rejected are being allowed through.

This leads to a general principle:

The stricter the filters before the final examination, the lower the apparent failure rate at the final examination.


Germany: substantial attrition, but a different examination culture

Germany illustrates another problem with international comparisons.

Doctoral education in Germany does not simply reproduce the British thesis-plus-viva model. The precise structure of the oral examination varies among universities and faculties.

Studies of German doctoral education nevertheless show substantial attrition during the doctoral process.

One longitudinal analysis, for example, found that approximately 71% of doctoral candidates had successfully completed, while around 13% had dropped out and another group remained unresolved at the end of the observation period.

The important point is not the precise percentage, which depends heavily on the cohort and definition of completion.

It is that the major selection process occurs throughout the doctoral programme, rather than being concentrated entirely in one final examination.


The United States: the qualifying examination changes everything

The United States provides perhaps the clearest example of why "PhD pass rate" is an ambiguous expression.

American doctoral programmes frequently contain substantial examinations before the dissertation stage.

Depending on the university and discipline, students may have to pass:

  • coursework requirements;
  • qualifying examinations;
  • comprehensive examinations;
  • candidacy examinations;
  • proposal defences;
  • and finally the dissertation defence.

A student who fails a qualifying examination and leaves the programme is obviously a doctoral non-completer.

But that person never had the opportunity to "fail the PhD viva."

Consequently, American statistics on doctoral completion can look quite different from statistics on dissertation-defence outcomes.

This is why one should be extremely careful when comparing an American "PhD completion rate" with a British "viva pass rate."

They are measuring different stages of the process.


What about India?

This is where the statistics become much more difficult.

There does not appear to be a reliable, publicly available national dataset that allows us to say:

"X% of Indian PhD candidates fail their final viva."

That number should therefore not be invented.

India has enormous numbers of doctoral students and doctoral graduates, but national higher-education statistics generally provide information about enrolment, degrees awarded and disciplines rather than a comprehensive national breakdown of thesis-examination outcomes.

There is, however, something very important in the UGC regulations.

India has a major quality-control filter before the viva

Under the UGC regulations, the PhD thesis is evaluated by the research supervisor and at least two external examiners.

The viva is conducted only when the required external examination recommendations support acceptance of the thesis after incorporation of the suggested corrections.

If one external examiner recommends rejection, the thesis is sent to an alternate external examiner. The viva can proceed only if the alternate examiner recommends acceptance.

If the alternate examiner also does not recommend acceptance, the thesis is rejected and the candidate is declared ineligible for the PhD.

This produces a crucial observation:

In India, a candidate can effectively fail the PhD examination before ever reaching the viva.

Therefore, an Indian university's "viva failure rate" would not necessarily tell us how often doctoral candidates ultimately fail to obtain the degree.


China has something even more unusual: checking the thesis after the PhD has been awarded

China offers one of the most interesting contrasts.

The Chinese system has a formal national mechanism for post-award random inspection of doctoral dissertations.

Under the national regulations, approximately 10% of doctoral dissertations awarded in the previous academic year are randomly selected for inspection each year.

This is fundamentally different from the way most people imagine a PhD examination.

Imagine the following sequence:

  1. The university examines the thesis.
  2. The candidate successfully defends it.
  3. The PhD is awarded.
  4. The dissertation can subsequently be selected for national-level quality inspection.

The Chinese system therefore contains a form of post-award quality control.

The inspection is not simply a second viva. Selected dissertations are reviewed by external experts according to defined criteria.

If experts identify serious problems, the dissertation can be classified as a "problematic dissertation," triggering institutional consequences and remedial measures.

For example, in the 2018 inspection, the Chinese authorities reported that 6,572 doctoral dissertations were randomly inspected, representing approximately 10.4% of all doctoral dissertations awarded that year.

That is an enormous national quality-control exercise.


India and China therefore illustrate two different philosophies of quality control

The comparison is particularly interesting.

In India, the formal examination sequence places substantial emphasis on:

  • external thesis examiners;
  • acceptance of the thesis;
  • corrections recommended by examiners;
  • and the subsequent viva voce.

The UGC regulations explicitly prevent the viva from proceeding in certain circumstances where the thesis has not received the necessary external acceptance.

China, meanwhile, combines institutional examination with a national post-award sampling system in which roughly one in ten doctoral dissertations is independently inspected.

Neither system can simply be described as having a particular "PhD failure rate."

The quality-control mechanisms are distributed across different stages.


So does the PhD pass rate vary by country?

Yes—but not in a way that permits a simple league table.

Consider the following simplified picture:

Country/system What we can reasonably say Major quality-control stage
United Kingdom About 96% of candidates in one large dataset succeeded at the viva Thesis examination + viva
Sweden Extremely rare formal dissertation rejection in one large humanities/social-science study Strong pre-defence filtering + public defence
Germany Substantial attrition during doctoral training Supervision + institutional milestones + dissertation/oral examination
United States Substantial doctoral attrition; final defence is only one of several filters Qualifying/comprehensive exams + dissertation defence
India No reliable national final-viva failure rate identified External thesis examination + viva
China No directly comparable national final-defence failure rate identified Institutional examination + national post-award thesis inspection

The table therefore looks less like a ranking and more like a map of different doctoral systems.


Does it vary by discipline?

Almost certainly—but again, the distinction between attrition and final examination failure is critical.

Different disciplines have very different doctoral cultures.

A laboratory-based biomedical PhD may involve:

  • multiple experiments;
  • large datasets;
  • long periods of troubleshooting;
  • shared equipment;
  • animal or clinical work;
  • and dependence on funding and laboratory infrastructure.

A theoretical mathematics PhD may have a completely different structure.

A humanities doctorate may take substantially longer and may involve a very different relationship between supervisor, candidate and thesis.

Engineering, computer science, social science and the natural sciences each have their own patterns of funding, publication, supervision and completion.

These differences can strongly influence attrition and time to degree.

But we have much less evidence that the probability of outright failure at the final examination differs dramatically between disciplines.

That may be because the final examination is itself the end product of a long selection process.


The great paradox of the PhD pass rate

We can now understand an apparent contradiction.

On the one hand, a substantial fraction of people who begin PhDs never receive a doctorate.

On the other hand, once someone actually reaches the final examination, outright failure can be remarkably uncommon.

There is no contradiction.

It is the consequence of selection.

Think of the PhD as a series of filters:

Admission

Coursework / qualifying examinations

Research progress

Thesis submission

External examination

Viva / defence

PhD

By the time a candidate reaches the last box, many of the people who would have failed earlier have already disappeared from the population being examined.

Therefore:

A 96% viva pass rate does not mean that 96% of people who start PhDs get doctorates.

It means that approximately 96% of the people who survived all the previous filters and reached the viva passed it in that particular dataset.


And this changes how we should interpret Heisenberg's story

Return to Werner Heisenberg.

He did not simply stroll into a university after an undergraduate degree and almost fail a random exam.

He had already completed an advanced research project under Arnold Sommerfeld.

He had demonstrated extraordinary mathematical ability.

He had already established himself as an exceptional young theoretical physicist.

And yet, during the final oral examination, Wilhelm Wien identified a genuine weakness so serious that he reportedly wanted to fail him.

That is precisely why the story is so remarkable.

Heisenberg's examination demonstrates that:

A candidate can be extraordinarily good overall and still have a serious, examination-relevant weakness.

Conversely, the statistics tell us something equally important:

A genuinely catastrophic final PhD examination is unusual because candidates have already passed through multiple filters before reaching it.


What should count as a serious PhD-viva failure?

This distinction is especially important for examiners.

A candidate who cannot remember a minor fact is not necessarily demonstrating doctoral-level incompetence.

A candidate who cannot reproduce a formula from memory is not necessarily unfit for a doctorate.

And a candidate who says "I don't know" to an unexpected question may actually be demonstrating scientific honesty.

Much more serious are situations in which a candidate cannot explain:

  • the central methodology of the thesis;
  • why a particular analysis was performed;
  • what an important parameter means;
  • the assumptions underlying a major conclusion;
  • the limitations of the principal results;
  • why the data support the conclusions;
  • or how the work differs from previous research.

These are not merely memory failures.

They can indicate that the candidate does not have sufficient intellectual ownership of the research.

That is fundamentally different from failing to answer a difficult peripheral question.


The examiner is also part of the experiment

There is another lesson from comparing countries.

A PhD examination is not a purely objective measurement like determining the melting point of a chemical.

The outcome depends on:

  • the candidate;
  • the thesis;
  • the supervisor;
  • the examiners;
  • institutional rules;
  • disciplinary culture;
  • and the examination system itself.

That is why two equally strong candidates could potentially experience very different doctoral examinations in different countries—or even in two departments within the same country.

A particularly important safeguard is therefore the use of multiple examiners.

India's regulations, for example, explicitly provide an additional examiner when one external examiner rejects a thesis.

China's national sampling system similarly introduces an additional layer of external scrutiny after the degree has been awarded.

These mechanisms acknowledge an uncomfortable truth:

Examiners can make mistakes too.


So what is the "real" PhD failure rate?

There isn't one.

And anyone who gives you a single worldwide percentage without defining exactly what they mean is probably comparing apples with oranges.

A more honest summary is:

  • PhD attrition can be substantial.
  • Completion rates vary considerably by country, discipline, institution and cohort.
  • Final thesis/viva failure is generally much rarer than PhD attrition.
  • Countries differ substantially in where they place the major quality-control filters.
  • India does not currently appear to have a robust national statistic for final-viva failure.
  • China has an unusually strong national post-award dissertation inspection system.
  • International comparisons are difficult because "PhD examination" means different things in different systems.

The surprising conclusion

Perhaps the most interesting conclusion is that the terrifying PhD viva is actually the least likely place for most doctoral journeys to fail.

The candidate has already survived years of research.

The supervisor has approved the thesis for submission.

The institution has accepted it for examination.

External experts have evaluated it.

Only then does the candidate walk into the room for the final defence.

That is why stories such as Heisenberg's are so memorable.

They represent an unusual event: a candidate who has survived the entire doctoral pipeline suddenly encountering a weakness that one examiner considers serious enough to threaten the degree.

And that is also why a difficult viva should not automatically be interpreted as a failed PhD.

A difficult viva may be exactly what a doctoral examination is supposed to be.

The real question is not:

"Did the candidate answer every question?"

It is:

"Did the examination provide sufficient evidence that this person has reached the level of an independent researcher?"

That is a much harder question.

And perhaps a much more interesting one.


Sources and further reading

  • DiscoverPhDs — analysis of 26,076 PhD candidates at 14 UK universities, based on Freedom of Information data, 2006–2017.
  • Stigmar, M. (2019), Learning from reasons given for rejected doctorates: drawing on some Swedish cases from 1984 to 2017, Higher Education.
  • University Grants Commission — UGC regulations concerning external thesis examination and the conditions under which the PhD viva may proceed.
  • Ministry of Education, China — regulations for random inspection of doctoral and master's dissertations.
  • Ministry of Education, China — report on the 2018 national doctoral dissertation inspection, including 6,572 dissertations sampled.

A note on the statistics: PhD completion, attrition and final examination failure are different quantities and should not be treated as interchangeable. The figures above therefore describe specific datasets or examination systems rather than providing a universal international "PhD pass rate."

Sunday, September 6, 2026

How to Benchmark Without Fooling Yourself

Ten timeless lessons for building trustworthy computational methods

Every year, thousands of new computational methods are published. New machine learning models. New statistical techniques. New optimization algorithms. New bioinformatics pipelines.

Almost every paper proudly claims:

"Our method outperforms the state of the art."

Yet, a curious paradox exists.

If every new method is better than every previous method, why do independent benchmarking studies often tell a very different story?

The answer lies in one word:

Benchmarking.

A benchmarking study is far more than running several algorithms on a dataset and producing a leaderboard. Done well, it becomes the scientific equivalent of a fair sporting competition. Done poorly, it becomes advertising disguised as science.

A wonderful review published in Genome Biology distills years of experience into ten practical principles for designing reliable computational benchmarks.

These principles apply not only to computational biology but to machine learning, robotics, AI, computer vision, signal processing, and virtually every computational discipline.

Let's explore them.


Why benchmarking matters

Imagine testing a new car.

If you only drive it downhill with the wind behind you, you'll conclude it's the fastest car ever built.

But that's not how people actually drive.

You need:

  • highways

  • traffic

  • hills

  • rain

  • fuel efficiency

  • braking distance

  • maintenance costs

Only then do you know whether it's actually a good car.

Computational methods are exactly the same.

A benchmark should answer:

"When should I use this method instead of another?"

not

"Can I find one dataset where my algorithm wins?"

That difference separates science from marketing.


The Ten Golden Rules


1. Clearly define the purpose

Not every benchmark serves the same goal.

The paper identifies three major categories:

Developer benchmark

Created by authors introducing a new algorithm.

Purpose:

"Is my new method better than existing ones?"


Neutral benchmark

Performed independently.

Purpose:

"Which methods actually work best?"

These are generally the most valuable because they reduce author bias.


Community challenge

Large collaborative competitions such as DREAM or CASP.

Purpose:

Push the entire field forward.

Before writing a single line of code, define which category your benchmark belongs to.


2. Compare against all relevant methods

Nothing weakens a paper faster than comparing against outdated competitors.

A benchmark should include:

  • current state-of-the-art methods

  • strong baseline methods

  • widely used methods

  • publicly available implementations

Excluding an important competitor simply because it performs well introduces obvious bias.

The goal is not to make your method look good.

The goal is to discover the truth.


3. Use realistic datasets

This is perhaps the most important lesson.

A benchmark is only as good as its datasets.

The paper recommends combining:

Simulated datasets

Advantages:

  • known ground truth

  • unlimited size

  • controlled experiments

Disadvantages:

  • may not resemble real-world data


Real datasets

Advantages:

  • realistic complexity

  • biological variability

  • genuine challenges

Disadvantages:

  • often lack known answers


Hybrid datasets

The authors particularly like semi-simulated data, where real data are combined with carefully inserted synthetic signals. This provides realistic variability while retaining a known ground truth.

The best benchmarks rarely rely on only one type of dataset.


4. Treat every method fairly

Parameter tuning can completely change performance.

Suppose:

Method A

  • carefully tuned for weeks

Method B

  • default settings

Method A wins.

But did it really?

Probably not.

The paper stresses that all methods should receive comparable effort during tuning. Otherwise, the benchmark measures the researcher's effort rather than the algorithm itself.


5. Measure what actually matters

Accuracy alone is almost never enough.

Depending on the problem, evaluate metrics such as:

  • precision

  • recall

  • F1-score

  • ROC-AUC

  • precision-recall curves

  • false discovery rate

  • correlation

  • RMSE

  • robustness

  • stability

The review's diagram (Figure 2) organizes evaluation metrics into quantitative and qualitative categories, emphasizing that different tasks require different measures.

A single score rarely captures the whole story.


6. Evaluate practical usability

Imagine two algorithms.

Algorithm A

  • 98% accurate

  • requires 128 GB RAM

  • runs for 14 hours

Algorithm B

  • 97% accurate

  • finishes in 30 seconds

  • installs with one command

Which would most researchers choose?

Probably Algorithm B.

The paper argues that good benchmarks should report:

  • runtime

  • memory usage

  • scalability

  • ease of installation

  • documentation quality

  • software maintenance

  • user friendliness

These often determine real-world adoption more than a small accuracy gain.


7. Avoid declaring a single winner

One of the paper's most refreshing messages is this:

There may not be a universally best method.

Instead of publishing one leaderboard, identify:

  • methods consistently performing well

  • strengths of each method

  • weaknesses of each method

  • situations where each excels

Different users care about different things.

Some value speed.

Others prioritize accuracy.

Others need scalability.

A benchmark should help users make informed decisions rather than crown a single champion.


8. Present results clearly

Good science is useless if nobody understands it.

The authors recommend:

  • summary tables

  • intuitive plots

  • interactive websites

  • decision flowcharts

  • open-access publication

One striking example in the paper is an interactive benchmarking website where users can filter methods by accuracy, scalability, stability, and memory requirements instead of relying on a static table.


9. Design benchmarks that can grow

Methods evolve rapidly.

A benchmark published today may become outdated within a year.

Instead of treating benchmarking as a one-time event, design it so others can extend it by adding:

  • new algorithms

  • new datasets

  • new evaluation metrics

  • updated software versions

Science progresses faster when benchmarks become living resources rather than frozen snapshots.


10. Make everything reproducible

Perhaps the most important principle of all:

If nobody can reproduce your benchmark...

...then nobody can trust it.

The paper recommends publishing:

  • source code

  • datasets

  • software versions

  • parameter settings

  • random seeds

  • workflow scripts

  • containerized environments (Docker, Singularity)

  • public repositories (GitHub, Zenodo, etc.)

Reproducibility transforms a benchmark from a claim into a scientific asset that others can verify, reuse, and improve.


Common benchmarking mistakes

The review also highlights pitfalls that quietly undermine many studies:

  • Choosing datasets that favor your method.

  • Comparing against weak or outdated competitors.

  • Tuning only your own algorithm.

  • Reporting a single metric instead of multiple perspectives.

  • Ignoring computational cost.

  • Overstating tiny performance differences.

  • Hiding code or datasets.

  • Failing to discuss benchmark limitations.

These mistakes can mislead both users and future research directions.


What this means beyond computational biology

Although written for computational biology, these recommendations apply remarkably well across fields.

Whether you're benchmarking:

  • large language models,

  • computer vision systems,

  • reinforcement learning algorithms,

  • robotics controllers,

  • signal processing pipelines,

  • optimization methods, or

  • autonomous navigation systems,

the same principles hold:

  • Compare fairly.

  • Use realistic data.

  • Measure multiple dimensions of performance.

  • Report limitations honestly.

  • Share everything needed to reproduce the work.

These habits build trust, accelerate progress, and make comparisons genuinely useful.


Actionable Checklist Before Publishing a Benchmark

Before submitting your next paper, ask yourself:

  • ✅ Have I clearly stated the purpose of the benchmark?

  • ✅ Have I included all strong competing methods?

  • ✅ Are my datasets representative of real-world applications?

  • ✅ Were all methods tuned with comparable effort?

  • ✅ Am I reporting multiple performance metrics instead of just accuracy?

  • ✅ Did I measure runtime, memory usage, and usability?

  • ✅ Am I discussing strengths, weaknesses, and tradeoffs instead of declaring one "best" method?

  • ✅ Are my figures and tables easy to interpret?

  • ✅ Can others extend my benchmark with new methods?

  • ✅ Have I released code, data, parameters, and software versions so others can reproduce the results?

If you can confidently check every box, you're much closer to producing a benchmark that informs the community rather than simply supporting a single method.


Final Thoughts

The central message of this review is deceptively simple: benchmarking is not about proving that one method wins. It is about helping the community make better decisions. A trustworthy benchmark values fairness over favoritism, transparency over selective reporting, and practical insight over flashy leaderboards.

The most influential benchmarks are not remembered because they produced the highest accuracy score. They are remembered because researchers trusted them, built upon them, and used them to move the field forward.

The Semmelweis Reflex: When Truth Knocks and the Mind Pretends Nobody Is Home 🧼🕯️

Anybody who has watched a false rumor travel through a crowd knows that people do not simply believe what is true. We believe what fits. We believe what preserves our dignity, our tribe, our training, our status, and our existing map of the world. That is why fake news is not merely an information problem. It is a human reflex problem.

One of the oldest names for this reflex comes from the tragic story of Ignaz Semmelweis, the Hungarian physician who discovered that doctors could save mothers’ lives by washing their hands, and was punished by history before being honored by it. The term Semmelweis reflex now refers to the tendency to reject new evidence because it contradicts established beliefs, habits, or professional norms. In one medical review, it is described as the human tendency to cling to preexisting beliefs and reject fresh ideas even when adequate evidence is available.

The story is not just about handwashing. It is about how truth can look offensive before it looks obvious.

The hospital with two doors

In the 1840s, Semmelweis worked at the Vienna General Hospital, one of the great medical institutions of Europe. Its maternity clinic had two divisions. The First Clinic was staffed by doctors and medical students. The Second Clinic was staffed largely by midwives. Women in Vienna knew the difference. Some reportedly begged not to be admitted to the doctors’ clinic because death from childbed fever was so much more common there.

The puzzle tortured Semmelweis. Same hospital. Same city. Same disease. Different deaths.

The key came after the death of his colleague Jakob Kolletschka, who died after being accidentally cut during an autopsy. Kolletschka’s illness resembled the childbed fever killing postpartum women. Semmelweis reasoned that doctors and students were carrying “cadaverous particles” from autopsy rooms to maternity wards on their hands. In 1847, he introduced mandatory handwashing with chlorinated lime before examining patients. Reports of his intervention describe a dramatic fall in maternal mortality, from roughly 16% to below 2% within months.

Today, this sounds almost painfully obvious. Doctors dissecting corpses in the morning and examining women in labor in the afternoon should wash their hands. But before germ theory, Semmelweis had data without the accepted explanatory machinery. He could show that handwashing worked, but he could not fully explain why in terms his peers accepted.

That was enough to make him dangerous.

Why doctors rejected the evidence

The resistance to Semmelweis was not simply stupidity. That would be a comforting fairy tale, and history rarely provides such tidy villains.

His claim attacked professional identity. It implied that respected physicians were not merely failing to save women but actively killing them. It contradicted dominant theories of disease, including miasmatic explanations based on bad air and constitutional weakness. It also demanded a behavioral change that was unpleasant, inconvenient, and humiliating. Chlorinated lime irritated the skin and smelled awful. But the deeper sting was moral: the doctor’s healing hand had become a weapon.

Semmelweis also did not help his case with diplomacy. Over time, he grew angry and accusatory, writing polemical letters to physicians and denouncing their failure to adopt his method. His major book appeared in 1861, but broad acceptance came only later, after Pasteur’s germ theory and Lister’s antiseptic surgery gave the medical world a framework in which Semmelweis’s results finally made sense. Semmelweis died in 1865, shortly after being confined to an asylum, a grim end to a life now often remembered through the phrase “savior of mothers.”

The tragedy is that the evidence was there. But evidence does not float into the mind like a feather. It must pass through gates guarded by reputation, theory, habit, and ego.

The mirror image of fake news

Fake news is usually discussed as the acceptance of falsehood. The Semmelweis reflex is the rejection of truth. They look opposite, but they are siblings.

Fake news succeeds when a false claim feels more emotionally satisfying, socially useful, or identity-protective than the truth. The Semmelweis reflex operates when a true claim feels too threatening to accept. In both cases, the mind is not acting as a neutral courtroom. It is acting as a border checkpoint, stamping passports according to familiarity, loyalty, fear, and cost.

Modern misinformation research shows that false news can spread faster and more widely than true news online. A major study in Science found that false news diffused farther, faster, deeper, and more broadly than true news on Twitter, in part because false stories often appeared more novel. Another Science review argued that addressing fake news requires understanding not only online platforms but also how people process information, trust sources, and form beliefs.

The Semmelweis story adds a darkly funny twist: novelty can help falsehood travel, but novelty can also make truth look suspicious. The new claim arrives wearing the wrong clothes. It does not sound like the old textbooks. It does not flatter the right authorities. It asks people to admit that yesterday’s confidence was today’s error.

That is why corrections are so difficult. Psychological research on misinformation has shown that false beliefs can continue to influence people even after correction, a phenomenon often called the continued influence effect. Retractions may fail when they leave a mental gap, when the original story remains easier to remember, or when the correction threatens identity.

Semmelweis faced a similar problem before social media existed. The false story was not a viral WhatsApp forward. It was the medical worldview itself.

The anatomy of the reflex

The Semmelweis reflex has a recognizable structure:

First, an observation appears that does not fit the reigning explanation.

Second, accepting the observation would impose a cost. Someone would lose face, authority, money, convenience, or theoretical elegance.

Third, critics attack the messenger, the method, the tone, or the incompleteness of the explanation rather than grappling with the pattern.

Fourth, the evidence waits for a new framework. Once the framework changes, the once-ridiculous claim becomes “obvious.”

This is why Semmelweis is so relevant to fake news. We often imagine misinformation as garbage entering a clean machine. But the machine is not clean. It has gears shaped by incentives. Institutions can reject truth when truth threatens their internal order. Communities can embrace falsehood when falsehood protects their shared story.

In one case, truth is blocked. In the other, falsehood is welcomed. The same gatekeeper is asleep in one direction and ferocious in the other.

Parallel 1: John Snow and the pump handle

John Snow’s cholera work offers a close public-health parallel. In mid-19th-century London, cholera was widely interpreted through miasma theory, the idea that disease spread through foul air. Snow argued that cholera was waterborne. During the 1854 Soho outbreak, he mapped deaths and linked them to the Broad Street pump. His work became foundational for epidemiology, but acceptance was not instantaneous because the dominant miasma framework was deeply embedded in public-health thinking.

Snow’s case resembles Semmelweis in one crucial way: both men had strong empirical patterns before the microbial mechanism was widely accepted. Their evidence was not weak. It was homeless. It had not yet found the theory that could house it.

Parallel 2: Barry Marshall, Robin Warren, and the ulcer bacterium

For much of the 20th century, peptic ulcers were commonly attributed to acid, stress, and lifestyle. Barry Marshall and Robin Warren argued that many ulcers were caused by the bacterium Helicobacter pylori. The idea was initially resisted because the stomach was considered too acidic for bacteria to play such a central role. Marshall later drank a culture of H. pylori to demonstrate that the organism could infect a healthy person and cause gastritis. He and Warren received the 2005 Nobel Prize in Physiology or Medicine for discovering H. pylori and its role in gastritis and peptic ulcer disease.

This is Semmelweis with better timing and better eventual reward. Again, the problem was not absence of evidence alone. It was the inertia of a medical story that already felt complete.

Parallel 3: Alfred Wegener and continental drift

Alfred Wegener proposed that continents had once been joined and later drifted apart. He assembled evidence from fossil distributions, geological continuity, and the fit of continental margins. Yet his hypothesis was largely rejected for decades, partly because he lacked a convincing mechanism for how continents could move. Only later, with ocean-floor exploration and plate tectonics, did the broad idea become accepted.

Wegener’s case shows an important nuance. Skepticism was not entirely irrational. A theory without a mechanism can be incomplete. But the Semmelweis reflex begins when incompleteness becomes an excuse to ignore a strong pattern. “You have not explained everything” quietly becomes “you have shown nothing.”

Parallel 4: James Lind and scurvy

James Lind’s 1747 citrus trial is often remembered as an early controlled clinical experiment showing that citrus fruits could treat scurvy. Yet the widespread adoption of lemon juice in the British Navy took decades. Historical reassessments note that Lind’s work preceded the Navy’s routine use of citrus by about half a century.

This delay reminds us that evidence does not implement itself. Between knowing and doing lies a swamp of logistics, hierarchy, economics, competing explanations, and institutional sleepiness. Truth may arrive by ship, but policy often travels by ox cart.

The dangerous misuse of the Semmelweis story

There is one trap here. Every rejected idea is not a future breakthrough. Most rejected ideas are rejected because they are wrong, weak, overclaimed, or unsupported. The phrase “they laughed at Semmelweis” should not become a shield for conspiracy theories.

This matters deeply in the age of fake news. Anti-vaccine influencers, pseudoscientists, miracle-cure sellers, and conspiracy entrepreneurs often borrow the costume of the persecuted genius. They say, “Semmelweis was rejected too.” But the lesson of Semmelweis is not “believe every outsider.” The lesson is “build systems that can evaluate uncomfortable evidence fairly.”

Semmelweis had a measurable intervention, a defined outcome, and a large mortality change. Snow had maps and natural comparisons. Marshall and Warren had pathology, culture, infection, treatment, and eventual clinical transformation. Wegener had converging geological evidence, though an incomplete mechanism. Lind had comparative treatment observations.

The real parallel is not rejection. The real parallel is evidence plus resistance.

Without evidence, persecution is just branding.

What Semmelweis teaches us about fake news

The Semmelweis reflex and fake news are both failures of epistemic hygiene. That phrase sounds academic, but the metaphor is deliciously literal here. Handwashing saved bodies. Mind-washing, in the honest sense of cleaning our belief-forming habits, may save public reasoning.

A healthy information culture needs three habits:

First, separate the claim from the discomfort it causes.
A fact that threatens your profession, political identity, religious instinct, or favorite theory is still allowed to be true.

Second, ask what evidence would change your mind.
If the answer is “nothing,” you are not reasoning. You are guarding a shrine.

Third, distinguish skepticism from immune rejection.
Skepticism examines evidence. The Semmelweis reflex swats it away before examination.

Fake news spreads because some falsehoods feel useful. Semmelweis-like truths fail because some truths feel unbearable. Both reveal the same unsettling fact: human beings are not only belief-forming animals. We are belief-protecting animals.

The final irony

Semmelweis asked doctors to wash their hands. The deeper request was harder: wash the profession’s assumptions.

That remains the harder demand today. We can fact-check a headline, debunk a rumor, retract a paper, or correct a statistic. But the more difficult work is changing the social conditions under which people decide whether truth deserves entry.

The Semmelweis reflex is not just a medical-historical curiosity. It is a warning label on the human mind:

The truth may arrive before we are ready for it.

And when it does, the first symptom is often not curiosity.

It is irritation.

When Geniuses Nearly Failed Their PhDs

When Geniuses Nearly Failed Their PhDs

Famous scientists, terrifying oral examinations, and what a viva can—and cannot—tell us

There is a particular kind of academic nightmare that is difficult to explain to anyone who has never experienced a PhD oral examination.

You spend years becoming one of the world's leading experts on a very small subject. You know your experiments, your equations, your literature and your thesis.

Then three professors sit across from you and ask:

"But why did you use that method?"

You answer.

They ask:

"What would happen if you changed that assumption?"

You answer again.

Then, suddenly:

"And what is the resolving power of a microscope?"

And you discover that the entire trajectory of your scientific career may apparently depend on something you last studied as an undergraduate.

It is easy to imagine that the greatest scientists in history would have sailed effortlessly through such examinations.

They did not.

Some performed spectacularly badly. Some were nearly failed. Some were so aggressive, eccentric or intellectually overpowering that the examination became uncomfortable for the examiner.

And a few stories are particularly revealing because the very question that exposed a candidate's weakness later became connected to their greatest scientific achievement.

The most famous example is Werner Heisenberg.

But he is far from the only one.


1. Werner Heisenberg: the PhD oral that almost ended in failure

This is the classic story, and unlike many internet anecdotes about famous scientists, it is exceptionally well documented.

In July 1923, a 21-year-old Werner Heisenberg appeared before the examination committee at the University of Munich.

He was already an extraordinary theoretical physicist.

His supervisor, Arnold Sommerfeld, regarded him as one of his most gifted students. His doctoral dissertation concerned the stability and transition of fluid flow from laminar to turbulent motion—a formidable mathematical problem.

Sommerfeld had deliberately chosen hydrodynamics rather than quantum theory because Heisenberg's interests in the emerging quantum physics were still regarded as unconventional.

The thesis itself was difficult enough that Sommerfeld later said he would not have assigned such a demanding problem to any of his other students.

The thesis passed.

Then came the oral.

And everything went wrong.

The microscope question

One of Heisenberg's examiners was Wilhelm Wien, a Nobel Prize-winning experimental physicist.

Wien was unimpressed by Heisenberg's laboratory skills.

During the examination he asked Heisenberg about the resolving power of a Fabry–Perot interferometer—an instrument Heisenberg had actually encountered in his experimental physics course.

Heisenberg could not derive it.

Wien then moved to something more familiar:

the resolving power of a telescope.

Heisenberg struggled again.

Then:

the resolving power of a microscope.

Again, Heisenberg could not provide the answer.

Wien apparently became increasingly exasperated.

Finally, he asked how a storage battery works.

Heisenberg was unable to give a satisfactory answer to that either.

This was not simply a case of an examiner asking irrelevant trick questions. Wien believed that a physicist should understand experimental physics. Heisenberg had taken Wien's laboratory course and had performed poorly in it.

From Wien's perspective, this was a genuine deficiency.

He wanted to fail him.

Sommerfeld disagrees

Arnold Sommerfeld strongly disagreed.

The committee was effectively confronted with two very different assessments of the same student.

Wien saw an experimental physicist who could not answer elementary questions about optical instruments or a battery.

Sommerfeld saw an extraordinarily gifted theoretical physicist with an exceptional command of mathematics and physical intuition.

The result was a compromise.

Heisenberg received the lowest passing grade in physics and the same overall grade for the doctorate.

He was devastated.

He left a small celebration at Sommerfeld's home, packed his things and took the midnight train to Göttingen.

The next morning he appeared in Max Born's office and asked whether Born still wanted him as an assistant.

Born had him work through the questions he had failed before deciding that Heisenberg's performance was not sufficient reason to withdraw the offer.

The irony gets almost too good

Four years later, the microscope came back.

In 1927, Heisenberg was developing what became the uncertainty principle.

To explain the physical meaning of quantum mechanics, he considered a hypothetical gamma-ray microscope.

The basic idea was that if you wanted to determine an electron's position extremely accurately, you would need radiation with an extremely short wavelength.

But the photon used to locate the electron would also transfer momentum to it.

Better positional resolution therefore came at the price of greater uncertainty in momentum.

The famous relation emerged:

Δx Δp ≥ ℏ/2

The historical irony is extraordinary:

The man who had nearly failed his doctorate because he could not explain the resolving power of a microscope subsequently used a microscope as part of the physical argument surrounding the uncertainty principle.

There is an important qualification, however. The popular version of this story sometimes suggests that Wien's question directly inspired the uncertainty principle. The evidence does not establish such a simple causal connection. Heisenberg's later microscope argument was also technically imperfect and was criticised by Niels Bohr. The modern uncertainty relation is more fundamentally a mathematical property of quantum states than simply a statement about disturbance by an imperfect measuring instrument.


2. David Hestenes: the physicist who actually failed his doctoral oral

Heisenberg nearly failed.

David Hestenes actually did.

Hestenes later became an influential American mathematician and physicist, particularly known for the development and advocacy of geometric algebra, a mathematical framework combining algebra and geometry.

But his path through graduate school included a spectacular setback.

In his autobiographical account of the development of geometric algebra, Hestenes recalls failing his doctoral oral examination.

The committee asked him to solve a standard problem in quantum mechanics.

Unfortunately, it happened to concern one of the gaps in his background.

He struggled.

Eventually he managed to solve the problem after about half an hour, but the committee was unimpressed. They required him to spend another year studying before repeating the examination.

Hestenes later came to regard the decision rather differently.

He thought that, with hindsight, he had actually performed better than many students might have, and that the committee could perhaps have passed him.

But he also concluded that the committee had been trying to protect him from progressing too rapidly.

That is an unusually mature interpretation of failing an oral exam.

The failure was not necessarily saying:

"You are not capable of doing research."

It was saying:

"There is a hole in your foundation, and we think you should fix it before proceeding."

The distinction matters enormously.


3. J. Robert Oppenheimer: when the examiner felt like he was being examined

Not every frightening oral examination ends in a poor grade.

Sometimes the candidate terrifies the examiner instead.

J. Robert Oppenheimer completed his PhD at the University of Göttingen in 1927 under Max Born.

The oral examination became legendary for a very different reason.

After the examination, James Franck—the Nobel Prize-winning physicist who administered the examination—was reportedly heard to say:

"I'm glad that's over."

The reason? Oppenheimer had apparently been so intellectually aggressive and probing that Franck felt that the candidate was getting close to questioning the examiner.

The story should be treated as a reported anecdote rather than a verbatim transcript of the examination. But the general picture fits documented descriptions of Oppenheimer's personality at Göttingen.

He was intellectually intense, exceptionally quick and prone to taking over discussions.

His fellow students eventually became sufficiently frustrated that some reportedly petitioned Max Born to do something about Oppenheimer's dominance in seminars.

Born left the petition where Oppenheimer could see it.

The message worked.

Oppenheimer's case illustrates a different kind of viva problem

A PhD oral examination is not simply an intelligence test.

A candidate can know an enormous amount and still perform badly.

Conversely, a candidate can be so intellectually forceful that the examination becomes difficult to control.

Oppenheimer seems to have belonged to the second category.

His problem was not lack of knowledge.

It was almost the opposite:

He could generate questions faster than the examination could contain them.

That is a very different viva disaster.


4. Hermann J. Muller: the Nobel laureate who was terrified before his oral

The story of geneticist Hermann J. Muller gives us a more human version of the same problem.

Muller would eventually win the 1946 Nobel Prize in Physiology or Medicine for his discovery of X-ray-induced mutations in Drosophila.

But before his doctoral examination at Columbia University in 1915, he was extremely nervous.

His mentor Edmund Beecher Wilson tried to calm him by telling him about his own nightmare before his doctoral examination decades earlier.

Wilson described dreaming that he was standing before his examiners—and that among them were Charles Darwin, Thomas Huxley and Ernst Haeckel.

Wilson announced that he had solved the problem of life.

He then drew two triangles on the blackboard.

His conclusion?

That was the secret of life.

In the dream, Darwin threw his hat into the air and congratulated him.

Muller later remembered the story.

The remarkable thing is that this is not merely an internet anecdote. The Cold Spring Harbor Laboratory Archives reconstructed the episode from archival material, including the diary of geneticist and historian Elof Carlson, who had heard the story directly from Muller.

Muller did not fail his oral.

But the story is worth including because it shows something that is often forgotten when we look backward at Nobel laureates:

They were once terrified graduate students too.


5. Hugh Huxley: a thesis examination at the edge of a scientific revolution

Another fascinating biology example involves Hugh Huxley, one of the founders of modern muscle biology.

Huxley was working on the fine structure of muscle at a time when the interpretation of muscle structure was undergoing a major transformation.

Researchers were beginning to realise that muscle contraction could not be explained simply by the shortening of individual protein filaments.

The emerging idea was that two sets of filaments might instead slide past one another.

Around this period, Huxley's work, together with the work of Andrew Huxley, John Kendrew, Francis Crick and others, helped establish the structural basis for what became the sliding-filament model of muscle contraction.

The important point for our discussion is that a PhD examination can take place at precisely the moment when the candidate is still figuring out what the evidence means.

The oral examination is therefore not necessarily the interrogation of a finished scientist.

It is an assessment of someone who is still becoming a scientist.


6. Sydney Brenner: when the examiner nearly failed the candidate because he disliked the thesis

Sydney Brenner provides perhaps the most entertaining perspective because he was himself an examiner.

In his recollections, Brenner describes examining a PhD thesis in biochemistry.

He was unimpressed by the dissertation.

The work, in his view, contained speculative sections unsupported by direct experimental evidence and repetitive experiments. He drafted an extraordinarily sarcastic report recommending that the candidate not receive the degree.

Then he reconsidered.

Brenner realised that he had no objective basis for allowing his personal irritation with the thesis's author to determine the candidate's fate.

So he discarded the devastating report and wrote a conventional one.

The candidate received the PhD.

The story is valuable precisely because it comes from the examiner's side.

A viva can be influenced by something surprisingly dangerous:

the examiner's personality.

An examiner can become annoyed by:

  • the writing;
  • the supervisor;
  • the choice of terminology;
  • the interpretation of the data;
  • the field itself;
  • or simply the candidate's manner.

The scientific merits of the thesis can become entangled with these reactions.

Brenner recognised that danger in himself.

And he corrected for it.


7. Philip Eaton: when the chemistry examiners themselves disagreed

Chemistry provides fewer famous examples of outright PhD-viva disasters, at least ones that are documented well enough to survive historical scrutiny.

But there are some wonderful examination stories.

One comes from organic chemist Philip Eaton, whose PhD research involved early applications of NMR spectroscopy.

During his oral examination, two giants of organic chemistry—John D. Roberts and Robert Burns Woodward—became involved in a vigorous argument over the interpretation of Eaton's spectra.

Eaton later realised that both examiners had been wrong about the interpretation.

This is a beautiful example of why a PhD oral is not necessarily an interrogation in which the examiner possesses the answer and the candidate is supposed to reproduce it.

Sometimes the candidate is closer to the truth than the examiner.

And sometimes the examiner's expertise is itself bounded by the technology and conceptual framework of the time.

NMR was still developing rapidly.

What seems obvious in a modern textbook can be genuinely uncertain at the frontier of a field.


8. Linus Pauling: a reminder that even great scientists have historical blind spots

Linus Pauling is an interesting chemistry counterexample.

There is no strong evidence that Pauling nearly failed his PhD oral. He successfully completed his PhD at Caltech in 1925.

But his examination is interesting from a historical perspective because the chemistry of his doctoral period was undergoing an enormous conceptual transformation.

A modern Pauling transported back to his doctoral examination with decades of subsequent knowledge would know quantum chemistry, spectroscopy and laboratory technologies that his examiners could scarcely have imagined.

This raises an important question:

How much of a PhD oral should test timeless knowledge, and how much should test the candidate's mastery of the scientific world available at the time?

That question is particularly important in rapidly changing fields.


9. The famous oral examinations that probably didn't happen the way the internet says

There is another lesson in researching these stories.

The internet is full of claims such as:

  • "Einstein failed his PhD oral."
  • "A Nobel laureate couldn't answer basic questions."
  • "The examiner tried to fail a famous scientist."
  • "A professor asked a trick question that changed the history of science."

Many of these stories are either exaggerated or involve a different examination altogether.

For example, one famous story concerns Mileva Marić, Einstein's future wife.

She failed the final ETH examination in 1900, whereas Einstein passed narrowly. But this was not a PhD oral examination. It was the final examination for the ETH diploma.

Similarly, many American scientists underwent qualifying examinations or comprehensive examinations that should not be confused with the final oral defence of a doctoral thesis.

This distinction matters.

A PhD student can fail:

  1. a qualifying examination;
  2. a comprehensive examination;
  3. a candidacy examination;
  4. a thesis proposal defence;
  5. a final thesis defence;
  6. or a separate oral examination in a particular subject.

These are not interchangeable.


Why are genuinely failed famous PhD vivas so rare?

There is an interesting historical reason.

If a truly famous scientist had an ordinary PhD oral, historians often simply never recorded it.

But if the examination was disastrous, humiliating or bizarre, it was much more likely to become part of the scientist's biography.

This creates a strong selection effect.

We remember:

Heisenberg could not answer the microscope question.

We do not remember the thousands of future professors who gave perfectly competent answers to every question.

There is another reason.

Most universities do not want the final PhD examination to function as a lottery.

By the time someone reaches a thesis defence, the supervisor, department and university have already invested years in the candidate.

The thesis has been submitted.

The examiners have been selected.

The candidate's research record is available.

A complete failure should therefore normally reflect something substantial:

  • the thesis does not constitute an original contribution;
  • the candidate cannot explain or defend central results;
  • the data do not support the conclusions;
  • there are major methodological flaws;
  • the candidate does not understand the work;
  • or the candidate cannot demonstrate the expected level of independent scholarship.

A candidate simply forgetting a formula is not necessarily enough.


The most frightening question is often not the hardest question

This may be the biggest lesson from these historical cases.

The question that causes a candidate the most trouble is not necessarily the one requiring the most advanced knowledge.

Heisenberg's problem was not an obscure equation in quantum field theory.

It was:

How does a microscope resolve two nearby objects?

For another candidate it might be:

Why did you choose this statistical test?

Or:

What is the parameter you defined in Chapter 4?

Or:

What would happen if this assumption were violated?

These questions are frightening because they probe whether the candidate actually understands the foundations of their own work.

An examiner who asks a difficult but relevant question may discover that the candidate knows the subject extremely well.

An examiner who asks a seemingly simple question can sometimes discover something much more serious:

The candidate has been operating a method without understanding it.


The Heisenberg lesson has a dangerous corollary

It would be easy to draw the wrong conclusion from these stories.

The lesson is not:

"Oral examinations are useless because Heisenberg almost failed one."

That would be absurd.

Wien identified a genuine weakness in Heisenberg.

Heisenberg really did have poor experimental knowledge.

The lesson is instead:

A PhD examination samples a multidimensional ability using a very small number of observations.

Heisenberg's oral revealed something true about him.

It simply did not reveal everything that was true about him.

That distinction is crucial.


What should a PhD oral actually determine?

A good doctoral examination should ideally establish several things.

1. Does the candidate understand the thesis?

Not just the conclusions.

The candidate should understand:

  • why the question was asked;
  • why particular methods were chosen;
  • what the assumptions were;
  • what the alternatives were;
  • what the limitations are;
  • and what the results actually demonstrate.

2. Can the candidate reason beyond the thesis?

This is where a good viva becomes much more than a presentation.

The examiner should be able to ask:

"What if...?"

and see whether the candidate can reason through the consequences.

3. Does the candidate understand the foundations?

A candidate doing evolutionary genomics should know more than the commands used to run an analysis.

A candidate doing molecular biology should understand more than the protocol.

A candidate doing theoretical physics should understand more than the equations in the thesis.

A doctorate represents independent scholarship, not merely successful completion of a research project.

4. Can the candidate recognise uncertainty?

This may be one of the most important signs of scientific maturity.

The strongest candidates can say:

"I don't know."

and then continue:

"But here is how I would find out."

That is much more impressive than inventing an answer.


And what should an examiner not conclude?

A candidate who cannot answer one question is not necessarily incompetent.

A candidate who becomes nervous is not necessarily unprepared.

A candidate who forgets a formula is not necessarily ignorant.

And a candidate who disagrees with an examiner is not necessarily wrong.

The history of science contains plenty of cases where examiners themselves were wrong.

The Philip Eaton story is a useful reminder of this.

The Heisenberg story is another.


The truly dangerous viva

The most interesting cases are therefore not necessarily those in which the candidate says:

"I don't know."

The genuinely dangerous cases are those in which the candidate repeatedly demonstrates that they do not understand what they have done.

There is a profound difference between:

"I did not consider that possibility."

and:

"I don't understand why this analysis was done."

There is also a difference between:

"I don't remember the exact derivation."

and:

"I cannot explain what this parameter means."

A PhD oral should be able to distinguish those situations.

That is much more difficult—and much more meaningful—than simply trying to catch a candidate out.


What these famous disasters ultimately tell us

Heisenberg's examination tells us that a future revolutionary can have a glaring weakness.

Hestenes's failed oral tells us that failure can sometimes be a useful intervention rather than a final judgement.

Oppenheimer's examination reminds us that intellectual intensity can be as challenging for an examiner as ignorance.

Muller's story reminds us that even Nobel laureates once stood outside examination rooms terrified.

Huxley's story shows that important scientific ideas can emerge while a scientist is still in training.

Brenner's story reminds examiners that their own emotions can influence judgement.

And Eaton's chemistry examination reminds us of something even more fundamental:

The examiner is not necessarily right.

Science is not a ritual in which an all-knowing professor interrogates an ignorant student.

At its best, a PhD oral is a meeting between scientists in which one side is testing whether the other has genuinely crossed the threshold from student to independent researcher.

Sometimes that process goes badly.

Sometimes spectacularly badly.

Sometimes the candidate deserves to fail.

Sometimes the examiner gets it wrong.

And occasionally, as in Heisenberg's case, a candidate can leave the examination room with the lowest passing grade—and go on to transform an entire field of science.

The most comforting lesson for a nervous PhD student is therefore not:

"Don't worry. Even geniuses fail exams."

It is something more useful:

An oral examination can reveal what you know, what you don't know, and how you think under pressure. But it is not a crystal ball.

A bad viva can be a serious warning.

It can be a deserved failure.

It can be a temporary setback.

Or, occasionally, it can be one afternoon in which an extraordinarily uneven human being happens to be asked exactly the questions he is least equipped to answer.

And that, as Werner Heisenberg discovered, is not necessarily the end of the story.


Sources and further reading

  • American Physical Society — historical account of Werner Heisenberg's doctoral oral examination.
  • American Institute of Physics — The Sad Story of Heisenberg's Doctoral Oral Exam.
  • Cold Spring Harbor Laboratory Archives — historical material concerning Hermann J. Muller and Edmund Beecher Wilson.
  • Historical accounts and autobiographical writings of David Hestenes concerning his doctoral oral examination.
  • Biographical accounts of J. Robert Oppenheimer's time at Göttingen.
  • Historical accounts of Philip Eaton and the early development of NMR spectroscopy.

Note: Stories about famous scientists and examinations are particularly vulnerable to embellishment. Where an anecdote is not supported by a contemporary examination record or a reliable retrospective source, it should be treated as a reported story rather than as an exact transcript of what happened.