Sunday, September 27, 2026

A Negative Pathway Result That Strengthens the Evidence for GPRC6A Protein in Cow

In biology, a negative result can sometimes be surprisingly valuable. A paper may test a gene, decide it is not part of the pathway under study, and in doing so provide unusually credible evidence that the gene’s protein product exists in the experimental system. That is exactly the case for GPRC6A in the bovine mammary epithelial cell paper:

“Taurine Promotes Milk Synthesis via the GPR87-PI3K-SETD1A Signaling in BMECs.”

The paper was authored by Mengmeng Yu, Yang Wang, Zhe Wang, Yanxu Liu, Yang Yu, and Xuejun Gao, with affiliations at Agricultural College of Guangdong Ocean University and The Key Laboratory of Dairy Science of Education Ministry, Northeast Agricultural University. It was published in the Journal of Agricultural and Food Chemistry in 2019, volume 67, pages 1927 to 1936, DOI 10.1021/acs.jafc.8b06532.

The headline conclusion of the paper is not about GPRC6A. The authors conclude that taurine promotes milk synthesis through GPR87-PI3K-SETD1A signaling. In the abstract, they state that gene-function approaches revealed GPR87-PI3K-SETD1A signaling was required for taurine to increase mTOR and SREBP-1c mRNA levels, and that taurine stimulated GPR87 expression and membrane localization.

But this is precisely why the GPRC6A result is interesting.

The authors did not build their story around GPRC6A. They tested GPRC6A as a plausible candidate receptor, knocked it down, found that taurine signaling still occurred, and then moved the pathway to GPR87. That makes the GPRC6A Western blot evidence less likely to be a pathway-confirmation artifact. The GPRC6A band was not needed to sell the final mechanism. In fact, the final mechanism explicitly excludes GPRC6A from taurine signaling.

Why GPRC6A was tested at all

The authors had a strong biological reason to consider GPRC6A. In the introduction, they explain that some GPCRs can sense extracellular amino acids and activate downstream signaling such as PI3K/mTOR. They specifically name GPRC6A as one of these amino-acid-sensing GPCRs.

They also state in the Results that previous mass-spectrometric data showed both GPRC6A and GPR87 were upregulated in BMECs treated with methionine. Therefore, they hypothesized that GPRC6A, GPR87, or both might be required for taurine-induced PI3K activation.

This gives the experiment a clean logic:

Taurine activates PI3K. GPRC6A and GPR87 are candidate GPCRs. Knock each down. See which one matters.

The key assay: Western blot detection of GPRC6A protein

The paper’s Methods section lists a specific antibody for GPRC6A detection: GPRC6A antibody ab138994 from Abcam. The same Western blot workflow also used antibodies against GPR87, SETD1A, H3K4Me3, PI3K pathway proteins, mTOR, SREBP-1c, and β-actin. The signals were visualized by chemiluminescence, quantified in ImageJ, and normalized to β-actin or histone H3.

This matters because the paper is not only mentioning GPRC6A in text. It directly measures a GPRC6A protein band by Western blot.

The strongest GPRC6A evidence: siRNA knockdown in Figure 7A

The crucial figure is Figure 7A.

In Figure 7A, BMECs were transfected with GPRC6A siRNA and treated with 0.24 mM taurine for 24 hours. The figure caption explicitly states that cells were transfected with GPRC6A siRNA and then analyzed by Western blot.

The Methods section gives the exact GPRC6A siRNA sequence used:

GPRC6A-siRNA: 5′-GCUCUGAGGUGUGUUUCUATT-3′.

This is important because the experiment is not simply “we saw a band and named it GPRC6A.” The authors used a targeted knockdown reagent against GPRC6A and then observed the Western blot signal under knockdown conditions.

That provides two layers of evidence:

  1. Baseline detection: a GPRC6A protein band is detectable in bovine mammary epithelial cells.
  2. Knockdown validation: the GPRC6A band is reduced after GPRC6A siRNA treatment.

This makes the protein-level evidence much stronger than a standalone antibody blot. A Western blot band that decreases after gene-specific siRNA behaves like the intended protein signal. It is not perfect proof, but it is one of the more persuasive practical validations used in cell-biology papers.

What Figure 7A actually shows

Figure 7A tests whether GPRC6A is required for taurine-induced PI3K activation.

The result is beautifully paradoxical for our purpose.

The authors report that GPRC6A knockdown did not suppress PI3K activation after taurine stimulation. In their interpretation, this means GPRC6A is not required for taurine to activate PI3K.

So Figure 7A says two things at once:

First: GPRC6A protein is detectable and knockdown-able in bovine BMECs.
Second: GPRC6A is not the receptor responsible for taurine-to-PI3K signaling.

That distinction is the whole treasure chest.

If the authors wanted to force a GPRC6A pathway story, Figure 7A would have been inconvenient. Instead, they used it to eliminate GPRC6A and support GPR87.

Why the negative result makes the GPRC6A blot more credible

One has to be careful here: we cannot know the authors’ intentions. But we can evaluate the evidentiary structure.

The GPRC6A Western blot is not being used to claim that GPRC6A mediates taurine signaling. The paper’s main pathway is GPR87, not GPRC6A. The Results section states that GPRC6A knockdown did not block taurine-induced PI3K activation, while GPR87 knockdown largely abolished taurine effects on p-PI3K, p-mTOR, and SREBP-1c.

The Discussion says this even more directly: “GPR87 but not GPRC6A” knockdown abolished taurine’s stimulatory effects on PI3K and downstream signaling.

That makes the GPRC6A detection valuable in a special way. The authors had no pathway-level incentive to exaggerate GPRC6A as functional in taurine signaling, because their conclusion goes the other way. Yet they still show a GPRC6A Western blot band and use GPRC6A siRNA in Figure 7A.

In other words, the GPRC6A signal is not decorative confetti thrown over the main claim. It is part of a receptor-exclusion experiment.

The contrast with GPR87 sharpens the argument

The paper’s positive receptor is GPR87.

In Figure 7B to 7F, GPR87 knockdown reduces GPR87 protein and blocks taurine-induced p-PI3K, p-mTOR, and SREBP-1c responses. The figure caption describes the GPR87 knockdown Western blots and quantification of GPR87, p-PI3K/PI3K, p-mTOR/mTOR, and SREBP-1c.

That contrast matters.

The authors did not merely say, “GPRC6A exists.” They ran a comparative receptor screen:

GPRC6A knockdown: taurine signaling survives.
GPR87 knockdown: taurine signaling collapses.

This makes the GPRC6A experiment a clean negative control against the GPR87 result.

But for protein existence, the GPRC6A portion remains valuable because it demonstrates that the GPRC6A protein signal was measurable and experimentally reducible in bovine cells.

What this paper can honestly be used to claim

This paper should not be cited as evidence that GPRC6A mediates taurine-induced milk synthesis. It says the opposite.

But it can be cited as strong evidence for this narrower claim:

Bovine mammary epithelial cells contain a detectable GPRC6A protein signal by Western blot, and this signal is reduced by GPRC6A-targeting siRNA.

That is a meaningful protein-level validation.

The strongest honest wording would be:

Yu et al. tested GPRC6A as a candidate taurine receptor in bovine mammary epithelial cells. Although GPRC6A knockdown did not block taurine-induced PI3K activation, Figure 7A provides protein-level evidence that GPRC6A is detectable by Western blot in BMECs and that the detected signal is responsive to GPRC6A siRNA knockdown. Thus, the paper is negative evidence for GPRC6A in taurine signaling, but positive evidence for the existence of GPRC6A protein in cow mammary epithelial cells.

Why this matters for bovine GPRC6A functionality

Functional annotation often asks several different questions:

  1. Is the gene present in the genome?
  2. Is the transcript expressed?
  3. Is the protein product detectable?
  4. Is the protein part of a biological pathway?
  5. Is it required for a particular phenotype?

This taurine paper helps mainly with question 3.

It does not show that GPRC6A mediates taurine signaling. It does not establish direct ligand binding. It does not prove a GPRC6A-dependent taurine phenotype. But it does show a GPRC6A protein band in bovine BMECs and a targeted knockdown experiment that reduces the band.

For the broader cow GPRC6A argument, this paper should be used as a brick, not the whole barn.

Together with other bovine papers where GPRC6A knockdown blocks lysine or palmitic-acid signaling, this taurine paper adds an independent piece of evidence: even in a pathway where GPRC6A is ruled out, the protein is still detected and experimentally manipulated.

That is why this paper is useful. Not because it makes GPRC6A the taurine receptor, but because it shows that GPRC6A was present enough, measurable enough, and knockdown-able enough to be tested and rejected.

Final interpretation

The most honest conclusion is:

This paper provides strong evidence for the existence of GPRC6A protein in bovine mammary epithelial cells, based on Western blot detection and siRNA knockdown validation in Figure 7A. However, it does not support GPRC6A as part of the taurine-induced milk-synthesis pathway. Instead, the authors use the GPRC6A knockdown result to exclude it and identify GPR87 as the functional taurine-responsive receptor.

That is not a weakness. It is exactly why the GPRC6A protein evidence is persuasive.

The paper clears GPRC6A from the taurine pathway, but in doing so, it leaves behind a useful fingerprint: GPRC6A protein exists in cow mammary epithelial cells and can be detected by Western blot. 🐄

Saturday, September 26, 2026

Variance and Standard Deviation: Why Squaring Won

Variance occupies a privileged position in statistics.

Why?

Not because squaring deviations is intuitively inevitable.

It is because squared deviations have extraordinarily convenient mathematics.


1. Population variance

For a random variable (X),

E[(X-\mu)^2].
]

Expanding,

X^2-2\mu X+\mu^2.
]

Taking expectations,

E[X^2]-2\mu E[X]+\mu^2.
]

Since

[
E[X]=\mu,
]

we obtain

[
\boxed{
\operatorname{Var}(X)=E[X^2]-E[X]^2
}
]

This identity makes variance enormously convenient computationally and theoretically.


2. Why deviations are measured around the mean

Consider

[
L(c)=\sum_i(x_i-c)^2.
]

Differentiate:

-2\sum_i(x_i-c).
]

Set this equal to zero:

[
\sum_i(x_i-c)=0.
]

Therefore,

[
nc=\sum_i x_i
]

and hence

[
c=\bar{x}.
]

So the arithmetic mean is precisely the value that minimizes total squared deviation.

By contrast, the median minimizes total absolute deviation.

This gives us a beautiful duality:

[
L_2 \text{ loss} \rightarrow \text{mean}
]

[
L_1 \text{ loss} \rightarrow \text{median}.
]


3. Sample variance and (n-1)

For a sample,

\frac{1}{n-1}
\sum_i(x_i-\bar{x})^2.
]

Why (n-1)?

Because after estimating the sample mean, only (n-1) deviations are free.

Indeed,

[
\sum_i(x_i-\bar{x})=0.
]

Once (n-1) deviations are known, the last is determined.

More formally,

\frac{n-1}{n}\sigma^2.
]

Dividing by (n-1) instead removes this downward bias for estimating population variance.


4. Why standard deviation exists

Variance has squared units.

If height is measured in centimeters,

[
\operatorname{Var}(X)
]

has units of

[
\text{cm}^2.
]

Taking the square root produces

[
SD=\sqrt{\operatorname{Var}(X)},
]

which returns us to centimeters.

That makes standard deviation substantially easier to communicate.


5. Python demonstration

import numpy as np

x = np.array([4, 5, 5, 6, 10])

print("Population variance:",
      np.var(x, ddof=0))

print("Sample variance:",
      np.var(x, ddof=1))

print("Sample SD:",
      np.std(x, ddof=1))

Verify the alternative variance formula:

var1 = np.mean(
    (x - np.mean(x))**2
)

var2 = np.mean(x**2) - np.mean(x)**2

print(var1, var2)

6. R demonstration

x <- c(4, 5, 5, 6, 10)

# R's var() uses denominator n - 1
var(x)
sd(x)

# Population variance
mean((x - mean(x))^2)

# Computational identity
mean(x^2) - mean(x)^2

7. Variances add

One of variance's greatest mathematical advantages is:

\operatorname{Var}(X)
+\operatorname{Var}(Y)
+2\operatorname{Cov}(X,Y).
]

If (X) and (Y) are independent,

[
\operatorname{Cov}(X,Y)=0
]

and therefore

\operatorname{Var}(X)
+
\operatorname{Var}(Y).
]

Standard deviations do not enjoy such a clean decomposition.

This property helped make variance foundational in experimental design, quantitative genetics, measurement theory, and signal processing.

Fisher's 1918 work made variance decomposition central to statistical thinking.


8. Variance's Achilles heel

Variance squares deviations.

Suppose our data are:

[
1,2,3,4,5.
]

Now replace 5 with 100.

import numpy as np

a = np.array([1,2,3,4,5])
b = np.array([1,2,3,4,100])

print(np.var(a, ddof=1))
print(np.var(b, ddof=1))

The variance explodes.

In fact, classical variance has a breakdown point effectively approaching zero: one sufficiently extreme observation can make it arbitrarily large.

This motivates an entirely different branch of statistics.

Robust statistics.

Friday, September 25, 2026

Range, IQR, and Absolute Deviations

Before reaching variance, it is worth exploring measures that are often simpler and sometimes more appropriate.


1. Range

The range is

[
R=x_{\max}-x_{\min}.
]

For

[
2,;3,;4,;5,;6
]

the range is

[
6-2=4.
]

Its greatest advantage is interpretability.

Its greatest weakness is equally obvious.

Only two observations determine it.

A dataset containing one million observations has its range determined entirely by its minimum and maximum.


2. Interquartile range

The interquartile range is

[
IQR=Q_3-Q_1.
]

It describes the width occupied by the middle half of the observations.

Because observations below (Q_1) and above (Q_3) do not directly affect the endpoints, it is much less sensitive to extremes than the range or standard deviation.

That makes the IQR especially useful for skewed distributions.

It is also the machinery behind the familiar boxplot.


3. Mean absolute deviation

Instead of squaring deviations, why not simply take their absolute values?

Around the mean:

\frac1n
\sum_{i=1}^{n}|x_i-\bar x|.
]

Or around the median:

\frac1n
\sum_{i=1}^{n}|x_i-\tilde x|.
]

These should not be confused with the median absolute deviation, which we will encounter later.

Absolute deviations have a useful property: extreme observations grow linearly rather than quadratically in influence.

Consider deviations of 2 and 20.

Under absolute loss:

[
2 \rightarrow 2,\qquad 20\rightarrow20.
]

Under squared loss:

[
2\rightarrow4,\qquad20\rightarrow400.
]

Squaring turns the second observation into a statistical megaphone.


4. Python comparison

import numpy as np
import pandas as pd

x = np.array([10, 11, 12, 13, 14, 15, 50])

mean = np.mean(x)
median = np.median(x)

results = {
    "Range": np.ptp(x),
    "IQR": np.percentile(x, 75)
           - np.percentile(x, 25),
    "Mean abs dev about mean":
        np.mean(np.abs(x - mean)),
    "Mean abs dev about median":
        np.mean(np.abs(x - median)),
    "SD": np.std(x, ddof=1)
}

print(pd.Series(results))

Now progressively increase the outlier.

outliers = np.arange(15, 101, 5)

rows = []

for o in outliers:
    z = np.array([10, 11, 12, 13, 14, 15, o])

    rows.append({
        "outlier": o,
        "range": np.ptp(z),
        "IQR": np.percentile(z, 75)
               - np.percentile(z, 25),
        "SD": np.std(z, ddof=1),
        "AAD": np.mean(
            np.abs(z - np.mean(z))
        )
    })

df = pd.DataFrame(rows)
print(df)

Plot the trajectories:

import matplotlib.pyplot as plt

plt.plot(df["outlier"], df["range"],
         label="Range")
plt.plot(df["outlier"], df["SD"],
         label="SD")
plt.plot(df["outlier"], df["IQR"],
         label="IQR")
plt.plot(df["outlier"], df["AAD"],
         label="Absolute deviation")

plt.xlabel("Extreme observation")
plt.ylabel("Dispersion")
plt.legend()
plt.show()

This plot is wonderfully revealing.

Range grows directly with the extreme observation.

SD grows rapidly.

Absolute deviation responds more gently.

IQR barely notices.


5. R version

x <- c(10, 11, 12, 13, 14, 15, 50)

aad_mean <- mean(abs(x - mean(x)))
aad_median <- mean(abs(x - median(x)))

c(
  Range = diff(range(x)),
  IQR = IQR(x),
  AAD_mean = aad_mean,
  AAD_median = aad_median,
  SD = sd(x)
)

Simulation:

outliers <- seq(15, 100, by = 5)

result <- data.frame(
  outlier = outliers,
  range = NA,
  IQR = NA,
  SD = NA,
  AAD = NA
)

for (i in seq_along(outliers)) {
  z <- c(10, 11, 12, 13, 14, 15,
         outliers[i])

  result$range[i] <- diff(range(z))
  result$IQR[i] <- IQR(z)
  result$SD[i] <- sd(z)
  result$AAD[i] <- mean(abs(z - mean(z)))
}

matplot(
  result$outlier,
  result[, c("range", "IQR", "SD", "AAD")],
  type = "l",
  lty = 1,
  xlab = "Extreme observation",
  ylab = "Dispersion"
)

legend(
  "topleft",
  legend = c("Range", "IQR", "SD", "AAD"),
  lty = 1,
  col = 1:4
)

6. Advantages and limitations

MeasureStrengthMain limitation
RangeExtremely intuitiveDetermined by two observations
IQRRobust and easy to interpretIgnores much tail information
Mean absolute deviationSame units and moderate tail sensitivityLess algebraically convenient
SDRich mathematical theoryHighly sensitive to tails

A useful habit is to report more than one.

For skewed or contamination-prone data, reporting

[
\text{median + IQR}
]

may be far more informative than

[
\text{mean + SD}.
]

Wednesday, September 23, 2026

What Does "Spread" Actually Mean?

 

1. Location is only half the story

Suppose two laboratories measure the same quantity.

Laboratory A reports:

[
10,;10,;10,;10,;10
]

Laboratory B reports:

[
2,;6,;10,;14,;18
]

Both have mean 10.

Yet describing them as statistically equivalent would clearly be absurd.

A location statistic tells us where the distribution sits.

A dispersion statistic tells us how broadly the observations occupy the space around that location.

This distinction appears everywhere:

  • biology: variability in gene expression,
  • ecology: variability in species abundance,
  • manufacturing: process consistency,
  • finance: volatility,
  • medicine: heterogeneity of patient responses,
  • genomics: read depth variability,
  • machine learning: variability of representations or prediction uncertainty.

Sometimes variability is noise.

Sometimes variability is the biological phenomenon.

That distinction matters enormously.


2. Four different ideas of dispersion

Consider:

[
x=(1,2,3,4,20).
]

There are several reasonable ways to describe its spread.

Extreme separation

[
\text{Range}=\max(x)-\min(x)
]

This asks:

How far apart are the most extreme observations?

Central spread

[
IQR=Q_{0.75}-Q_{0.25}
]

This asks:

How wide is the middle 50%?

Typical deviation from a center

For example,

[
\frac{1}{n}\sum |x_i-\bar{x}|
]

or

[
\operatorname{median}|x_i-\operatorname{median}(x)|.
]

These ask:

How far does a typical observation lie from some center?

Pairwise spread

We can instead ask how far apart observations are from one another:

[
\frac{1}{\binom n2}
\sum_{i<j}|x_i-x_j|.
]

This idea leads to Gini's mean difference.

None of these questions is inherently more correct than the others.

They simply measure different geometries of variability.


3. A short historical detour

The modern vocabulary emerged gradually.

Karl Pearson introduced the term standard deviation in lectures in 1893 and used it in print in 1894, replacing older terminology such as "mean error" and "error of mean square."

R. A. Fisher introduced the statistical term variance in his famous 1918 work on the resemblance between relatives, where variation could be decomposed into meaningful components.

Corrado Gini introduced his mean-difference approach to variability in 1912, providing an alternative family of ideas based on absolute pairwise differences rather than squared deviations.

So even historically, variance was never the only road through the forest.


4. A first experiment in Python

import numpy as np

x = np.array([1, 2, 3, 4, 20])

mean = np.mean(x)
median = np.median(x)
data_range = np.ptp(x)
variance = np.var(x, ddof=1)
sd = np.std(x, ddof=1)
iqr = np.percentile(x, 75) - np.percentile(x, 25)
mad_raw = np.median(np.abs(x - median))

print("Mean:", mean)
print("Median:", median)
print("Range:", data_range)
print("Variance:", variance)
print("SD:", sd)
print("IQR:", iqr)
print("MAD:", mad_raw)

Now remove the extreme observation:

y = np.array([1, 2, 3, 4])

for name, z in [("with outlier", x),
                ("without outlier", y)]:
    print("\n", name)
    print("SD =", np.std(z, ddof=1))
    print("IQR =", np.percentile(z, 75) -
                   np.percentile(z, 25))
    print("MAD =", np.median(
        np.abs(z - np.median(z))
    ))

The measures react very differently.

That reaction is not a bug. It tells us what each measure cares about.


5. The same experiment in R

x <- c(1, 2, 3, 4, 20)

mean(x)
median(x)
diff(range(x))
var(x)
sd(x)
IQR(x)

mad_raw <- median(abs(x - median(x)))
mad_raw

Compare with:

y <- c(1, 2, 3, 4)

metrics <- function(x) {
  c(
    SD = sd(x),
    IQR = IQR(x),
    MAD_raw = median(abs(x - median(x)))
  )
}

metrics(x)
metrics(y)

6. Properties we should demand from dispersion measures

A useful dispersion measure might possess several properties.

Non-negativity

[
D(X)\geq0.
]

Zero for constant data

If every observation is identical,

[
D(X)=0.
]

Translation invariance

Adding a constant should usually not change spread:

[
D(X+c)=D(X).
]

Variance, SD, IQR and MAD all satisfy this.

Scale equivariance

Multiplying the data by (a) should change a scale measure proportionally:

[
D(aX)=|a|D(X).
]

SD and MAD satisfy this.

Variance instead satisfies:

[
\operatorname{Var}(aX)=a^2\operatorname{Var}(X).
]

Robustness

How much can one pathological observation alter the answer?

This turns out to be one of the central questions in the entire series.


7. A crucial lesson

There is no universally best dispersion measure.

Choosing one involves deciding what kind of variation deserves influence.

Variance says:

Large deviations deserve disproportionately large influence.

MAD says:

The behavior of the majority matters more than extreme observations.

Range says:

I care specifically about the extremes.

IQR says:

I care about the central half.

Gini mean difference says:

I care about distances among all pairs.

Those are scientific choices disguised as formulas.

And that is why dispersion deserves more thought than simply typing sd(x).

What Should Science Learn From the Career Effects of Retractions?

Retractions are necessary.

That should be the starting point.

Scientific knowledge is valuable partly because science contains mechanisms for identifying and correcting unreliable claims. A literature in which papers can never be withdrawn would not be more trustworthy. It would be less trustworthy.

But the Nature Human Behaviour study shows that retraction systems do more than modify the literature.

They affect people.

The study's findings can be summarized as a sequence.

Retraction is associated with earlier departure from scientific publishing.

The effect appears particularly concerning for researchers with less-established careers.

Greater public attention surrounding a retraction is associated with a wider attrition gap.

Among researchers who remain, collaboration networks often grow rather than shrink.

Yet those networks change in composition, with important differences in collaborator seniority, productivity and impact.

What should institutions do with this information?

First, distinguish correction from culpability.

A retraction tells us something went wrong with a publication. It does not necessarily tell us that every author committed misconduct.

Second, make retraction notices more informative.

Readers should be able to distinguish honest error, plagiarism, fabrication, methodological failure and author-initiated correction whenever the evidence allows such distinctions.

Third, pay particular attention to junior researchers.

Because early-career scientists possess less accumulated reputational capital, institutions and mentors may need procedures ensuring that involvement in a retracted paper is evaluated according to actual contribution and responsibility.

Fourth, rethink how self-correction is rewarded.

If scientists believe voluntarily retracting erroneous work will permanently damage their careers, the system creates incentives to defend questionable results rather than correct them.

The authors themselves identify self-retraction, scientific-community support and changes in collaboration strategies as important mechanisms that future studies should examine.

Fifth, study the role of publicity.

A correction that receives almost no public attention and one that becomes an international scandal may have radically different career consequences. Future work needs to distinguish attention from condemnation and scientific discussion from personal exposure.

Finally, research integrity should be evaluated as a system rather than as a collection of individual retraction events.

The ideal system has to accomplish two things simultaneously:

correct science aggressively and assign responsibility accurately.

Those goals are not in conflict.

Indeed, both are necessary for a culture in which researchers are willing to acknowledge mistakes while deliberate misconduct remains consequential.

The deeper message of this study is therefore not that retractions are too harsh or too lenient.

It is that retraction is a much more powerful institutional intervention than simply placing a warning label on a PDF.

A retraction changes the scientific record.

It changes how researchers see one another.

It can reshape collaboration networks.

And, for some scientists, it can mark the point at which a publishing career ends.

Understanding those consequences is essential if science wants its mechanisms of self-correction to be both rigorous and fair.

Tuesday, September 22, 2026

Beyond the Standard Deviation

 

A Practical Series on Measuring Variability, Spread, and Statistical Dispersion

Most introductory statistics courses teach a familiar sequence:

mean → variance → standard deviation.

That sequence is useful, but it can accidentally suggest that once we know the standard deviation, the problem of measuring variability has been solved.

It has not.

There are many legitimate meanings of "spread":

  • How far apart are the extremes?
  • How wide is the central half of the data?
  • How far is a typical observation from the center?
  • How different are two randomly selected observations?
  • How variable is the quantity relative to its magnitude?
  • How much of the spread is caused by rare observations?
  • How dispersed is a multidimensional cloud?
  • What does dispersion even mean for angles, compositions, probability distributions, networks, images, or embeddings?

Different measures answer different questions.

This series explores those questions mathematically, historically, computationally, and practically.

The posts are:

  1. What Does "Spread" Actually Mean?
  2. Range, IQR, and Absolute Deviations
  3. Variance and Standard Deviation: Why Squaring Won
  4. Relative Dispersion: CV, Fano Factor, and Scale-Free Measures
  5. Robust Dispersion: MAD, Qn, Sn, and Gini Mean Difference
  6. Comparing Dispersion Between Groups
  7. Multivariate and High-Dimensional Dispersion
  8. When Ordinary Variance Stops Making Sense
  9. Where Dispersion Research Could Go Next

Are Retraction Systems Fair to Co-Authors?

A paper may have one author.

It may also have fifty.

Yet when that paper is retracted, every author becomes permanently associated with the word “retracted.”

This creates a difficult fairness problem.

Authorship is collective.

Responsibility is often not.

Consider a hypothetical paper containing fabricated microscopy images.

The researcher who created those images may be directly responsible.

Another researcher may have performed an unrelated computational analysis.

A junior student may have contributed samples.

A senior principal investigator may have supervised the project.

All appear on the same paper.

Should they experience the same reputational consequences?

The study cannot determine the precise culpability of every author, and the authors explicitly acknowledge this limitation. Researchers associated with retracted papers differ in their awareness of the problems that ultimately produced the retraction, and mentors or colleagues may respond differently depending on those circumstances. The paper calls for future research capable of distinguishing authors according to their involvement in the reasons behind retraction.

This is more than a methodological problem.

It is an institutional-design problem.

Retraction notices often function as both corrections to the scientific literature and reputational documents.

Those two roles should perhaps be separated more clearly.

A good correction notice could explain:

what part of the paper is unreliable;

why it is unreliable;

whether misconduct was established;

whether the retraction was initiated by authors or the journal;

which contributions were implicated;

and, where investigations have established it, which individuals bear responsibility.

There are obvious legal and procedural challenges. Journals cannot simply accuse individual authors without adequate evidence.

But ambiguity also has consequences.

When notices provide too little information, readers may infer collective culpability.

This is particularly concerning given the paper's finding that less-established researchers appear more vulnerable to career exit after retraction.

Research-integrity systems therefore face two simultaneous obligations.

First, protect the scientific record.

Second, avoid converting the correction of a publication into indiscriminate punishment of everyone associated with it.

Those goals are compatible.

In fact, greater specificity in retraction notices could strengthen both.

Clearer explanations would help readers understand why a result should no longer be trusted while allowing the scientific community to distinguish error, negligence and deliberate misconduct.

A mature research-integrity system should be capable of saying:

“This paper is unreliable.”

without automatically implying:

“Every scientist whose name appears on it is unreliable.”

That distinction may be one of the most important lessons to emerge from this research.