ResearchOS/Wiki

Outlier tests

Now and then a data point sits far from the rest and you wonder whether it is a genuine result or a mistake, a mispipette, a bubble, a transcription slip. An outlier test gives you a principled way to flag a point as statistically extreme. But flagging is the easy part. The honest and harder part is that statistical extremeness alone is never a good enough reason to delete data, and this page is as much about that caution as about the test.

What an outlier test does

An outlier test asks whether the most extreme point in your sample is farther from the others than you would expect from the normal spread of the data. The common one is the Grubbs test, which takes the single most extreme value, measures how many standard deviations it sits from the mean, and computes a p-value for whether a spread that large would arise by chance from a normal distribution. A small p-value flags that point as a statistical outlier.

When a lab reaches for it

You run an assay in triplicate and one well reads wildly off from the other two. Before you average, you want a defensible way to ask whether that well is an aberration rather than real biology. The Grubbs test is built for exactly this, a single suspected outlier in an otherwise roughly normal small sample.

How to read the result

The Data Hub reports, for each pass of the sweep, the suspected value, the Grubbs G statistic (how many standard deviations it lies from the mean of the remaining sample), the critical G for the sample size and alpha, and whether the value is flagged. The engine runs iteratively by default: after flagging and removing the most extreme point, it recomputes the mean and standard deviation on what remains and tests again, repeating until no point is flagged or the sample drops below 3. Each pass is shown in the result table in order, so you can see which values were removed and which pass cleared the bar. This iterative sweep is the standard approach for removing more than one potential outlier, and it uses Bonferroni-corrected critical values so that testing multiple points does not inflate the false-flagging rate. A cluster of extreme points that all flag may mean the data simply are not normally distributed rather than that several replicates are errors.

Grubbs runs an iterative sweep by default. The table shows each pass: the most extreme value in that pass, its G statistic, and whether G exceeded the Bonferroni-corrected critical value. The sweep stops when a pass clears or the sample is too small to test.

The honest rule for removing data

The defensible reason to drop a point is a documented problem with that measurement, not its value. A note that the well had a visible bubble, an instrument error logged at that timestamp, a sample you know was compromised. That reason should exist independently of how the number came out, and it should be recorded. The right habit is to decide your handling rule before you see the data where you can, report that you removed a point and why, and ideally show the result both with and without it so a reader can judge. A genuinely surprising point can be the most important thing in the dataset.

A worked example

Triplicate readings come back as 4.9, 5.1, and 9.8. Grubbs flags 9.8 (G = 1.15, p = 0.03). On its own that is not permission to delete it. You check the run log, find that well was flagged for a pipetting error, and on that documented basis you exclude it, reporting "one of three replicates was excluded due to a logged pipetting error (Grubbs p = 0.03); results are shown for the remaining two." Without that logged reason, you would keep the point and report the spread honestly.

ResearchOS validates the Grubbs test against scipy and R on the transparency page.