ResearchOS/Wiki

Repeated measures, mixed models, and nested designs

Ordinary ANOVA assumes every data point is an independent sample. Real experiments often break that assumption. You measure the same animal before and after, or you read three wells from the same dish, or you count cells within mice within litters. These designs share structure that you have to account for, otherwise you fool yourself into thinking you have far more independent data than you really do. This page covers the three common cases.

The problem these designs solve

Independence is the quiet assumption behind most simple tests. Two measurements are independent when knowing one tells you nothing about the other. The moment they share a source, a subject, a dish, an animal, they are correlated, and treating them as independent makes your sample look bigger and your p-values look smaller than they should. The fix is not to throw data away. It is to use a method that knows about the structure.

Repeated-measures ANOVA

Use a repeated-measures ANOVA when the same subjects are measured under every condition. A classic case is a crossover, the same ten patients receive drug and placebo in turn, or the same cell line is read at four timepoints. Because each subject acts as its own control, the analysis can subtract out the steady differences between subjects (one patient just runs high across the board) and look only at how each subject moves across conditions. That makes the test more sensitive than treating the measurements as unrelated groups.

The result reports an F statistic and p-value for the within-subject factor, plus a partial eta-squared effect size. Partial eta-squared is the fraction of the within-subject variance that the condition factor explains, computed as the condition sum of squares divided by the sum of condition and error sums of squares (the subject-to-subject baseline is excluded from the denominator). A partial eta-squared of 0.30 means the condition accounts for 30% of the within-subject variation, the part you can actually change by manipulating the condition.

Sphericity and its corrections

Repeated-measures ANOVA relies on an assumption called sphericity: the variances of the differences between every pair of conditions should be equal. When you have only two conditions the assumption is automatically met. With three or more, it can fail, and when it does the standard F test's p-value is too small, giving more false positives than it should.

The Data Hub reports two corrected p-values alongside the standard one. The Greenhouse-Geisser correction adjusts the degrees of freedom by a factor epsilon (between 1/(k minus 1) and 1) estimated from the covariance matrix of the conditions. The smaller epsilon is, the worse the sphericity violation. The Huynh-Feldt correction uses a slightly less conservative epsilon that corrects the downward bias in the Greenhouse-Geisser estimate. Both corrected p-values are reported with their epsilon. When the standard and corrected p-values agree you have little to worry about; when they diverge, report the corrected one. If either epsilon is well below 1, a nonparametric Friedman test is worth considering as an alternative.

Three timepoints measured on the same subjects. The table splits the variation into the condition effect, the subject-to-subject differences, and the leftover error, and F is tested against that within-subject error. Partial eta-squared below it captures how much of the within-subject spread the condition explains. The Greenhouse-Geisser and Huynh-Feldt corrections adjust the p-value when sphericity is in doubt.

Mixed models

A mixed model is the flexible generalization. The name comes from mixing two kinds of effect. A fixed effect is the thing you care about and chose deliberately, the drug, the genotype, the dose. A random effect is a source of variation you are sampling from rather than studying, the particular animals, the particular plates, the particular days. The model estimates your fixed effect while explicitly accounting for the wobble each random effect adds.

Mixed models earn their keep when the design is unbalanced or has gaps, a patient missed a visit, a well failed. Repeated-measures ANOVA gets awkward with missing cells; a mixed model handles them gracefully because it works from the data you have rather than requiring a perfect grid. The result reports, for each fixed effect, an estimate with its standard error, confidence interval, and p-value. It also reports, for each random effect, the estimated variance component, which tells you how much of the total spread each random source (animals, plates, days) is responsible for. A large random-effect variance on "animal" means animals differ a lot from each other, and collecting more replicates within each animal will not help much; you need more animals.

Nested designs and the replicate trap

This is the one that quietly invalidates a lot of published work, so it is worth being plain about. A nested design is when your observations sit inside a hierarchy. You image 30 cells, but those cells come from 3 mice, 10 cells each.

The honest analysis respects the nesting. Either you summarize each animal to one number and compare animals, or you use a nested t-test or nested ANOVA (a mixed model with animal as a random effect) that pools the cell-level data while still counting animals as the unit of replication. The Data Hub frames the verdict for a nested test at the biological-replicate level on purpose, so the conclusion is about your mice, not your microscope.

The nested result reports the effect estimated at the group level (the difference between conditions across animals), its confidence interval, and a p-value based on the number of animals. The within-animal spread is used to weight the estimate, not to inflate the sample size.

A worked example

You treat 4 mice with a drug and 4 with vehicle, imaging 25 cells per mouse. A naive t-test on all 200 cells gives p < 0.0001, which looks spectacular and is wrong. The nested analysis, counting mice as the unit, returns a mean difference of 8% (95% CI -2 to 18, p = 0.10). You would report the nested result and say the effect is suggestive but not statistically clear with four animals per arm, which is the truth. The first p-value was an artifact of pretending 200 cells were 200 mice.

ResearchOS validates repeated-measures ANOVA (including sphericity corrections), the mixed-model fits, and the nested tests against statsmodels, pingouin, and R on the transparency page.