Group comparison
Repeated Measures ANOVA
Compares the means of three or more continuous measurements taken from the same participants in a single test.
Method summary
Repeated measures ANOVA is for designs where one measurement is repeated on the same people: a panel study with three waves, interim tests across a training programme, or several products each participant rates in turn. Because the measurements come from the same people they are not independent, and stable differences between individuals are kept as a separate source in the model. That shrinks the error term against which change is tested and makes the analysis more sensitive than one based on the same number of independent observations. The output is a single F value telling you whether at least one difference exists among the measurement occasions. YouReply Analyze accepts either layout: in wide format you tick the measurement columns, in long format you point to the subject identifier and the within-subject factor yourself.
Which research questions does it answer?
- Does brand trust shift between the wave before an advertising campaign, one month after it and three months after it?
- Does satisfaction with a service stay stable across four quarters among the same panel members?
- Do the ratings each participant gives to three packaging prototypes tried in sequence differ from one another?
- Does knowledge rise across the entry, interim and exit measurements of a training programme?
When should you use it?
- When there are three or more measurement occasions; with two occasions the paired t-test answers the same question more directly.
- When every measurement comes from the same participants, or participants are arranged in matched blocks.
- When the measured value is continuous and its mean is interpretable.
- When each participant has a value at every occasion, since anyone missing a wave breaks the balanced design.
Required variable types
- Dependent variable: a continuous numeric measurement. One column per occasion in wide format, a single value column in long format.
- Within-subject factor: the categorical information saying which occasion or condition a measurement belongs to. In wide format this comes from the column names themselves.
- Subject identifier: in long format, the column tying one person's rows together. In wide format the engine treats the row position as the subject id, so a row must belong to exactly one participant.
Key assumptions
- Balanced, complete design
- Every participant needs a value at every occasion. Someone who misses a wave cannot contribute their remaining measurements, because the within-subject model treats a person as a whole.
- Independence between participants
- Dependence among one person's own measurements is handled by the model itself, but different participants must not influence each other. Samples drawn from the same household or the same work team strain this condition.
- Continuous measurements
- The F test is built on means and sums of squares, so measurement at interval or ratio level is expected. Reading a mean difference on a single ordinal item is contentious.
- Sphericity
- The variances of the difference scores between pairs of occasions are expected to be similar. When this is violated the F test returns significant results more often than it should, which the literature addresses with corrections that shrink the degrees of freedom.
- Distribution of the measurements
- The distribution at each occasion should not be severely skewed. On items with a pronounced floor or ceiling, difference scores pile up in one direction.
How YouReply checks these assumptions
- Balanced, complete design: You can declare which codes stand for missing values on the Variable tab. The panel does not fill in absent observations for you, so preparing a balanced table remains your work.
- Independence between participants: The panel does not check this assumption automatically; the researcher evaluates it.
- Continuous measurements: The method picker looks for numeric columns here, dims the method where a column does not fit and states the reason. The data requirements box lists which of your columns satisfy the condition.
- Sphericity: The panel does not check sphericity: Mauchly's test is not run and no Greenhouse-Geisser or Huynh-Feldt correction is applied. The F and p values you see are uncorrected, so as the number of occasions grows you need to justify this assumption yourself.
- Distribution of the measurements: Normality is not tested while this method runs. To inspect the distributions you can run the Normality tests method separately.
How the analysis is run
- 1Drop your data file onto the Data tab; once it loads you can go over the table in the editable grid on that same tab.
- 2On the Variable tab, confirm that each measurement column is numeric with the right measurement level, and declare your missing value codes there.
- 3Open the method list on the Analysis tab. With the 'Only ones that fit my data' switch on, unsuitable methods are not hidden but dimmed, with the reason written in plain language.
- 4Fill in the parameters according to your layout: tick the measurement columns in wide format, or select the value column together with the subject id and within-subject factor columns in long format.
- 5Run it. The F value, both degrees of freedom, the p value and partial eta squared appear in one section, with the descriptives for each occasion in another.
- 6A bar chart comparing the occasion means is drawn above the tables and can be saved as a PNG; the tables themselves export to Excel.
- 7The computation credits card names the library call and version behind the numbers, and the run is kept on the history tab.
Statistics and tables produced
- F statistic
- The ratio of variability between occasions to the residual within-person variability. The larger it grows, the harder it becomes to attribute the differences between occasions to chance.
- Two degrees of freedom
- The numerator degrees of freedom are one less than the number of occasions; the error (denominator) degrees of freedom derive from the number of participants together with the number of occasions. Both appear as separate fields in the result.
- Significance (p) value
- The probability of an F value at least this large if no occasion differed from any other. It is an uncorrected value, since no sphericity correction is applied.
- Partial eta squared
- An effect size computed from the F value and the degrees of freedom; it states how much of the residual within-person variability is associated with the measurement occasion.
- Descriptives per occasion
- Mean, median, mode, standard deviation and number of observations for every occasion. This is where you read the direction of change, as the F value says nothing about direction.
- Occasion means chart
- A bar chart placing the occasion means side by side. It is drawn above the result sections and downloads as a PNG.
Effect size and confidence intervals
- Partial eta squared
- The share of variability attributed to the occasion relative to that effect plus the residual variability. In the behavioural sciences the values quoted are 0.01 for a small, 0.06 for a medium and 0.14 for a large effect. Because the figure belongs to a within-subject model, it should not be compared directly with an eta squared from a between-groups ANOVA.
No confidence interval is computed in this method. None is returned for the occasion means, for the differences between them, or for partial eta squared; the descriptive table gives means and standard deviations, not standard errors or intervals. If your report needs intervals around the occasion means, you have to calculate them yourself.
Example research question and example result
The numbers below are a representative example, not data from a real study or a real user.
- Research question
- Representative example: in a panel study, brand trust was measured on a 1-7 scale from the same 92 people before a campaign, one month after it and three months after it. Does trust change across the occasions?
- Variables
- Dependent variable: brand trust score (1-7, a scale mean treated as continuous) · Within-subject factor: measurement occasion (before, month one, month three) · Layout: wide format, one column per occasion, one row per panel member
- Example result
- The means were 4.12 (SD = 1.02), 4.68 (SD = 0.95) and 4.55 (SD = 0.98) in order. The repeated measures ANOVA gave F(2, 182) = 11.46, p < .001, partial eta squared = 0.11.
- Interpretation
- At least one of the three occasions differs from another, and the effect size falls between the conventional medium and large benchmarks. The descriptives suggest the rise happened by month one and eased back slightly by month three, but that reading comes from the means and the chart rather than from a test. Which two occasions differ statistically cannot be settled from this output, since the method returns no post-hoc comparisons. And because the F value receives no sphericity correction, a design whose difference variances diverge noticeably may yield a p value smaller than it should be.
Real output on a sample dataset
The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.
General customer survey (synthetic)
A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.
- Rows
- 300
- Columns
- respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.
Research question: Do the means of the three repeated measurements differ over time?
- Analyzed columns
- measure_1,measure_2,measure_3
- Dependent variable
- measure_1
- Subject identifier column
- respondent_id
- Within-subject factor column
- measure_1
- F statistic
- 137.01
- p value
- < 0.001
- Degrees of freedom (numerator)
- 2
- Degrees of freedom (denominator)
- 598
- Partial eta squared
- 0.314
- dependent_var
- _dependent_var
- subject_id
- _subject_id
- within_factor
- _within_factor
Descriptive statistics
| Row | Mean | Median | Standard deviation | Count |
|---|---|---|---|---|
| measure_1 | 3.44 | 3.50 | 0.994 | 300 |
| measure_2 | 3.87 | 3.95 | 1.21 | 300 |
| measure_3 | 4.28 | 4.25 | 1.37 | 300 |
Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914
How to report the result
Brand trust changed significantly across the three occasions, F(2, 182) = 11.46, p < .001, partial η² = 0.11, with means of 4.12 (SD = 1.02) before the campaign, 4.68 (SD = 0.95) at month one and 4.55 (SD = 0.98) at month three.
An example sentence close to APA style; the numbers are representative.
When you should not use it
- It does not say which pair of occasions differs; with no post-hoc comparisons returned, you have to examine the pattern with separate tests.
- Since sphericity is neither tested nor corrected for, designs with diverging difference variances can carry a Type I error rate above the nominal level.
- Participants with missing observations break the balanced design, and the remaining measurements of someone who skipped a wave contribute nothing.
- When the order of the occasions is fixed, learning, fatigue and habituation cannot be separated from the effect of time. That is a limit of the design rather than of the test.
- On single ordinal items, or on scales with pronounced floor and ceiling effects, a mean difference can mislead.
- With no confidence interval returned, the output says nothing about how precisely the effect size is estimated.
What to use when the assumptions are not met
- Paired Samples t-TestWhy: When only two occasions are compared; it answers the same question and reports the size and direction of the mean difference.
- Friedman TestWhy: When the measurements are ordinal or the distributional assumptions are strained; it compares within-person ranks instead of means.
- One-Way ANOVAWhy: When the groups being compared consist of different people, so no within-subject structure is needed.
- Two-Way ANOVAWhy: When a between-groups factor sits alongside the occasion and the interaction of the two factors is the question.
Frequently asked questions
- Is a correction applied for sphericity?
- No. The panel does not run Mauchly's test and returns no Greenhouse-Geisser or Huynh-Feldt correction; the F value, degrees of freedom and p value are uncorrected. With three or more occasions it is up to you to justify that the difference variances are comparable, or to state the issue in your report.
- Should I prepare my data in wide or long format?
- Either works. Survey data usually lives in wide format, with one column per occasion, and ticking those columns in the parameter form is enough: the engine reshapes the table into long form itself. In long format each measurement is its own row and you point to the subject identifier and the within-subject factor. Because row position serves as the subject id in wide format, it matters that one row belongs to a single participant.
- The F value is significant, so which occasions differ?
- This output only says that at least one difference exists. To see pairwise comparisons you can run the paired t-test separately on the pairs that interest you, but every extra test raises the Type I error rate, so you need to adjust your alpha level for the number of comparisons yourself and report which adjustment you used.
- How many participants and occasions do I need?
- The method sets no strict minimum beyond having at least three occasions. Still, the error degrees of freedom come from the number of participants, so the test loses power quickly in small samples. And as the number of occasions grows, so does the amount of data expected from each participant, which raises the chance of missing observations and makes a balanced table harder to assemble.
References
- Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics
- Tabachnick, B. G., & Fidell, L. S. (2019). Using Multivariate Statistics
- Montgomery, D. C. (2019). Design and Analysis of Experiments
- statsmodels.stats.anova.AnovaRM documentation
Try it with your own data
The free plan includes 50 analysis runs a month and needs no card.