Group comparison

One-Way ANOVA

Compares the means of three or more independent groups in a single test and reports whether at least one group stands apart from the others.

Method summary

One-way analysis of variance divides the variability between groups by the variability within them to produce an F value. The question it answers is whether the observed group means lie further apart than samples drawn from one and the same population would. Because it answers that question for all groups at once, it avoids the accumulation of false positives that repeated pairwise t-tests produce. A significant F does not identify which group differs from which, so pairwise comparisons are needed afterwards. When F is significant and there are more than two groups, YouReply Analyze computes Tukey HSD, Bonferroni, Scheffe and Games-Howell comparisons together and lists them in a separate block.

Which research questions does it answer?

  • Do digital literacy scores differ across participants with four levels of education?
  • Is there a difference between satisfaction scores collected in three regions?
  • Does purchase intention differ between groups exposed to five different advertising messages?
  • Do three age cohorts differ in monthly time spent in the app?

When should you use it?

  • When three or more independent groups are compared.
  • When each participant belongs to a single group, with no repeated measurements from the same person.
  • When the dependent variable is continuous and the grouping variable is categorical.
  • When the first question is whether any difference exists at all, rather than which pair differs.

Required variable types

  • Dependent variable: a continuous numeric column, with the measurement level set accordingly on the Variable tab.
  • Independent variable: a single grouping column with three or more categories.
  • One row per participant; if repeated measurements sit in separate columns, this method does not apply.

Key assumptions

Independence of observations
The groups consist of different participants and no observation influences another. This is secured by the design of the data collection.
Continuous dependent variable
Because the F test compares group means, the dependent variable is expected to be measured at interval or ratio level.
Normality within groups
Each group's distribution should be approximately normal. Larger groups make the test more tolerant of departures.
Homogeneity of variance
The population variances are assumed equal. When they are not, the F test can mislead, especially with unequal group sizes.

How YouReply checks these assumptions

  • Independence of observations: The panel does not check this assumption automatically; the researcher evaluates it.
  • Continuous dependent variable: The method picker requires a numeric column for this field; columns that do not fit dim the method and the reason is shown.
  • Normality within groups: The panel does not test normality while running ANOVA. Run the Normality tests method separately if you want to inspect the distributions.
  • Homogeneity of variance: Levene's test runs on every run and is reported in its own block, but it does not change the test that is run: if the variances are unequal you choose the Welch ANOVA method yourself.

How the analysis is run

  1. 1Upload your data file and check on the Data tab that the columns were read correctly.
  2. 2On the Variable tab confirm that the grouping column is categorical and the dependent variable numeric.
  3. 3On the Analysis tab pick One-Way ANOVA; the data requirements box shows which of your columns have three or more groups.
  4. 4In the parameter form select the grouping column and the dependent variable.
  5. 5Run the analysis. The result opens the F value, degrees of freedom, p value, effect sizes, group statistics and the Levene block as separate sections.
  6. 6If F is significant, open the post-hoc block to see which pairs differ, then download the results as Excel and the chart as PNG.

Statistics and tables produced

F statistic and degrees of freedom
Between-groups degrees of freedom are the number of groups minus one; within-groups degrees of freedom are the total number of observations minus the number of groups.
p value
The probability of an F value at least this large under the null hypothesis that all group means are equal.
Group statistics
Mean, median, mode, standard deviation and sample size for every group.
Homogeneity of variance block
Levene's p value and whether the variances are treated as homogeneous; the test itself does not change with this result.
Post-hoc comparisons
When F is significant and there are more than two groups, Tukey HSD, Bonferroni, Scheffe and Games-Howell results are listed together.
Group means chart
A bar chart comparing the group means is drawn in the results panel and can be downloaded as a PNG.

Effect size and confidence intervals

Eta squared
The proportion of total variance in the dependent variable associated with the grouping variable. The common benchmarks are 0.01 small, 0.06 medium and 0.14 large; the statistic tends to overstate the effect in the population.
Omega squared
A bias-corrected version of eta squared. In small samples it comes out noticeably lower and gives a more cautious estimate of the population effect.

No confidence interval is returned for the group means or for the effect sizes; the values are point estimates. In the panel a confidence interval is computed only for the one-sample t-test. If a journal requires intervals you need to compute them separately.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
In an internal study, do organisational commitment scores (1-7) differ between employees in four departments?
Variables
Dependent variable: organisational commitment score (continuous) · Independent variable: department (four categories)
Example result
Group means were 4.62 (n = 61), 5.18 (n = 58), 5.44 (n = 55) and 4.87 (n = 63). Levene's p = 0.331. F(3, 233) = 5.12, p = 0.002, eta squared = 0.062, omega squared = 0.050. Tukey HSD found significant differences between the second and first department (p = 0.031) and between the third and first (p = 0.001).
Interpretation
The omnibus test shows that at least one department differs. Eta squared indicates that roughly six per cent of the variation in commitment is associated with department, leaving the rest to other factors. The pairwise comparisons locate the difference, though the design does not support a causal reading.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

Variable examined

Training method trial (synthetic)

A hypothetical trial in which one hundred and fifty participants were assigned to three training arms. Pre test and post test are scored out of one hundred and study hours is a continuous measure. Because each person is measured twice, a paired comparison is also possible.

Rows
150
Columns
participant_id, training_arm, pre_test, post_test, study_hours
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: Does the post test mean differ across the three training arms?

Grouping variable
training_arm
Measured variable
post_test
F statistic
3.73
p value
0.026
Eta squared
0.048
Omega squared
0.035
Degrees of freedom (between groups)
2
Degrees of freedom (within groups)
147
Number of groups
3
Total observations
150

Homogeneity of variance (Levene)

p value
0.033
Homogeneous
No

Group statistics

Group statistics
RowMeanMedianStandard deviationValid observations
Blended61.8961.858.2450
Control56.0855.6011.8250
Online59.4559.8511.6350

Pairwise comparisons (post hoc)

Tukey HSD

Tukey HSD
First groupSecond groupMean differenceStandard errorq statisticp valueConfidence interval lowerConfidence interval upperSignificant
BlendedControl5.811.513.850.0200.75310.88Yes
BlendedOnline2.441.511.620.490-2.627.50No
ControlOnline-3.371.512.230.259-8.431.69No

Bonferroni

Bonferroni
First groupSecond groupFirst group meanSecond group meanMean differencet statisticp value (uncorrected)p value (adjusted)Cohen's dSignificant
BlendedControl61.8956.085.812.850.0050.0160.571Yes
BlendedOnline61.8959.452.441.210.2280.6850.242No
ControlOnline56.0859.45-3.37-1.440.1540.461-0.288No

Scheffe

Scheffe
First groupSecond groupMean differenceF statisticp valueSignificant
BlendedControl5.817.400.027Yes
BlendedOnline2.441.300.522No
ControlOnline-3.372.490.291No

Games-Howell

Games-Howell
First groupSecond groupMean differenceStandard errorq statisticDegrees of freedomp valueSignificant
BlendedControl5.812.044.0487.500.015Yes
BlendedOnline2.442.011.7188.280.449No
ControlOnline-3.372.342.0397.970.326No

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260913

How to report the result

Organisational commitment differed significantly by department, F(3, 233) = 5.12, p = .002, eta squared = .062; Tukey HSD comparisons placed the difference between the first department and the second and third.

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • Unnecessary when only two groups are compared; the independent samples t-test answers that question directly.
  • Not for repeated measurements taken from the same participants; that design calls for repeated measures ANOVA.
  • When variances differ markedly and group sizes are unequal, the F test can mislead and Welch's ANOVA is preferable.
  • A significant F alone does not say which groups differ; without post-hoc comparisons the interpretation is incomplete.
  • With an ordinal dependent variable, interpreting means is contentious and a rank-based test suits better.
  • Outliers shift group means and distort both the F value and the effect size.

What to use when the assumptions are not met

  • Welch's ANOVAWhy: When homogeneity of variance does not hold; it adjusts the degrees of freedom and is more reliable with unequal group sizes.
  • Kruskal-Wallis H TestWhy: When normality does not hold or the dependent variable is ordinal; it compares ranks rather than means.
  • ANCOVA (Analysis of Covariance)Why: When groups must be compared while holding a continuous variable such as a pre-test score or age constant.

Frequently asked questions

Why not simply run several t-tests?
Four groups require six pairwise tests, and with each running at a five per cent error rate the chance of at least one false positive grows substantially. ANOVA controls that rate in a single test, and adjusted pairwise comparisons follow it.
Levene's test came out significant, what now?
The panel reports the result but does not switch tests. If the variances are not homogeneous and the group sizes are unequal, run the Welch ANOVA method as well, compare the two results and state which one you report.
Which post-hoc result should I report?
The panel computes four of them. Tukey HSD is the usual choice when variances are homogeneous, Games-Howell when they are not. You are expected to state which one you used and why; reporting all four is not required.
What is the difference between eta squared and omega squared?
Both express the proportion of variance explained, but eta squared describes only the sample at hand and tends to overstate the population effect. Omega squared corrects that bias, and the gap between them widens in small samples.

References

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.