Group comparison

MANOVA

Combines several continuous outcomes into a single test of whether groups differ on that set of outcomes taken together.

Method summary

Multivariate analysis of variance treats several continuous outcomes measured on the same grouping as one object rather than as separate analyses. The question becomes whether the groups differ in their average profile across those outcomes. Because outcomes are usually related to one another, running a separate test for each both ignores that relationship and inflates the chance of a false positive as the number of tests grows. One multivariate test eases both problems and can detect a separation even when the differences on each individual outcome are small but point the same way. YouReply Analyze fits the model, reports four multivariate criteria side by side, tests the equality of the covariance matrices and opens a follow-up analysis for each outcome.

Which research questions does it answer?

  • Do customers on three subscription tiers differ in the profile formed by satisfaction with price, ease of use and support?
  • Taking brand awareness, preference and willingness to recommend together, is there a regional difference across four regions?
  • Do seniority bands separate on the pair of outcomes made up of job satisfaction and organisational commitment?
  • Do groups shown three advertising treatments show the same pattern across recall, liking and purchase intent?

When should you use it?

  • When more than one continuous outcome belongs to the same research question and a single verdict is wanted over them.
  • When the outcomes form a theoretical whole, for instance complementary subscales of one instrument.
  • When the outcomes are moderately related. Very weak relationships make separate analyses easier to read, and very strong ones suggest the outcomes measure the same thing.
  • When there are three or more groups. It works with two as well, although the interpretation there comes close to a set of mean comparisons.

Required variable types

  • Dependent variables: two or more continuous numeric columns, chosen as a list in the parameter form.
  • Group column: a single column carrying categorical values, which may have more than two categories.
  • A row missing a value on any one of the outcome columns leaves the whole analysis, so the usable row count shrinks as the outcome list grows.

Key assumptions

The outcome set hangs together
The chosen outcomes should represent one theoretical whole. Putting columns about unrelated topics into the same list makes it unclear which outcome produced any separation that appears.
Equality of covariance matrices
Groups should show a similar pattern of variances and covariances among the outcomes. If the outcomes are tightly linked in one group and loosely linked in another, the error rate of the multivariate test drifts.
Multivariate normality
The outcome set is expected to be roughly jointly normal within each group. Each outcome looking normal on its own does not guarantee this, although marked skewness usually shows up at both levels.
Similar variances on each outcome
The univariate analyses read after the multivariate test expect the spread of each outcome to be similar across groups.
Independence of observations
Each row must belong to a different participant. This method does not suit data where one person's answers at several time points appear as separate rows.

How YouReply checks these assumptions

  • The outcome set hangs together: The panel does not check this assumption automatically; the researcher evaluates it.
  • Equality of covariance matrices: Box's M test runs on every run, its verdict is judged at the STRICT level of 0.001, and the group sizes are printed beside it. The lenient threshold is used because the test turns significant easily in large samples. Whatever it returns, the model is fitted the same way.
  • Multivariate normality: The panel runs no normality check inside this method. To inspect the distributions you can run the normality tests method on the relevant columns separately.
  • Similar variances on each outcome: The follow-up returns a Levene p value and a homogeneity flag for every dependent variable. The flag is informational: no alternative test is substituted when homogeneity fails.
  • Independence of observations: The panel does not check this assumption automatically; the researcher evaluates it.

How the analysis is run

  1. 1Drag and drop your data file into the panel; both CSV and Excel files are read.
  2. 2In the grid on the Data tab, check that every outcome column parsed as numeric and that no value arrived as text because of a separator or a unit suffix.
  3. 3On the Variable tab set the outcome columns to scale level and the group column to nominal, and write value labels for the categories.
  4. 4Pick MANOVA on the Analysis tab; the search box finds it by name. With the switch for only the methods that fit your data turned on, unsuitable methods are dimmed with a plain-language reason rather than removed from the list.
  5. 5In the parameter form name the list of dependent variables and the group column, then run the analysis. This method has no optional fields.
  6. 6The result opens as collapsible sections: the four multivariate criteria, the Box's M block, a follow-up analysis per outcome and the descriptive table.
  7. 7No method-specific chart is drawn here. You can export the tables to Excel and read the call and version from the computation credits card. The run is written to the history tab; the free plan allows fifty runs a month.

Statistics and tables produced

Wilks' lambda
The most commonly reported multivariate criterion. It comes with its value, an F equivalent, a p value and two degrees of freedom; the closer to zero, the stronger the group separation.
Pillai's trace
A criterion answering the same question with a different aggregation rule. It is held to be the most robust of the four when covariance matrices diverge and group sizes are unbalanced.
Hotelling-Lawley trace
A criterion giving the total share of group separation across the dimensions, reported with its value, F, p and degrees of freedom.
Roy's greatest root
A criterion looking only at the strongest dimension of separation. It is the most sensitive when the separation runs in one direction and the most optimistic when it spreads over several.
Box's M block
The test of equality of covariance matrices, its verdict at the 0.001 level and the group sizes. The sizes sit next to the verdict because imbalance makes a violation of this assumption matter more.
Follow-up analysis per outcome
A univariate analysis of variance for each dependent variable: F value, p value, partial eta squared and a Levene p value with a homogeneity flag. These p values are NOT adjusted for the number of outcomes tested.
Descriptive table
Mean, median, mode and standard deviation for every group and every outcome. This is where you read the direction of the separation the multivariate test pointed to.

Effect size and confidence intervals

Partial eta squared (follow-up analyses)
Reported separately for each outcome, it gives the share of variability in that outcome associated with group membership. The thresholds quoted are 0.01 small, 0.06 medium and 0.14 large. Comparing the values shows where the separation concentrates: if eta squared is clearly high on one outcome and near zero on the rest, the multivariate verdict largely comes from that single outcome.
Multivariate effect size
None is produced. The output carries no effect size for the outcome set as a whole, and although the literature mentions measures such as multivariate eta squared derived from Wilks' lambda, the panel does not return them. To report a magnitude for the whole set you would have to work it out outside the panel from the lambda value that is returned.

No value in this method comes with a confidence interval. No interval accompanies the four multivariate criteria, the F and partial eta squared figures of the follow-up analyses arrive as point estimates, and the means in the descriptive table carry no standard error or interval either. If you are writing for an outlet that expects uncertainty reported as intervals, those figures have to be computed elsewhere.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
Representative example: satisfaction among customers on three subscription tiers of a software product was measured on three separate dimensions, and the question is whether the tiers differ on the profile those three dimensions form. The starter, professional and enterprise tiers contributed 98, 105 and 87 complete responses, giving 290 usable rows.
Variables
Dependent variables: price satisfaction, ease-of-use satisfaction, support satisfaction (each 1-7, continuous) · Group column: subscription tier (starter / professional / enterprise) · Unit of observation: one customer per row
Example result
Wilks' lambda = 0.864, F(6, 570) = 7.23, p < .001. Pillai's trace = 0.138, F(6, 572) = 7.08, p < .001. The Hotelling-Lawley trace was 0.155 and Roy's greatest root 0.149, both with p < .001. Box's M returned p = .024, so the assumption was treated as met at the 0.001 threshold. Follow-ups: price satisfaction F(2, 287) = 14.92, p < .001, partial eta squared = 0.09, Levene p = .214; ease of use F(2, 287) = 3.05, p = .049, partial eta squared = 0.02, Levene p = .402; support satisfaction F(2, 287) = 1.18, p = .310, partial eta squared = 0.01, Levene p = .067. The price satisfaction means were 4.12, 4.58 and 5.21.
Interpretation
All four criteria point the same way: the tiers differ as profiles. Reading the follow-ups to locate the source, the weight falls mostly on price satisfaction, whose partial eta squared is medium and whose means rise steadily with the tier. Ease of use returned p = .049, which means that tightening the threshold yourself for three tested outcomes puts it outside the cut; the printed p value is unadjusted, and accepting it as it stands ignores the cost of having run three tests. Support satisfaction shows no separation. Box's M raises no concern at the 0.001 threshold, though with group sizes not quite balanced that note offers lenient reassurance. Which two tiers differ cannot be read from this output.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

General customer survey (synthetic)

A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.

Rows
300
Columns
respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: Taken together, do satisfaction and loyalty differ by gender?

Dependent variables
satisfaction,loyalty
Grouping variable
gender
Wilks' lambda value
0.993
F statistic
1.09
p value
0.337
Degrees of freedom (numerator)
2
Degrees of freedom (denominator)
297

Multivariate test statistics

Multivariate test statistics
RowStatistic valueF statisticp valueDegrees of freedom (numerator)Degrees of freedom (denominator)
Wilks' lambda0.9931.090.3372297
Pillai's trace0.0071.090.3372297
Hotelling-Lawley trace0.0071.090.3372297.00
Roy's greatest root0.0071.090.3372297

Homogeneity of variance (Levene)

Test statistic
2.22
p value
0.083
Homogeneous
Yes

Group sizes

Female
159
Male
141

Univariate follow-up tests

Univariate follow-up tests
RowF statisticp valuePartial eta squaredLevene p valueHomogeneous
satisfaction1.150.2840.0040.010No
loyalty0.7580.3850.0030.907Yes

Descriptive statistics

Descriptive statistics
Rowsatisfactionloyalty
Female--
Male--

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914

How to report the result

Subscription tier had a significant effect on the set of three satisfaction dimensions, Wilks' lambda = 0.864, F(6, 570) = 7.23, p < .001; in the follow-up analyses the difference was clear only for price satisfaction, F(2, 287) = 14.92, p < .001, partial η² = 0.09 (means 4.12, 4.58 and 5.21), while the difference in ease of use sat at the edge of the threshold once the multiple-testing burden was allowed for, F(2, 287) = 3.05, p = .049.

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • No multivariate effect size is returned, so the output offers nothing to report as a magnitude for the outcome set as a whole.
  • The p values of the follow-up analyses are not adjusted for the number of outcomes tested; tightening the threshold is entirely the researcher's decision, and reporting the unadjusted figures as they stand manufactures false positives.
  • There are no pairwise comparisons showing which pair of groups differs; the direction can only be inferred from the descriptive means.
  • Multivariate normality is not tested and no method-specific chart is produced, so nothing in the output lets you eyeball the distribution of the outcome set.
  • No confidence interval is given for any figure.
  • A row missing a value on just one outcome column is dropped too, so with long outcome lists the usable sample can shrink markedly below what you expect.

What to use when the assumptions are not met

  • One-Way ANOVAWhy: When the outcomes measure unrelated topics, or a single outcome is being asked about; each outcome then reports more readably under its own heading.
  • MANCOVAWhy: When a continuous variable affecting the outcome set has also been measured; the group effect is then tested net of its share and adjusted means are returned.
  • Two-Way ANOVAWhy: When there is one outcome but two categorical factors under study together, so the interaction term can be tested as well.
  • ANCOVA (Analysis of Covariance)Why: When there is one outcome and one continuous variable to control for; the covariate enters the model and adjusted means are estimated.

Frequently asked questions

Why a single multivariate test instead of one ANOVA per outcome?
There are two reasons. First, outcomes are usually related, and separate tests discard that relationship, whereas a multivariate test brings it inside the model so that differences too small to reach significance one at a time can reveal a separation when they align. Second, running five tests for five outcomes raises the false positive rate appreciably, and a single decision point limits that cost. Once the multivariate verdict is significant, though, the question returns to the individual outcomes, and the follow-up analyses the panel provides are read at exactly that point. Remember that those follow-up figures carry no adjustment.
Box's M came out significant. What should I do?
Look at the threshold first. The printed verdict is judged at 0.001 because this test turns significant over tiny differences as samples grow; at 0.05 it would declare a violation in almost any large dataset. If the verdict still reads as a violation at 0.001, turn to the group sizes printed beside it. When the sizes are close together the multivariate test is comparatively robust to the violation, and leading with Pillai's trace is a common way to report it. When the sizes are markedly unbalanced you should present the result cautiously, note the violation in the write-up and consider falling back to analysing the outcomes one at a time. The panel does not change the model when the assumption fails.
Which of the four criteria should I report?
All four test the same null hypothesis and usually reach the same verdict; where they part company is where the assumptions are under strain. Wilks' lambda is the one most often reported in the literature and is a reasonable default when the assumptions hold. If the covariance matrices diverge or the group sizes are unbalanced, Pillai's trace is considered more robust. Roy's greatest root, looking only at the strongest direction of separation, is the most sensitive when the separation sits on one dimension and too optimistic when it is spread out. If the four disagree, that disagreement is itself a finding and a reason to revisit the assumptions.
How many outcome variables should I choose?
A small number of outcomes that belong to the same theoretical whole gives the soundest analysis. Adding every continuous column hurts in two ways: the complete-rows requirement shrinks the sample, and the sensitivity of the multivariate test declines as outcomes carrying no separation are added. More outcomes also mean more p values in the follow-up stage and a heavier adjustment burden. Columns that correlate very highly may be measuring the same thing, in which case choosing one or building an index first yields a more interpretable result.

References

  • Tabachnick, B. G., & Fidell, L. S. (2019). Using Multivariate Statistics
  • Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics
  • Johnson, R. A., & Wichern, D. W. (2007). Applied Multivariate Statistical Analysis
  • statsmodels.multivariate.manova.MANOVA documentation

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.