Group comparison
MANCOVA
Takes several continuous outcomes together and tests group separation once the share of continuous variables that also explain those outcomes has been removed.
Method summary
This method joins two ideas: testing outcomes as a set rather than one at a time, and keeping in the model the continuous variables already known to explain that set. A covariate absorbs an influence that would otherwise be entangled with the grouping, and tenure, age, a prior-period measurement or a pre-test score are typical examples. Removing its share changes two things: the test gains sensitivity because the error term shrinks, and the group means are re-estimated on the assumption that the covariate sits at its sample average. YouReply Analyze reports multivariate statistics both for the group effect and for every covariate, so you can see whether a covariate genuinely relates to the outcome set, and then opens a follow-up analysis for each outcome that also carries the adjusted means.
Which research questions does it answer?
- With tenure controlled for, do office, hybrid and remote teams differ on the set formed by commitment and intention to stay?
- Once household monthly income enters the model, do three lifestyle segments differ in spending and saving tendencies?
- Holding baseline scores constant, do three intervention arms differ on the profile of knowledge, attitude and behavioural intent?
- Allowing for store size, do three layout formats produce the same result on customer satisfaction and basket satisfaction?
When should you use it?
- When a continuous variable known to influence the outcome set has been measured and the group comparison should be cleared of its effect.
- When the groups already differ on the covariate, since a comparison without it would present that difference as a group effect.
- When the covariate was chosen on theoretical grounds and measured at a moment grouping could not alter.
- When the outcomes are theoretically linked. If they are not, one covariate-adjusted analysis per outcome reports more readably.
Required variable types
- Dependent variables: two or more continuous numeric columns, chosen as a list.
- Group column: a single column carrying categorical values.
- Covariates: one or more continuous numeric columns; each gets its own multivariate row and its own adjustment share inside the follow-up blocks.
- For a row to enter the analysis it must carry a value on every outcome column and on every covariate column.
Key assumptions
- The covariate is independent of grouping
- A covariate's value must not be a consequence of group assignment. Using a variable measured after the intervention strips out part of the very effect under test and artificially erases the group separation.
- Linear relationship between covariate and outcomes
- The adjustment assumes the covariate relates linearly to each outcome. Where the relationship curves, the adjusted means come out biased and the shortfall in adjustment leaks into the group difference.
- Homogeneity of regression slopes
- The slope between the covariate and each outcome should be similar across all groups. If the slopes diverge, the group difference depends on the level of the covariate and a single comparison of adjusted means conceals that variation.
- Equality of covariance matrices
- Groups are expected to show a similar pattern of variances and covariances among the outcome variables.
- Similar variances on each outcome
- The follow-up analyses rest on the group spreads being close to one another on each outcome.
How YouReply checks these assumptions
- The covariate is independent of grouping: The panel does not check this assumption automatically; the researcher evaluates it.
- Linear relationship between covariate and outcomes: The shape of the relationship is not examined in this method, though the Wilks' lambda and p value returned for each covariate do show whether the covariate relates to the outcome set at all. To see the shape, run the correlation or regression methods separately.
- Homogeneity of regression slopes: This assumption is not tested. With no group-by-covariate interaction term in the fitted model, slope differences cannot be measured, so you must justify the assumption from your design or inspect the relationship group by group.
- Equality of covariance matrices: Box's M test runs and its verdict is given at the 0.001 level, with the group sizes beside it. The test is computed on the dependent variables ONLY; covariate columns take no part in it. The result does not change the model.
- Similar variances on each outcome: A Levene result and a homogeneity flag are reported for every dependent variable. The flag is informational; no alternative estimation is used when homogeneity fails.
How the analysis is run
- 1Drag your data file into the panel; CSV and Excel are both read.
- 2In the editable grid on the Data tab, confirm that the outcome and covariate columns resolved to numbers, correcting any cell that became text because of a decimal separator.
- 3On the Variable tab mark the outcomes and covariates as scale and the group column as nominal, and declare your missing-value codes there, otherwise a code enters the model as if it were a measurement.
- 4Choose MANCOVA on the Analysis tab. The data requirements box shows which columns meet the outcome, group and covariate conditions, and methods that do not fit stay dimmed in the list with a reason written out.
- 5In the parameter form name the outcome list, the group column and the covariates, then run it. The analysis will not start until all three are filled.
- 6The output opens in order: the multivariate criteria for the group effect, a multivariate row per covariate, the Box's M block, a follow-up analysis per outcome and the raw descriptives.
- 7Look for the adjusted means inside the follow-up block of the relevant outcome rather than in the descriptive table; the two tables show the same groups with different figures.
- 8No method-specific chart is produced; the tables export to Excel. The computation credits card names the library call and version, and the run is written to the history tab.
Statistics and tables produced
- Four multivariate criteria for the group effect
- Wilks' lambda, Pillai's trace, the Hotelling-Lawley trace and Roy's greatest root, each with its value, F equivalent, p value and two degrees of freedom. They are computed with the covariates' share held in the model.
- A multivariate row per covariate
- A separate Wilks' lambda and p value for every covariate chosen. This is where you learn whether the covariate relates to the outcome set; a large p value suggests it may be contributing nothing.
- Box's M block
- The test of equality of covariance matrices, its verdict at the 0.001 level and the group sizes. The calculation uses the dependent variables alone.
- Follow-up analysis per outcome
- A covariate-adjusted univariate analysis for each dependent variable: F value, p value, partial eta squared, the Levene result and the covariate-adjusted group means. These p values carry no adjustment for the number of outcomes tested.
- Adjusted group means
- Values estimated on the assumption that the covariates sit at their sample average. They live inside the follow-up block of each outcome and are not gathered into a separate summary table.
- Raw descriptives
- Mean, median, mode and standard deviation per group and per outcome. These figures are UNADJUSTED: they describe the data as observed and are not expected to match the adjusted means.
Effect size and confidence intervals
- Partial eta squared (follow-up analyses)
- Given for each outcome separately, it expresses the portion of the variability left after the covariates that falls to group membership. The usual thresholds are 0.01 small, 0.06 medium and 0.14 large. The figure belongs to the covariate-adjusted model: because the covariate shrinks the denominator, it can exceed the value from a covariate-free analysis of the same data and should not be compared directly across the two models.
- Strength of the covariate relationship
- A covariate's own Wilks' lambda indicates how strongly it ties to the outcome set, with values approaching one meaning it explains less. The figure is not a standardised effect size measure, so use it only to argue whether the covariate earns its place in the model.
- Multivariate effect size
- None is computed. No effect size for the outcome set as a whole is returned, either for the group effect or for a covariate. If your write-up calls for such a figure, you will have to derive it outside the panel from the lambda values returned.
No section of the output carries a confidence interval. The adjusted group means arrive as point estimates with no standard error or interval, and nothing accompanies the multivariate criteria, the covariate rows or the partial eta squared figures either. Conveying the uncertainty around an adjusted mean is a calculation you have to do yourself.
Example research question and example result
The numbers below are a representative example, not data from a real study or a real user.
- Research question
- Representative example: two outcomes were measured among office-based, hybrid and remote teams in one organisation. Because tenure was distributed unevenly across the groups and is known to relate to both outcomes, it entered the model as a covariate. The three arrangements contributed 81, 112 and 76 complete responses, giving 269 usable rows.
- Variables
- Dependent variables: commitment score (1-5, continuous) and intention-to-stay score (1-5, continuous) · Group column: working arrangement (office / hybrid / remote) · Covariate: tenure with the organisation (years, continuous)
- Example result
- Group effect: Wilks' lambda = 0.912, F(4, 528) = 6.21, p < .001; Pillai's trace = 0.089, F(4, 530) = 6.17, p < .001. For the tenure covariate, Wilks' lambda = 0.871, p < .001. Box's M treats the assumption as met at the 0.001 level (p = .058). In the follow-ups, commitment gave F(2, 265) = 9.84, p < .001, partial eta squared = 0.07, Levene p = .511, and intention to stay F(2, 265) = 2.41, p = .092, partial eta squared = 0.02, Levene p = .338. The raw commitment means were 3.74, 4.02 and 3.58, while the adjusted means in the same order were 3.81, 3.97 and 3.64. Mean tenure was 6.2, 4.8 and 4.1 years respectively.
- Interpretation
- Tenure's own lambda is smaller than the lambda for the group effect, meaning the covariate explains the outcome set more strongly than the grouping does and earns its place in this model. Setting the raw and adjusted means side by side shows what the adjustment did: the office team has the longest tenure, so its raw mean was pulled upward, the remote team shifted slightly the same way, and the widest gap across the three groups fell from 0.44 to 0.33 points. The figures in the two tables are therefore not interchangeable, and your write-up has to say which set you are quoting. The separation persists on commitment, whereas on intention to stay the difference between adjusted means stays below the threshold. With two outcomes tested, the follow-up p values should be judged against a tightened threshold rather than taken as printed. This output does not say which two arrangements differ.
Real output on a sample dataset
The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.
General customer survey (synthetic)
A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.
- Rows
- 300
- Columns
- respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.
Research question: With age held constant, do satisfaction and loyalty jointly differ by gender?
- Dependent variables
- satisfaction,loyalty
- Grouping variable
- gender
- Covariates
- age
- Wilks' lambda value
- 0.994
- F statistic
- 0.942
- p value
- 0.391
- Degrees of freedom (numerator)
- 2
- Degrees of freedom (denominator)
- 296
Multivariate test statistics
| Row | Statistic value | F statistic | p value | Degrees of freedom (numerator) | Degrees of freedom (denominator) |
|---|---|---|---|---|---|
| Wilks' lambda | 0.994 | 0.942 | 0.391 | 2 | 296 |
| Pillai's trace | 0.006 | 0.942 | 0.391 | 2 | 296 |
| Hotelling-Lawley trace | 0.006 | 0.942 | 0.391 | 2 | 296.00 |
| Roy's greatest root | 0.006 | 0.942 | 0.391 | 2 | 296 |
Covariate results
| Row | Wilks' lambda value | p value |
|---|---|---|
| age | 0.995 | 0.464 |
Homogeneity of variance (Levene)
- Test statistic
- 2.22
- p value
- 0.083
- Homogeneous
- Yes
Group sizes
- Female
- 159
- Male
- 141
Univariate follow-up tests
| Row | F statistic | p value | Partial eta squared | Levene p value | Homogeneous | Adjusted means |
|---|---|---|---|---|---|---|
| satisfaction | 0.941 | 0.333 | 0.003 | 0.010 | No | - |
| loyalty | 0.700 | 0.403 | 0.002 | 0.907 | Yes | - |
Descriptive statistics
| Row | satisfaction | loyalty |
|---|---|---|
| Female | - | - |
| Male | - | - |
Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914
How to report the result
With tenure entered as a covariate, working arrangement had a significant effect on the outcome set of commitment and intention to stay, Wilks' lambda = 0.912, F(4, 528) = 6.21, p < .001, and the covariate itself related to the set, Wilks' lambda = 0.871, p < .001. The follow-up difference appeared on commitment, F(2, 265) = 9.84, p < .001, partial η² = 0.07, with adjusted means of 3.81 for office, 3.97 for hybrid and 3.64 for remote, while no significant difference emerged on intention to stay, F(2, 265) = 2.41, p = .092.
An example sentence close to APA style; the numbers are representative.
When you should not use it
- Homogeneity of regression slopes is not tested; with no group-by-covariate interaction in the model, the adjusted means are biased if the slopes truly diverge.
- Box's M is computed on the dependent variables alone, so any difference in how the covariates spread across groups enters no test and is checked nowhere in the output.
- The follow-up p values are not adjusted for the number of outcomes tested, leaving the responsibility for tightening the threshold with the researcher.
- There are no pairwise comparisons of adjusted means, so a significant group effect leaves the question of which groups differ unanswered.
- No effect size is returned for the outcome set as a whole, and no value anywhere carries a confidence interval.
- Controlling for a covariate does not stand in for random assignment; where groups formed on their own, unmeasured differences remain outside the model and adjusted means support no causal claim.
What to use when the assumptions are not met
- MANOVAWhy: When there is no continuous variable to control for, or a covariate's lambda came out very close to one; the model answers the same question without the extra term.
- ANCOVA (Analysis of Covariance)Why: When the outcome is a single continuous variable; the covariate-adjusted analysis reads more directly under its own heading.
- One-Way ANOVAWhy: When there is one outcome and nothing to control for; it is the plainest group comparison available.
Frequently asked questions
- Why do the adjusted means differ from the means in the descriptive table?
- The descriptive table reports the data as observed, each group with its own tenure and its own age distribution, simply as it is. The adjusted means are the product of an assumption instead: an estimate answering what the outcome average would be if every group sat at the sample average of the covariate. When groups differ on the covariate, the two figures move apart, and the direction depends on where that group stands: a group above the covariate average is pulled down, one below it is pulled up. The panel keeps the two sets in separate places, raw values in the descriptive section and adjusted values inside the relevant outcome's follow-up block. Your write-up has to state plainly which one it quotes.
- What does a non-significant covariate mean?
- The Wilks' lambda and p value returned for each covariate say whether that covariate relates to the outcome set. A large p value means it is not explaining the outcomes, in which case keeping it carries a cost with no gain. The cost is that every covariate spends degrees of freedom and that a needless term makes the estimates less stable. Faced with such a result, dropping the covariate and running the analysis without it is usually easier to defend, although if you chose the covariate for theoretical reasons and the design demands it, the decision should not rest on a p value alone.
- Do the covariates enter Box's M test?
- They do not. The test runs on the covariance matrices of the dependent variables only, and whether the covariate columns spread similarly across groups is checked nowhere. Knowing the distinction matters: if a covariate's variance is markedly different in one group, the adjustment rests on weaker ground there, and a reassuring Box's M block will not catch it. To see how a covariate is distributed by group, run the descriptive statistics method with a group breakdown separately.
- How many outcomes and how many covariates should I choose?
- Think about the two together, because both tighten the complete-rows requirement: a row enters the analysis only if it carries values on every outcome and every covariate, so the sample erodes quickly as the lists grow. Two or three theoretically linked outcomes with one or two clearly justified covariates serve most survey designs. Adding covariates that correlate highly with each other destabilises the estimates, and padding the outcome list with columns that carry no separation lowers the sensitivity of the multivariate test.
References
- Tabachnick, B. G., & Fidell, L. S. (2019). Using Multivariate Statistics
- Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics
- Hair, J. F., Black, W. C., Babin, B. J., & Anderson, R. E. (2019). Multivariate Data Analysis
- statsmodels.multivariate.manova.MANOVA documentation
Try it with your own data
The free plan includes 50 analysis runs a month and needs no card.