Group comparison

Dunn's Test

Compares the ranks of three or more independent groups pair by pair, ranking the sample once and adjusting the p values for multiplicity.

Method summary

Dunn's test is what you reach for once an overall rank test has shown that groups separate and the question becomes which pairs are responsible. What sets the procedure apart is that it ranks the entire sample a single time: every group's mean rank is read off that one shared ranking, and pairs are never re-ranked among themselves. The comparisons therefore sit on a scale consistent with each other and rest on the same information the overall test used. For each pair the procedure reports the difference in mean ranks, a z value standardising that difference, and the matching p value, with an explicit correction for tied values built into the variance term. Because false positives accumulate as comparisons multiply, the p values are adjusted by the Holm or Bonferroni procedure, and turning the adjustment off is also available.

Which research questions does it answer?

  • Which pairs of four branches separate on service quality ratings?
  • Between which of five age cohorts does reported frequency of technology use differ?
  • Which of three education levels stand apart on an ordinal attitude item?
  • When an overall rank test is significant, does the difference come from one group or from several pairs?

When should you use it?

  • After an overall rank test has shown separation among three or more groups and the source needs locating.
  • When the dependent variable is ordinal, or normality cannot be defended for a continuous one.
  • When the pairwise comparisons should rest on one shared ranking and share a basis with the overall test.
  • When the accumulation of false positives across many pairs will be held in check by an adjustment.

Required variable types

  • Dependent variable: a single measurement column that is ordinal or coded numerically.
  • Independent variable: one grouping column with at least three categories; below that the engine will not run.
  • Each row stands for a single participant and belongs to exactly one group.

Key assumptions

Independence of observations
The groups consist of different participants. If several conditions were measured on the same person, pairwise work calls for a matched procedure.
At least ordinal measurement
Putting the whole sample into one ranking requires the values to be meaningfully comparable.
Three or more groups
The procedure exists to handle the multiplicity problem, so it needs three groups at minimum. With two there is nothing left to adjust.
Similar distribution shapes
Reading a result as a difference in location requires the groups to have similarly shaped distributions. Where the shapes clearly diverge, a significant pair may be reflecting a difference in spread as well.

How YouReply checks these assumptions

  • Independence of observations: The panel does not check this assumption automatically; the researcher evaluates it.
  • At least ordinal measurement: The method picker asks for a numeric column in the value field and a column with at least three categories in the grouping field; selections that do not fit dim the method.
  • Three or more groups: The engine rejects the analysis when the grouping column holds fewer than three categories and states the reason.
  • Similar distribution shapes: The panel does not inspect distribution shapes; the researcher makes that judgement from descriptive statistics.

How the analysis is run

  1. 1Upload your data file and check on the Data tab that the category labels in the grouping column are spelled consistently.
  2. 2On the Variable tab set the level of the value column to ordinal and declare the category labels.
  3. 3Choose Dunn's test in the method picker on the Analysis tab; the data requirements box shows which columns carry three or more categories.
  4. 4In the parameter form name the grouping column, the value column and the adjustment procedure, which can be Holm, Bonferroni or switched off.
  5. 5Run the analysis. The comparison table, the mean rank table, the list of significant pairs and the name of the adjustment used open as separate sections.
  6. 6Download the comparison table as Excel and state in your write-up which adjustment you used.

Statistics and tables produced

Comparison table
For each pair: the two group names, the difference in mean ranks, the z value, the raw p value, both sample sizes, the adjusted p value and a significance marker.
Mean rank per group
Mean ranks read off the single ranking built across the whole sample; these values carry the direction of every difference.
Group sizes and total observations
The sample size of each group and the total entering the analysis; imbalance influences how large the z values come out.
Comparison counts
How many pairwise comparisons were run and how many of them count as significant on their adjusted p value.
List of significant pairs
The pairs whose adjusted p value falls below the threshold, collected into a separate list.
Adjustment used
Which multiple comparison procedure was applied in the calculation, a detail that belongs in the report.

Effect size and confidence intervals

z value per comparison
The procedure returns no separate measure of effect; the standardised quantity you have is the z value computed for each pair. It divides the difference in mean ranks by a standard error derived from the group sizes and the tie correction, which makes it dependent on sample size and unsuitable for use as an effect measure across studies. To judge practical importance, read the difference in mean ranks together with the group medians.

The procedure produces no confidence interval for any comparison; differences in mean ranks and z values arrive as point estimates. An interval for a difference in location has to be computed separately.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
In a service study, which pairs of four branches separate on a five-point willingness-to-recommend item?
Variables
Value variable: willingness to recommend (1-5, ordinal) · Grouping variable: branch (four categories)
Example result
Total observations: 164. Mean ranks: first branch 68.0 (n = 41), second branch 79.0 (n = 42), third branch 89.2 (n = 40), fourth branch 94.1 (n = 41). Six comparisons, Holm adjustment. First versus fourth branch: rank difference -26.1, z = -3.37, raw p = 0.0008, adjusted p = 0.005. First versus third branch: rank difference -21.2, z = -2.74, raw p = 0.0062, adjusted p = 0.031. Second versus fourth branch: z = -1.95, raw p = 0.051, adjusted p = 0.204. The remaining three comparisons have adjusted p values of 0.468 and above. Significant comparisons: 2.
Interpretation
The separation traces back to the first branch, whose willingness-to-recommend ranks sit clearly below both the third and the fourth branch. The gap between the second and the fourth branch touches the threshold on its raw p value but loses significance once six comparisons are accounted for, a textbook illustration of why switching the adjustment off would mislead. Nothing separates the third branch from the fourth, so the picture is one branch lagging rather than a gradient running from one end to the other.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

General customer survey (synthetic)

A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.

Rows
300
Columns
respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: Which pairs of the three education levels differ from each other in income?

Grouping variable
education
Measured variable
income
Multiple comparison correction
holm
adjust
holm
Total observations
300
Number of comparisons
3
Number of significant comparisons
3

Pairwise comparisons

Pairwise comparisons
First groupSecond groupMean rank differencez statisticp valueFirst group sizeSecond group sizep value (adjusted)Significant
BachelorHigh school105.408.68< 0.001100104< 0.001Yes
BachelorPostgraduate-38.26-3.090.002100960.002Yes
High schoolPostgraduate-143.65-11.70< 0.00110496< 0.001Yes

Mean ranks

Bachelor
174.80
High school
69.40
Postgraduate
213.05

Group sizes

Bachelor
100
High school
104
Postgraduate
96

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914

How to report the result

Holm-adjusted Dunn comparisons showed the first branch separating significantly from the fourth (z = -3.37, adjusted p = .005) and from the third (z = -2.74, adjusted p = .031) on willingness to recommend, with no difference in the remaining four pairs (N = 164).

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • Scanning every pair before an overall rank test is significant amounts to letting the data set the research question, which is hard to report.
  • Not for repeated measurements from the same participants; a matched design needs a signed-rank procedure for its pairs.
  • As the number of groups grows the number of comparisons grows quickly, and the adjustment can push genuine differences back above the threshold.
  • Switching the adjustment off leaves the false positive rate uncontrolled and should serve nothing beyond an exploratory glance.
  • Since the output carries no effect size beyond the per-comparison z value, judging practical importance falls to reading medians and rank differences.
  • In very small groups the z value rests on a normal approximation, which makes the adjusted p values fragile.

What to use when the assumptions are not met

  • Kruskal-Wallis H TestWhy: When the overall verdict is needed before any pairs; in the panel that method also produces Bonferroni-adjusted pairwise Mann-Whitney comparisons of its own once it is significant.
  • Mann-Whitney U TestWhy: When a theoretical expectation concerns one specific pair of groups and there is no need to scan every pair.
  • One-Way ANOVAWhy: When the measurement is continuous and the distributional assumptions hold; four post-hoc procedures are computed together once F is significant.

Frequently asked questions

Is this the same as the pairwise comparisons the Kruskal-Wallis method produces?
It is not. The panel's Kruskal-Wallis method treats each pair as a separate Mann-Whitney run once its overall test is significant and applies a Bonferroni adjustment to those runs, with every run re-ranking its own two groups. Dunn's test instead ranks the whole sample once and works from the mean ranks read off that shared ranking, so it rests on the same basis as the overall test. That is why it is the standard post-hoc choice, and why you need to run this method separately.
Should I pick Holm or Bonferroni?
Both control the family-wise error rate. Bonferroni multiplies every p value by the number of comparisons and gives the most conservative answer; Holm sorts the p values and applies a stepped multiplier, offering the same protection at a smaller cost in power. Holm is the usual preference, but the classical Bonferroni option is there if a reviewer asks for it.
When is it reasonable to switch the adjustment off?
In practice only when you want to see the raw table. With no adjustment every comparison runs at its own five per cent error rate, and across six comparisons the chance of at least one false positive approaches twenty-five per cent. Published results are expected to use the adjusted p values and to name the procedure applied.
Is the result trustworthy on short scales with many identical answers?
Ties are accounted for: the standard error behind the z value carries an explicit correction for them, so p values on short scales are not inflated systematically. What does suffer is information, because the more ties there are the less the ranking distinguishes, and the power to detect real differences falls with it.

References

  • Dunn, O. J. (1964). Multiple Comparisons Using Rank Sums
  • Conover, W. J. (1999). Practical Nonparametric Statistics
  • Holm, S. (1979). A Simple Sequentially Rejective Multiple Test Procedure
  • scipy.stats.rankdata documentation

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.