Categorical

Cochran's Q Test

Evaluates in a single test whether the acceptance rates of three or more yes and no items measured in the same participants differ meaningfully from one another.

Method summary

Cochran's Q test carries the logic of a two-column paired binary comparison out to three or more columns. Two uses are typical: several yes and no items answered by the same participants, such as which services from a list they have used; or a single item measured repeatedly at three or more points in time. The calculation works within each participant, which means that people who answer yes to everything, or yes to nothing, contribute nothing to the differences between columns, and the separation comes from those whose answer pattern is mixed. The Q statistic summarises how far the columns' counts of yes answers move apart, scaled by each participant's own total, and is judged against a chi-square distribution with degrees of freedom equal to the number of columns minus one. What comes out is a single overall verdict: does at least one column stand apart from the others. YouReply Analyze reports the Q value, the degrees of freedom, the p value and, for each column, the count or proportion of yes answers.

Which research questions does it answer?

  • Do the usage rates of four listed digital services differ within the same set of participants?
  • Has the rate of compliance with a rule stayed level across three consecutive measurement waves?
  • Is trust in five different information sources, asked in yes and no form, at the same level throughout?
  • Do purchase intention rates differ across three packaging options shown to the same people?

When should you use it?

  • When there are three or more binary columns and every one of them was measured in the same participants.
  • When a single item measured at three or more points in time is compared and the answer is a yes or no.
  • When the question is whether the options ticked in a multiple response question separate in their acceptance rates.
  • When you intend to run an overall test first and move on to pairwise comparisons only if the overall test calls for it.
  • When what is compared is repeated binary answers from the same people rather than independent groups.

Required variable types

  • Three or more columns, each carrying the same two levels, such as 1 and 0 or ticked and not ticked.
  • The coding must be identical across all columns; a no shown as an empty cell in one column and as a zero in another reads as two different things.
  • Each row stands for one participant, so that the comparison between columns can be made within the person.
  • Where only two columns are involved, the two-column test defined for paired binary measures is the right tool instead of this one.
  • Continuous or multi-category measures cannot feed this test. Reducing them to binary happens through derived columns on the Variable tab, and the reduction belongs in your report.

Key assumptions

All columns come from the same participants
The columns being compared must be measured in the same people. When different items were put to different respondents, no within-row comparison exists and Q answers the wrong question.
Participants independent of one another
The link between one participant's columns is what the design intends, but no such link should exist between two different participants. In clustered samples the p value looks smaller than it should.
Two-level responses
Every column has to be reduced to exactly two levels. A three-level response scale cannot enter the Q calculation.
Complete rows
Since the comparison works within the row, a participant missing an answer in one column drops out of all of them. With long item lists this can narrow the sample noticeably.
Enough mixed patterns
Participants who say yes to every item, or to none, carry no information about differences between columns. When most of the sample piles up at those two extremes, the test effectively rests on a small subset.

How YouReply checks these assumptions

  • All columns come from the same participants: Because the pairing comes from how your file is laid out, the panel cannot verify it; confirming that the columns line up with the right person in each row is your work on the Data tab.
  • Participants independent of one another: The panel does not check this assumption automatically; the researcher evaluates it.
  • Two-level responses: The columns are read as binary during the run; the analysis will not proceed when more than two levels are present or when a column collapses to a single level.
  • Complete rows: Rows with a missing answer are excluded and the remaining number of observations appears in the result; no diagnosis is offered for why they went missing.
  • Enough mixed patterns: No separate warning is raised about the number of participants with mixed patterns. Judging how far the distribution is stacked at the extremes is something you do by reading the per-column counts.

How the analysis is run

  1. 1Drop the file onto the panel; a CSV or XLSX uploaded by drag and drop opens in the grid on the Data tab.
  2. 2Look over every binary column in the grid. Where unticked options were left blank, you have to decide whether those blanks mean no or mean missing.
  3. 3On the Variable tab set the measurement level and the missing-value codes for the columns; if you want blanks read as no, make that explicit with a derived column.
  4. 4On the Analysis tab pick the method from the categorical group; the data requirements box shows which columns satisfy the binary condition.
  5. 5In the parameter form tick the three or more columns to be compared and run the analysis.
  6. 6Read the per-column counts of yes answers first, then the Q value, the degrees of freedom and the p value in the collapsible sections; no chart is produced for this method.
  7. 7If Q comes out significant, plan the pairwise comparisons as separate runs, export the tables to Excel, and use the history tab to keep track of which set of columns you used.

Statistics and tables produced

The Q statistic
The test statistic that summarises how far the columns' counts of yes answers lie from one another, scaled by each participant's own pattern of answers.
Degrees of freedom
The number of columns compared, minus one. For a list of four items the degrees of freedom come to three.
Significance value
The p value from the chi-square distribution under the hypothesis that every column has the same acceptance rate. The verdict is overall and does not identify which column stands apart.
Per-column counts or proportions of yes answers
How many participants ticked each item, given as a proportion where that reads better. This is the descriptive material you pass to a reader when you describe the finding.
Observations entering the analysis
The number of participants with a valid answer in every column; as the item count grows, it is normal for this to drift away from the row count of the file.

Effect size and confidence intervals

No effect size is computed
No coefficient expressing the size of the separation between columns is returned. To convey how much the finding weighs you report the per-column rates, the gap in percentage points between the highest and the lowest, and the number of participants analysed. A measure summarising agreement across repeated measures has to be computed outside the panel.

No confidence interval is returned for this method. The Q value, the degrees of freedom and the p value come as point results, with no interval for the columns' proportions or for the differences between them. An interval estimate for a single item's proportion can be had by running the proportion test separately in its one-sample mode.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
In an illustrative public services study, do the usage rates of three digital channels put to the same participants differ?
Variables
First column: use of the mobile application (yes, no) · Second column: use of the website (yes, no) · Third column: service obtained by telephone (yes, no)
Example result
150 participants with an answer in all three columns entered the analysis. Yes counts came to 96 for the mobile application, 78 for the website and 61 for the telephone, putting the rates at 64.0, 52.0 and 40.7 per cent respectively. Of the participants, 40 use all three channels, 45 use two, 25 use only one and 40 use none. The test returned Q = 26.26 on 2 degrees of freedom, p < 0.001.
Interpretation
The three channels are not used at the same rate: a spread of 23.3 percentage points separates the highest rate from the lowest, and by the p value such a spread would rarely arise by chance. What the result withholds is which pair is responsible. Whether the 12-point gap between the application and the website is itself meaningful cannot be read off Q. Seeing that requires paired tests comparing the channels two at a time as separate runs, with the significance level adjusted for the three comparisons. Note also that the 40 people using all three channels and the 40 using none contribute nothing to Q, so the separation rests on the mixed patterns of the remaining 70 participants. The figures are illustrative.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

General customer survey (synthetic)

A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.

Rows
300
Columns
respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: Did the usage rate stay constant across all three measurements?

Analyzed columns
use_before,use_after,use_followup
q statistic
89.31
Degrees of freedom
2
p value
< 0.001
Valid observations
300
Significant
Yes

Success rates

use_before
0.303
use_after
0.553
use_followup
0.587

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914

How to report the result

Usage rates differed significantly across the three digital channels, Cochran's Q(2, N = 150) = 26.26, p < .001, with usage at 64.0 per cent for the mobile application, 52.0 per cent for the website and 40.7 per cent for the telephone.

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • No follow-up analysis is returned: a significant Q says at least one column stands apart without showing which pair is behind it.
  • When pairwise comparisons are assembled by hand, the multiple-comparison correction has to be applied by hand too, since an uncorrected sweep raises the false positive risk appreciably even across three items.
  • Participants answering the same way to every item contribute nothing, so the test is less sensitive than expected on lists where answers pile up at the extremes.
  • As the number of items grows, the requirement for complete rows narrows the sample, and on long lists the number of people analysed can fall well below the row count of the file.
  • With no effect size and no confidence interval returned, a reporting table cannot be completed from the statistic and the p value alone.
  • The items need to be conceptually comparable: even when all of them are binary in form, equating the acceptance rates of items that ask about different things leaves the interpretation empty.

What to use when the assumptions are not met

  • McNemar TestWhy: When the number of paired binary columns comes down to two, or when pairs are to be tested one at a time after a significant Q.
  • Friedman TestWhy: When the repeated measures are ordinal or continuous rather than binary; it compares three or more measurements through ranks.
  • Chi-Square Test of IndependenceWhy: When the sets being compared are independent groups rather than the same people; it addresses the association between two categorical variables.
  • Multiple Response AnalysisWhy: When describing how often each option in a multiple response question was ticked comes before any significance test.

Frequently asked questions

Q is significant. How do I find out which item stands apart?
The panel runs no follow-up for this method: pairwise comparisons are not carried out for you and no corrected p values are produced. The practical route is to compare the column pairs you care about with the paired binary test in separate runs, then apply a correction that accounts for how many comparisons you made. Remember that three items yield three possible pairs and four items yield six; the count climbs quickly, and every additional run counts against the free plan's monthly limit of fifty.
Why do participants who say yes to everything have no influence?
The test judges the standing of the columns relative to one another inside each participant. Someone who answers yes to every item separates no column from any other and therefore says nothing about differences between columns, and the same holds for anyone answering yes to none. This is why the count of mixed-pattern participants deserves reading alongside the per-column totals: if most of the sample sits at the extremes, Q rests on a small subset.
In a multiple response question, can I treat unticked options as no?
That is a data preparation decision and it bears directly on the result. Coding a blank as no is defensible when the participant saw the whole list and leaving an option unticked was a deliberate choice; where the list was split across pages or some respondents skipped the question, the same blank is a missing value. Make the decision explicit with a derived column on the Variable tab, because treating blanks as missing narrows the sample through the complete-rows requirement, while treating them as no pulls the acceptance rates down.
How many items can I compare?
The constraint lies not in the item count but in how many participants have an answer in every column. Degrees of freedom equal the item count minus one, so power spreads thin on long lists while the complete-rows requirement narrows the sample at the same time. The longer the list, the more the result turns into an overall verdict that is hard to interpret. Grouping conceptually close items into separate runs produces a more readable report than equating seven or eight items in one test.

References

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.