Categorical

Multiple Response Analysis

Reads questions where a respondent may tick several options from the option columns that hold them, and reports a count per option alongside two percentages built on two different bases.

Method summary

A check-all-that-apply question cannot be tabulated like a single-choice question, because one person may tick nothing at all while another ticks five options, which means the number of ticks and the number of people are two different quantities. Data of this kind is usually stored as one column per option, each column recording whether that option was selected. Multiple response analysis takes those columns together, reduces each to selected or not selected, and brings the options into one table. That table carries two percentage columns resting on separate bases: percent of responses gives an option's share of all the ticks, while percent of cases gives the share of respondents who ticked it. The first column adds to one hundred, the second does not and climbs above one hundred as the average number of selections per person rises. YouReply Analyze reports both percentages together with the selection counts, the number of respondents and of responses, the mean selections per respondent and the number of people who ticked nothing, and repeats the whole breakdown group by group when a grouping column is supplied.

Which research questions does it answer?

  • Which news sources did respondents use in the past month, and where does the weight sit across those sources?
  • Which product attributes were ticked as reasons to buy, and how many people gave more than one reason?
  • How many brands does the average respondent recall from an awareness list?
  • How do the contact channels people tick vary across age groups?

When should you use it?

  • When the question permits several options to be ticked and the answers sit in one column per option.
  • When a single table should show both the relative weight of the options and the share of people each option reached.
  • When you need to know how many options the average respondent ticked and how many found nothing on the list.
  • When describing the pattern of ticks across the levels of a grouping variable is enough and the differences need no test.
  • Not needed where only one option can be ticked, since a frequency breakdown answers that case.

Required variable types

  • Option columns: at least two must be supplied, each column recording whether one particular option was ticked.
  • In a numeric option column any value above zero counts as selected, while zero and blank count as not selected.
  • In a text option column only a limited set of truthy tokens is recognised: 1, true, evet, yes and x, along with their case variants. Anything else reads as not selected.
  • Grouping column (optional): a single nominal column. When supplied, the breakdown is repeated for each of its categories.
  • One row per respondent; data that packs the selections into a single comma-separated field has to be split into columns first.

Key assumptions

Options held in separate columns
The method expects the layout in which each option has a column of its own. Data that keeps every selection in one cell separated by commas does not match that layout.
Selection coding in a recognised form
Whether a cell counts as a selection depends on its content matching the binarisation rule. Numeric coding is safe in this respect, while text coding requires the spelling to be one of the recognised tokens.
Options belonging to one question
The table is read on the understanding that the columns supplied are the options of a single question. Columns assembled from different questions make the base of the percent of responses column conceptually meaningless.
What ticking nothing means
Ticking nothing on a list is not necessarily the same as skipping the question. The first says the list found no purchase with that person, the second says the answer is missing, and only the questionnaire design separates them.

How YouReply checks these assumptions

  • Options held in separate columns: The number of columns you supply is checked and the analysis will not start with fewer than two. Nothing recognises or splits selections packed into a single cell, so the splitting is yours to do on the Data and Variable tabs.
  • Selection coding in a recognised form: THE PANEL REPORTS NOTHING ABOUT AN UNRECOGNISED SPELLING. A value such as ticked, present, checked or true written in another language is quietly treated as not selected and the counts come out low. Inspect the values held in these columns on the Data tab before you run, and convert them to a numeric derived column on the Variable tab where needed.
  • Options belonging to one question: Whether the columns come from one question cannot be verified; which columns you bring together is your decision and belongs in the report alongside the question wording.
  • What ticking nothing means: The number of respondents who ticked nothing is reported as a field of its own, but it is not distinguished whether that was a deliberate answer or a missing one. The base for the percent of cases column excludes these people.

How the analysis is run

  1. 1Drop the file onto the panel and confirm on the Data tab that all of the option columns came through.
  2. 2Scan the values inside the option columns in the grid, since an unrecognised selection spelling has to be spotted at this point.
  3. 3On the Variable tab give the option columns readable labels, because the result table is written with the column names, and build numeric derived columns where needed.
  4. 4Pick this method from the categorical group on the Analysis tab; the data requirements box restates the two-column minimum.
  5. 5In the parameter form tick every option column belonging to one question and decide whether you want a breakdown.
  6. 6Run the analysis. The result presents the option table and the overall counts as collapsible sections, and no chart is drawn for this method.
  7. 7While reading the table, keep track of which percentage column you are looking at, and note how many respondents ticked nothing.
  8. 8Export the breakdown to Excel; if you supplied a grouping column, the group tables sit in the same file.

Statistics and tables produced

Selection count per option
How many people each option counted as selected for after binarisation. Both percentage columns derive from this number.
Percent of responses
An option's share of all the ticks placed by all respondents. This column adds to one hundred apart from rounding and expresses the relative weight of the options.
Percent of cases
The share of respondents with at least one selection who ticked this option. The column does not add to one hundred; its total exceeds one hundred by the mean number of selections per person.
Respondents and responses
The two bases of the breakdown are printed separately: how many people entered the analysis and how many ticks were placed in total. They let you verify where each percentage column comes from.
Mean selections per respondent
Total ticks over the number of respondents. It shows how generously the list was ticked and predicts how far the percent of cases column will overshoot one hundred.
Respondents who ticked nothing
How many respondents selected none of the options on the list. They do not appear in the base for the percent of cases column.
Breakdown by group
When a grouping column is supplied, the selection count and the percent of cases are listed for each group. Percent of responses does not appear in the group tables.

Effect size and confidence intervals

The base of percent of responses
No effect size is computed here, so the weight of the reading falls on keeping the two bases apart. Percent of responses is based on the total number of ticks, which makes an option's share dependent on how heavily the other options were ticked. Adding an option to the list or removing one shifts every figure in the column, which is why it cannot be compared across studies that used different lists. The column totals one hundred.
The base of percent of cases
Percent of cases is based on the number of people with at least one selection, so an option's figure is independent of the other options and states directly how many people it reached. Questions about reach are answered from this column. Its total grows with the mean number of selections: an average of two options means a total of roughly two hundred per cent, and that is not an error.
Mean selections per respondent
It falls between one and the number of options and summarises how the list was ticked. An average sitting barely above one suggests respondents handled the question as though only one answer were allowed, which can point to unclear wording. An average approaching the number of options says the list has lost its power to discriminate, since almost everybody ticked almost everything.

No confidence interval is returned by this method. The selection counts and both percentage columns arrive as single values, with no interval around an option's proportion. Nor is any significance test run: no p value and no effect size is produced for the gap between two options or for an option's difference across the groups of a breakdown. To test whether a selection rate really varies between groups, take that option's binary column and the grouping variable into a separate method.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
In an illustrative brand awareness study, respondents were shown five energy drink brands and asked to tick every brand they had heard of.
Variables
Option columns: five separate selection columns, one per brand · Grouping column: age band (18-29, 30-44, 45 and over)
Example result
500 rows were read and 38 respondents ticked no brand at all, leaving 462 people as the base for percent of cases. The ticks totalled 996, a mean of 2.16 selections per person. The brands came out as follows: the first brand 331 ticks, 33.2 per cent of responses and 71.6 per cent of cases; the second 258 ticks, 25.9 and 55.8; the third 197 ticks, 19.8 and 42.6; the fourth 124 ticks, 12.4 and 26.8; the fifth 86 ticks, 8.6 and 18.6. In the age breakdown the fifth brand reached 29.4 per cent of cases in the youngest group against 7.1 per cent in the oldest.
Interpretation
The response percentages add to one hundred within rounding and place the weight of the list on the first two brands. The case percentages add to 215.4, which is no inconsistency but the arithmetic consequence of 2.16 selections per person. The two columns answer two questions: the first brand collects a third of all the ticks and at the same time reaches about seventy two per cent of the people who recalled at least one brand. The 38 respondents who ticked nothing are what separates awareness read over the whole sample from awareness read over those who recalled something. The 22-point spread in the age breakdown is a descriptive observation, and this method does not test it. The figures are illustrative.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

Variable examined

Market research survey (synthetic)

Two hundred and fifty respondents: the brand in use, three demographic breaks, six choose all that apply option columns and six purchase intent questions. The option columns overlap, so the extra reach an item brings can genuinely be measured.

Rows
250
Columns
respondent_id, gender, region, age_band, main_brand, opt_price, opt_quality, opt_service, opt_brand, opt_speed, opt_design, buy_a, buy_b, buy_c, buy_d, buy_e, buy_f
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: What share of respondents marked each feature and how many features did each person select?

Analyzed columns
opt_price,opt_quality,opt_service,opt_brand,opt_speed,opt_design
Number of respondents
232
Total selections
572
Mean selections per respondent
2.47
Respondents with no selection
18

Options

Options
RowCountPercent of responsesPercent of cases
opt_price13022.7356.03
opt_quality11219.5848.28
opt_service9115.9139.22
opt_brand7513.1132.33
opt_speed9015.7338.79
opt_design7412.9431.90

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260916

How to report the result

Across a five-brand awareness list, the 462 respondents recalling at least one brand placed 996 ticks in total (a mean of 2.16 each), with the first brand taking 33.2 per cent of responses and reaching 71.6 per cent of cases.

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • The two percentage columns are not interchangeable, and reading percent of responses as reach understates the share of people an option actually touched.
  • Adding an option changes every response percentage, so two measurements taken with different option lists cannot be compared through that column.
  • An unrecognised selection spelling in a text-coded column is quietly treated as not selected and raises no error, which can leave the counts too low.
  • Bias arising from the order in which options appeared on the questionnaire passes straight into the table, as options near the top tend to attract more ticks and the breakdown does not correct for it.
  • Group tables carry no test, confidence interval or effect size for differences between groups, so any visible spread remains descriptive.
  • The pattern of co-selection, meaning which pairs of options tend to be ticked together, is invisible in this breakdown; option sets that build reach cumulatively need a different method.

What to use when the assumptions are not met

  • Frequency DistributionWhy: When the question allows only one option to be ticked; the category breakdown of a single column is more direct.
  • Banner TableWhy: When the question is whether an option's selection rate differs significantly across the groups of a breakdown, since it marks significance with letters.
  • TURF AnalysisWhy: When the question is which set of options together reaches the widest audience, since it accounts for overlap in computing cumulative reach.
  • Cross-TabulationWhy: When the joint distribution of a single option's binary column and a nominal variable should be examined with row and column percentages.

Frequently asked questions

My case percentages add to more than one hundred. Is something wrong?
Nothing is wrong. The base of percent of cases is the number of people with at least one selection rather than the number of ticks, and because one person can tick several options the same person is counted on several rows. The mean number of selections tells you roughly where the total will land: at 2.16 selections the column totals around two hundred and fifteen per cent. The column that adds to one hundred is percent of responses, and stating in your report which column you published keeps readers from confusing the two bases.
My option columns hold yes and no. Will the breakdown be right?
Yes is among the recognised truthy tokens and so counts as selected, while no is not recognised and counts as not selected, which is the correct outcome. The trouble lies elsewhere: a value such as ticked, present or checked reads silently as not selected and no warning appears. The safest route is to inspect the distinct values in those columns on the Data tab before running and, if in any doubt, recode the columns to zero and one.
Should I drop the respondents who ticked nothing from the base?
They are already outside the base for percent of cases and their number is reported separately. The decision depends on your questionnaire: if finding nothing on the list is a meaningful answer, you may want to report awareness over the whole sample as well, and in that case the percentages have to be recomputed by hand. People who simply skipped the question sit inside the same number, so separating them requires your own fieldwork records.
Is an option ticked significantly more often in one age group?
This method does not say. The breakdown gives a selection count and a percent of cases for each group, but no test is run between groups and no effect size is produced. Where significance is required, take that option's binary column together with the grouping variable into a banner table or a categorical test of independence. Bear in mind too that the group tables carry only counts and percent of cases, so percent of responses is not available there.

References

  • Agresti, A. (2018). An Introduction to Categorical Data Analysis
  • Malhotra, N. K. (2019). Marketing Research: An Applied Orientation
  • Oppenheim, A. N. (1992). Questionnaire Design, Interviewing and Attitude Measurement
  • pandas.DataFrame.groupby documentation

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.