Choice-based
MaxDiff
Turns a design that asks respondents to mark the best and worst option in each small set into a preference score and a rank for every item.
Method summary
Rather than asking people to rate dozens of items one at a time, MaxDiff shows four or five at once and asks two questions of each screen: which of these matters most to you, and which matters least. The design arranges the screens so that every item appears several times and in different company, which compares the items against one another indirectly. Its appeal is that it sidesteps scale use: the habit of rating everything highly, or of clustering answers in the middle, has no purchase here, because each screen forces a trade-off. The counting approach then subtracts how often an item was picked as worst from how often it was picked as best and divides by how often it was shown, which makes items comparable. YouReply Analyze reports that score, a preference share rescaled to a hundred, the rank and how evenly the design was built. The scores come straight from the counts rather than from a fitted choice model.
Which research questions does it answer?
- Which of twenty candidate product features genuinely stand out for the target audience?
- Which three of twelve benefit statements should lead the communication?
- Which criteria do business buyers treat as least important when choosing a supplier?
- How can a long list of priorities be ordered without everyone calling everything important?
When should you use it?
- When many items need to be put in priority order and one-by-one rating fails to separate them.
- When a tendency to rate every statement highly leaves the ordering useless.
- When a genuine trade-off exists between the items, because not all of them can be chosen.
- In studies where scale use differs between cultures or between subsamples.
- When the design is already fielded and the data were collected in long format, one row per item shown.
Required variable types
- Respondent column: an identifier separating each person, tying together all rows belonging to them.
- Set column: an identifier for the screen or task, numbered within each respondent.
- Item column: the name or code of the item shown on that row.
- Choice column: a value marking whether the item on that row was picked as best, as worst, or not at all.
- Which values stand for best and worst is declared in the parameter form, with the strings best and worst as the defaults.
Key assumptions
- Long-format data layout
- Each row has to stand for a single item shown to a single respondent in a single task. A wide layout that squeezes a whole task into one row cannot be read here, because times shown are counted off the rows.
- A balanced design
- Every item should appear roughly the same number of times and meet different companions. An item shown markedly less often rests on fewer observations and its place in the ordering becomes unstable.
- Exactly one best and one worst per task
- Several items marked best in one task, or none marked at all, signals a problem in the questionnaire flow and corrupts the counts.
- Correctly declared best and worst values
- Which value in the choice column means best and which means worst has to be declared. A swapped mapping turns the ordering upside down while the computation proceeds smoothly.
- Items pitched at the same level
- Putting a sweeping statement up against a narrow detail hands the sweeping one an unearned advantage. The internal consistency of the list is a design matter rather than a statistical one.
How YouReply checks these assumptions
- Long-format data layout: The panel does not reshape data and will not offer to pivot a wide file into long format; that conversion belongs before the upload.
- A balanced design: The design balance section returns the smallest and largest number of times shown, the ratio between them, and a balanced flag stating whether that ratio clears 0.8. The flag informs and does not halt the analysis.
- Exactly one best and one worst per task: No separate check or warning is produced for the number of choices per task. Verifying the expected count for every respondent and task is work to do before uploading.
- Correctly declared best and worst values: The defaults are the strings best and worst; if your data use other labels, typing them into the parameter form is your responsibility and a typo passes in silence.
- Items pitched at the same level: Nothing assesses the scope or the wording level of the items; that judgement is left to the researcher.
How the analysis is run
- 1Drop the long-format file onto the panel, with each row carrying a single item seen by a single respondent in a single task.
- 2In the grid on the Data tab check how many rows belong to the same respondent and task and that the choice column carries the labels you expect.
- 3On the Variable tab mark the respondent, task and item columns as nominal and tidy up spelling differences in item names, since two spellings of one item count as two items.
- 4On the Analysis tab pick MaxDiff from the choice-based group; with Only ones that fit my data switched on, unsuitable methods dim and state why.
- 5In the parameter form map the four columns and type the values that stand for best and worst in the choice column.
- 6Run the analysis. The item table and the design balance section arrive as collapsible headings.
- 7Open the design balance section before reading the item table, and treat the lower half of the ordering cautiously if the balanced flag is negative.
- 8Export the result to Excel; the computation credits card names the library call and its version, and the history tab keeps earlier runs.
Statistics and tables produced
- Times shown
- How often each item reached the screen across all respondents and tasks. Comparability of the scores depends on this figure being similar across items.
- Best and worst choice counts
- How often the item was marked most important and how often least important, listed separately; these two counts are the raw material of the score.
- Best-worst score
- Worst choices subtracted from best choices and divided by times shown. It runs between minus one and plus one, with zero meaning an item was praised and dismissed equally often.
- Preference share
- Produced by exponentiating the best-worst score and rescaling so the items add to a hundred. It is a transformation that aids reading, not an estimated choice probability.
- Rank
- Items are ordered by their best-worst score and each receives a rank number.
- Design balance
- The smallest and largest number of times shown, the ratio between them and a balanced flag reporting whether that ratio clears the threshold.
Effect size and confidence intervals
- Best-worst score
- The principal quantity this method yields, and it comes straight from counting: the difference between two tallies divided by times shown. Positive values say the item was chosen more often than discarded, negative values the reverse. No settled threshold exists for its magnitude, so interpretation runs through comparison with the other items on the list. It is not a utility parameter and has no per-person value.
- Preference share
- Obtained by exponentiating the scores and scaling them to add to a hundred, which restates the gaps between items in proportional terms. An item's share is not the probability that a choice model would assign to it, and with no model fitted it cannot be read that way. On a list of twelve items an even spread corresponds to about 8.3 per cent, and shares are read against that anchor.
- Ratio of times shown
- The smallest number of times shown over the largest summarises how even the footing beneath the ordering is. A ratio of 0.8 or above counts as a balanced design. This is not an effect size but an indicator of how trustworthy the scores are.
No confidence interval is returned for any quantity and no standard error is computed. That follows from the approach rather than from a preference: with nothing estimated, no measure of uncertainty arises around which an interval could be built. Whether two items sitting next to each other in the ordering genuinely differ cannot be settled from this output, and treating adjacent scores as equivalent is the more defensible reading.
Example research question and example result
The numbers below are a representative example, not data from a real study or a real user.
- Research question
- In an illustrative enterprise software study, which of twelve benefit statements should lead the communication?
- Variables
- Respondent column: person identifier across a sample of 400 · Task column: 8 tasks per person, 4 statements per task · Item column: 12 benefit statements · Choice column: best, worst or blank
- Example result
- A total of 12,800 rows were analysed. The fast deployment statement was shown 1,072 times, chosen best 512 times and worst 41 times: best-worst score 0.439, preference share 12.9 per cent, rank 1. The flexible reporting statement in the middle of the list was shown 1,065 times and, with 268 best and 252 worst choices, scored 0.015 for a share of 8.5 per cent and rank 6. The brand recognition statement at the bottom was shown 1,058 times, with 27 best and 498 worst choices, scoring minus 0.445 for a share of 5.3 per cent. The design balance section reported a minimum of 1,041, a maximum of 1,089, a ratio of 0.956 and a balanced design.
- Interpretation
- The two ends of the list pull clearly apart: fast deployment is chosen best in roughly one of every two tasks where it appears and is almost never discarded, while brand recognition is mostly discarded when it appears. A score near zero for the middle statement means it was both praised and dismissed, so the audience is split on it rather than mildly interested, and opposing tendencies may be cancelling out. Preference shares are read against the even-spread anchor of 8.3 per cent, which puts the leading statement at about one and a half times that. Because the shares come from no choice model, they cannot be read as purchase probabilities, and nothing quantifies the small gap between ranks five and six. A positive balance flag says the ordering rests on an equal number of observations per item. The figures are illustrative.
Real output on a sample dataset
The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.
Choice experiment, long format (synthetic)
One hundred and twenty respondents, six tasks each and four options per task: 2,880 rows in total. One row is one exposure. Each task marks one best and one worst option and has exactly one chosen option; brand, price and size vary by exposure. MaxDiff and conjoint require this layout, and the panel does not reshape data into it.
- Rows
- 2,880
- Columns
- respondent, task, item, choice, chosen, brand, price, size
The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.
Research question: Which of the six purchase criteria matters most to respondents?
- Respondent identifier column
- respondent
- Choice set column
- task
- Item column
- item
- Best and worst choice column
- choice
- Value marking the best pick
- best
- Value marking the worst pick
- worst
- Number of respondents
- 120
- Number of choice sets
- 6
- Number of items
- 6
Items
| Item | Times shown | Times chosen best | Times chosen worst | Best minus worst score | Preference share (sums to 100) | Rank |
|---|---|---|---|---|---|---|
| Kalite | 490 | 243 | 24 | 0.447 | 24.68 | 1 |
| Fiyat | 486 | 231 | 32 | 0.409 | 23.77 | 2 |
| Hiz | 474 | 110 | 96 | 0.030 | 16.26 | 3 |
| Marka | 483 | 76 | 134 | -0.120 | 14.00 | 4 |
| Destek | 478 | 40 | 197 | -0.328 | 11.36 | 5 |
| Tasarim | 469 | 20 | 237 | -0.463 | 9.94 | 6 |
Design diagnostics
- Display balance ratio
- 0.957
- Minimum times shown
- 469
- Maximum times shown
- 490
- Design is balanced
- Yes
Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260917
How to report the result
Best-worst scaling placed fast deployment first (1,072 exposures, 512 best, 41 worst, score 0.439, preference share 12.9 per cent), with a design balance ratio of 0.956 (N = 400, 8 tasks each).
An example sentence close to APA style; the numbers are representative.
When you should not use it
- No individual-level utilities are produced; each item carries one sample-level score, which rules out per-person preference patterns and segment building.
- No hierarchical Bayes or conditional logit estimation is applied, so preference shares cannot be reported as model estimates.
- With no standard errors, intervals or significance tests, two adjacently ranked items cannot be shown to differ.
- The data must arrive in long format and the panel will not convert a wide file, so reshaping falls before the analysis.
- Scores mean something only relative to the items on the list, and adding or removing one item changes every share.
- What is measured is where an item stands against the others rather than how important it is, so no absolute level of importance or acceptability threshold emerges.
What to use when the assumptions are not met
- Conjoint AnalysisWhy: When the items are feature levels of a product rather than free-standing statements, handling trade-offs within and across attributes together.
- TURF AnalysisWhy: When the question is how many people are covered rather than how items rank, searching for the set that reaches the most people.
- Multiple Response AnalysisWhy: When no trade-off design is possible and only a multiple-choice question is available, reporting tick rates directly.
- Descriptive StatisticsWhy: When the items were each rated separately, since describing them with means and spread needs no trade-off design.
Frequently asked questions
- Is the preference share the probability that an item gets chosen?
- It is not. The share comes from exponentiating the best-worst score and scaling the results to add to a hundred, with no fitted choice model behind it. A value of twelve per cent therefore does not mean a respondent would pick that item twelve per cent of the time. The purpose of the share is to make the gaps between scores readable on a common total. Genuine choice probabilities would require estimating a choice model, which this panel does not do.
- Why are there no per-person utility values?
- Because the arithmetic rests on counting. One score is produced per item at the sample level, summing the choices of all respondents. Getting a value per person calls for an estimation method such as hierarchical Bayes, and with none applied there are no individual utilities, no segmentation and no simulation. If you need to compare segments, the practical route is to filter the data by segment, run the analysis once per segment and read the orderings side by side.
- The balance flag came back negative. Can I still use the result?
- The flag reports rather than blocks. When the ratio of the smallest to the largest number of exposures falls below 0.8, some items carry scores resting on fewer observations than others. The top of the list usually survives that, while the middle and lower ranks become unstable. Read with the times-shown column in view, avoid treating the rank of a rarely shown item as a finding, and where possible repair the design and collect again.
- How do I convert my data from wide to long format?
- The panel will not do this conversion. The layout needed is this: one row per item shown to one respondent in one task, so a task of four items yields four rows, one marked best, one marked worst and two left blank. If your survey tool writes a whole task on one row, you have to unstack those rows, moving the item name into one column and the choice information into another. Once converted, a good check is that a respondent's row count equals the number of tasks times the number of items per task.
References
- Louviere, J. J., Flynn, T. N., & Marley, A. A. J. (2015). Best-Worst Scaling: Theory, Methods and Applications
- Flynn, T. N., Louviere, J. J., Peters, T. J., & Coast, J. (2007). Best-Worst Scaling: What It Can Do for Health Care Research and How to Do It
- Orme, B. K. (2020). Getting Started with Conjoint Analysis
- numpy documentation
Try it with your own data
The free plan includes 50 analysis runs a month and needs no card.