Choice-based
Conjoint Analysis
Turns a design in which respondents pick one of several product profiles per task into part-worth utilities for feature levels and relative importance for the features.
Method summary
Choice-based conjoint stops asking which feature matters and makes people choose instead. Each task presents a handful of profiles that combine different levels of attributes such as price, delivery frequency or warranty length, and the respondent picks one. Because the attributes vary systematically across tasks, the pattern of choices exposes trade-offs that nobody states out loud, such as whether a longer wait is accepted in exchange for a lower price. The value of the exercise lies in weighing features against each other rather than in measuring how much each is liked on its own. YouReply Analyze recovers those weights by counting: how often every level appeared and how often a profile containing it was picked are both known, that ratio becomes a part-worth utility, and the importance of an attribute follows from how far its utilities spread. The approach is deliberately count-based, with no choice model fitted.
Which research questions does it answer?
- Does price or delivery frequency weigh more heavily in a subscription package?
- Improving which feature level lifts the choice rate the most?
- How much extra are respondents willing to accept in price for a longer warranty?
- Which of the four features on the product card barely moves the decision at all?
When should you use it?
- When a product or service is built from several discrete features whose relative weight is in question.
- When stated importance rankings fail to separate anything because everybody calls everything important.
- When the design is already built, the tasks and profiles planned, and fieldwork complete.
- When the data were collected in long format, one row per profile shown.
- When the goal is the ordering of feature weights rather than precise utility parameters and simulation.
Required variable types
- Respondent column: a person identifier tying together all the tasks belonging to one person.
- Task column: an identifier separating tasks within a respondent, since the choice happens inside that unit.
- Choice column: a binary value saying whether the profile on that row was picked.
- Attribute columns: one column per attribute, holding the level that profile carried, stored as categorical labels.
- Numeric attributes such as price are handled as levels too, so with no continuous coefficient estimated, price is read as a list of categories that happen to be numbers.
Key assumptions
- Long-format data layout
- Each row has to represent one profile shown to one respondent in one task, so a task of three profiles yields three rows with one of them marked as chosen.
- Exactly one choice per task
- Every task stands for one decision, so a count of chosen profiles other than one means the task was recorded incompletely or something went wrong in the flow.
- A balanced design with no confounding
- Feature levels should appear about equally often and vary independently of one another. If two attributes move together across tasks, their effects cannot be separated.
- Consistent spelling of feature levels
- Two spellings of one level count as two levels and can inflate the spread of utilities within an attribute, and with it that attribute's importance percentage.
- Plausible and feasible profiles
- Profiles that pair the best features with the lowest price, or that could not exist in reality, distort choices through the design itself and inflate the apparent spread of utilities.
How YouReply checks these assumptions
- Long-format data layout: The panel will not pivot a wide file into long format and offers no such suggestion; the layout has to be right before the upload.
- Exactly one choice per task: The design section reports how many tasks lack exactly one chosen profile. That is an informational line: those tasks are not separately excluded, and judging a large number is left to you.
- A balanced design with no confounding: The design section reports the smallest and largest number of profiles per task. No measure is produced for how evenly levels appear or for dependence between attributes, so comparing the exposure counts in the level table is your job.
- Consistent spelling of feature levels: Consistency of spelling is not checked. Scanning level names in the grid on the Data tab and merging them through value labels on the Variable tab falls to the researcher.
- Plausible and feasible profiles: Nothing assesses whether the profiles are realistic; prohibited combinations have to be prevented when the design is built.
How the analysis is run
- 1Drop the long-format file onto the panel, with each row carrying a single profile seen by one respondent in one task.
- 2In the grid on the Data tab check the number of rows inside each task and that the choice column arrives as zeros and ones.
- 3On the Variable tab mark the attribute columns as nominal and unify spelling differences in level names through value labels.
- 4On the Analysis tab pick conjoint analysis from the choice-based group; with Only ones that fit my data switched on, unsuitable methods dim and state why.
- 5In the parameter form map the respondent, task and choice columns, then tick the attribute columns to include.
- 6Run the analysis. The level table, the attribute importances and the design check arrive as separate collapsible sections.
- 7Open the design section before reading the importance percentages, and go back to the data if the count of tasks without exactly one choice is high.
- 8Export the result to Excel; the computation credits card names the library call and its version, and the history tab keeps earlier runs.
Statistics and tables produced
- Times shown and times chosen per level
- How many profiles carried each level and how many of those were picked, listed separately; the whole calculation rests on these two counts.
- Choice rate
- Times chosen over times shown for a level. It reads directly, but takes no account of which rival profiles stood beside it in the task.
- Part-worth utility
- The logarithm of the chosen-to-shown ratio with a half-count correction, centred so that the levels of an attribute average zero. Levels are then read around zero with positive and negative values.
- Attribute importance
- The spread between the highest and lowest utility within each attribute, with those spreads normalised to add to a hundred.
- Design check
- The smallest and largest number of profiles per task, plus the number of tasks that do not have exactly one chosen profile.
Effect size and confidence intervals
- Part-worth utility
- Shows how a feature level fares against the average level of the same attribute, with positive values sitting above that average and negative values below. It is produced by counting: a logarithmic transformation followed by centring within the attribute, rather than an estimated coefficient. Differences between levels carry meaning inside one attribute, and rather than racing utilities across attributes you should turn to the importance percentages.
- Attribute importance
- The distance between an attribute's best and worst level expresses its weight in the decision, and the percentages add to a hundred. An importance above fifty per cent says one feature largely drives the decision. These percentages also come from counting rather than from a model, and they depend on the design: widening the range of levels in an attribute mechanically enlarges its share.
- Choice rate
- The plainest description available, being the proportion of profiles containing a level that were picked. It is easy to read and can mislead: a level shown mostly against weak rivals comes out higher than it deserves. A balanced design shrinks that bias without removing it.
No confidence interval is returned for any part-worth or importance percentage and no standard error is computed. With nothing estimated in a counting approach, there is no uncertainty around which an interval could be built. Nothing therefore shows that the gap between two importance percentages, or between two levels' utilities, exceeds sampling variability. Treating two close importance percentages as equivalent is sounder than ranking them.
Example research question and example result
The numbers below are a representative example, not data from a real study or a real user.
- Research question
- In an illustrative coffee subscription study, which feature drives choice most among price, delivery frequency and bean sourcing?
- Variables
- Respondent column: person identifier across a sample of 350 · Task column: 10 tasks per person, 3 profiles per task · Attributes: monthly price (79, 99, 129 lira), delivery frequency (weekly, fortnightly, monthly), bean sourcing (single origin, blend)
- Example result
- A total of 10,500 profile rows were analysed. Among price levels, 79 lira appeared 3,512 times and was chosen 1,596 times (rate 0.454, utility 0.372); 99 lira took 3,498 exposures and 1,211 choices for a rate of 0.346 and a utility of minus 0.100; 129 lira took 3,490 exposures and 681 choices for a rate of 0.195 and a utility of minus 0.473. For delivery frequency, weekly returned 0.167, fortnightly minus 0.074 and monthly minus 0.243. For sourcing, single origin returned 0.067 and blend minus 0.067. With spreads of 0.845 for price, 0.410 for delivery frequency and 0.134 for sourcing, the importances normalised to 60.9, 29.5 and 9.6 per cent. The design section reported a minimum and maximum of 3 profiles per task and 12 of the 3,500 tasks without exactly one chosen profile.
- Interpretation
- Price dominates the decision, weighing more than the other two features combined. Its utilities fall steadily from 79 to 129 lira, meaning respondents move away consistently as price rises, and that regularity can be taken as a sign the design behaved sensibly. In delivery frequency the step from weekly to fortnightly is smaller than the step from fortnightly to monthly, so the real break comes with monthly delivery. Sourcing landing last at 9.6 per cent says its presence on the product card does not visibly change the decision. Because the counts ignore which profiles competed inside a task, the utilities cannot be read as coefficients of a conditional logit model; the ordering of the features would usually agree with such a model while the values would not. The 12 tasks without exactly one choice amount to less than three in a thousand and cannot carry the result, though their origin still deserves a look. The figures are illustrative.
Real output on a sample dataset
The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.
Choice experiment, long format (synthetic)
One hundred and twenty respondents, six tasks each and four options per task: 2,880 rows in total. One row is one exposure. Each task marks one best and one worst option and has exactly one chosen option; brand, price and size vary by exposure. MaxDiff and conjoint require this layout, and the panel does not reshape data into it.
- Rows
- 2,880
- Columns
- respondent, task, item, choice, chosen, brand, price, size
The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.
Research question: How is the weight of brand, price and size distributed in the choice decision?
- Respondent identifier column
- respondent
- Choice task column
- task
- Chosen indicator column
- chosen
- Attribute columns
- brand,price,size
- Number of respondents
- 120
- Number of choice tasks
- 720
- Number of profile rows
- 2,880
Attributes
| Row | Levels | Utility range | Importance (percent) |
|---|---|---|---|
| brand | - | 0.612 | 25.37 |
| price | - | 1.31 | 54.16 |
| size | - | 0.494 | 20.47 |
Importance ranking
| Attribute | Importance (percent) |
|---|---|
| price | 54.16 |
| brand | 25.37 |
| size | 20.47 |
Design diagnostics
- Minimum profiles per task
- 4
- Maximum profiles per task
- 4
- Tasks with a single choice
- 720
- Total tasks
- 720
Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260917
How to report the result
Count-based conjoint analysis returned attribute importances of 60.9 per cent for monthly price, 29.5 per cent for delivery frequency and 9.6 per cent for bean sourcing, with price utilities running from 0.372 at 79 lira to minus 0.473 at 129 lira (N = 350, 10 tasks each).
An example sentence close to APA style; the numbers are representative.
When you should not use it
- With no conditional logit model fitted, the utilities and importance percentages cannot be reported as model estimates, and their count-based origin has to be stated.
- Counting a level's choices ignores which profiles competed in the same task, so a level shown against strong rivals can look unfairly weak.
- Interactions between attributes cannot be estimated, leaving cases where a price is accepted only together with a particular delivery level invisible.
- With no standard errors, intervals or significance tests, the gap between two importance percentages cannot be compared statistically.
- Since no individual-level utilities are produced, market share simulation, segmentation and the choice probability of a specific profile are all out of reach.
- Importance percentages depend on the design: widening an attribute's range of levels enlarges its share mechanically, so percentages compare only within the same design.
What to use when the assumptions are not met
- MaxDiffWhy: When the items are free-standing statements rather than feature levels of a product, calling for a simpler design to rank one list by preference.
- Van Westendorp Price SensitivityWhy: When only the bounds of price perception are wanted, deriving an acceptable band from four price questions with no experimental design to build.
- Logistic RegressionWhy: When the binary outcome should be modelled with coefficient estimates, standard errors and intervals, supplying the uncertainty measures counting withholds.
- TURF AnalysisWhy: When the question concerns how many people a limited set of options covers rather than how features are weighted.
Frequently asked questions
- Will these results match a conditional logit model?
- The ordering of the attributes usually will: which feature drives the decision and which barely registers tend to agree across the two approaches. The values will not. Utilities here are computed straight from the ratio of times chosen to times shown, whereas a logit model estimates its coefficients by considering the profiles that competed in each task together. Expecting the numbers to coincide with a logit output is therefore a mistake, and the values produced here should not be reported as model coefficients.
- Why does attribute importance depend on the design?
- Because importance comes from the distance between the highest and lowest utility inside an attribute. Set the price range from 79 to 129 lira and you get one distance; set it from 79 to 299 lira for the same product and the distance grows, pushing price's share up and everything else down. No feature is inherently forty per cent important, and the percentages hold only within the level ranges you chose. Before comparing importance percentages from two studies, check whether the level ranges match.
- Can I compute the choice probability of a profile?
- Not with this method. Probability calculations and market share simulation require an estimated choice model, and none is estimated here, nor are per-person utilities produced. What you hold is the relative weight of feature levels at the sample level. The work those weights support is ranking which level is worth improving first; it does not extend to putting a number on the share a particular product configuration would win.
- Tasks without exactly one choice were reported. What should I do?
- Start from their share of all tasks. A small share usually comes from a few abandoned questionnaires and cannot carry the result. A large share points back at data collection or reshaping: check that the choice column holds a one only on the chosen row, that every row of a task carries the same task identifier, and that no rows were lost in conversion. Since those tasks are not separately excluded, the fix is to repair the data and rerun the analysis.
References
- Green, P. E., & Srinivasan, V. (1978). Conjoint Analysis in Consumer Research: Issues and Outlook
- Louviere, J. J., Hensher, D. A., & Swait, J. D. (2000). Stated Choice Methods: Analysis and Applications
- Orme, B. K. (2020). Getting Started with Conjoint Analysis
- numpy documentation
Try it with your own data
The free plan includes 50 analysis runs a month and needs no card.