Choice-based

TURF Analysis

Counts how many different people a set of options covers and finds the combination of a given size that reaches the largest number of them.

Method summary

TURF stands for total unduplicated reach and frequency, and it answers the question of what to put on a shelf, on a menu or in a feature list when only a few slots are available. Its arithmetic is plain counting: for every respondent it is known, as a yes or no, which options appeal to them, and the reach of a set is simply the number of people who like at least one option in it. Stacking the individually strongest options rarely adds up, because the same people tend to like several of them at once, and TURF surfaces options that cover each other's gaps instead. Frequency looks at the other dimension: among the people a set reaches, how many of its options they actually liked. YouReply Analyze reports the reach of every option on its own, the strongest combinations up to the size you ask for, and a step-by-step portfolio that shows what each additional option buys you.

Which research questions does it answer?

  • If only three new flavours can be added to the range, which three appeal to the most customers?
  • How many of eight campaign messages, and which ones, cover the bulk of the target audience?
  • What does a third product add on top of two that are already on sale?
  • How much coverage is lost if the bundle shrinks from five items to four?

When should you use it?

  • When the number of slots is limited and a subset has to be picked out of a larger pool.
  • On liking, interest or purchase-intent questions where respondents may tick more than one option.
  • When the goal is to maximise the number of people served by at least one option rather than an average rating.
  • When the overlap between options matters, since individual percentages reveal nothing about it.
  • When the decision is a portfolio decision and the question is what each added item contributes.

Required variable types

  • One binary column per option: liked or not liked, chosen or not chosen.
  • At least two option columns are required, and the number of combinations to scan grows quickly with the pool.
  • One row per respondent; tables of pre-summarised percentages cannot be used here.
  • There is no independent or grouping variable; the analysis runs over a single sample.
  • To use a five-point purchase-intent item you first have to settle on a cut-off and turn it into a derived binary column on the Variable tab.

Key assumptions

Binary coding
The columns must carry a chosen or not chosen value. Feeding a continuous liking score into the analysis without a cut-off robs the reach counts of meaning and buries the cut-off decision where nobody can see it.
Options from a single decision set
The columns you compare have to be genuinely interchangeable options. Pulling items from unrelated questions into one scan yields a portfolio that cannot be put into practice.
Enough options and a ceiling on set size
Since every combination up to the chosen size really is tried, the amount of work multiplies as the pool and the set size grow.
A sample that represents the target audience
Reach percentages invite being read as market size. When the sample does not represent the audience, the percentages stay internally consistent but cannot be carried outside it.
How missing answers are handled
Treating a blank in an option column as not liked produces different reach counts from dropping that respondent altogether.

How YouReply checks these assumptions

  • Binary coding: Nothing verifies that the coding is genuinely binary. You set the cut-off, you build the derived column, and you state in the report which cut-off you used.
  • Options from a single decision set: No check asks whether the columns conceptually belong to the same decision, which makes this entirely a matter of study design.
  • Enough options and a ceiling on set size: The run requires at least two option columns and caps the requested set size at six, applying the ceiling when a larger value is asked for.
  • A sample that represents the target audience: Representativeness is not assessed and no weighting is applied here; reach percentages come from raw row counts.
  • How missing answers are handled: Declaring missing-value codes on the Variable tab is up to you, and you are expected to confirm which code counts as blank before running.

How the analysis is run

  1. 1Drag the file into the upload area; the panel reads both CSV and XLSX.
  2. 2On the Data tab confirm in the grid that the option columns arrive as zeros and ones and that text labels have been converted to numbers.
  3. 3On the Variable tab give every option column a readable label, since the result tables are written with those labels and long column names make the report hard to follow.
  4. 4On the Analysis tab pick TURF analysis from the choice-based group; switching on Only ones that fit my data dims unsuitable methods with a plain reason instead of removing them from the list.
  5. 5In the parameter form tick the option columns to scan, set the largest set size and say how many combinations to list.
  6. 6Run the analysis. Reach per option, the strongest combinations, the incremental portfolio and total reach arrive as separate collapsible sections.
  7. 7Compare the top row of the exhaustive ranking with the portfolio row of the same size, and if the two diverge, write down which one you chose and why.
  8. 8Export the result to Excel; the computation credits card names the library call and its version.

Statistics and tables produced

Reach per option
How many people each option reaches on its own, as a count and a percentage; this is where reading the portfolio starts.
Best combinations
Every combination up to the chosen size is tried and the leaders are listed in order of reach, with the length of the list set in the parameter form.
Frequency of a combination
For each combination, the average number of its options liked by the people it reaches. Reach measures breadth while frequency measures depth.
Incremental portfolio
A greedy path adds, at each step, the option that grows reach the most, and shows how many new people each addition brings in.
Total reach of all options
The number of people covered if the whole pool were offered, reported as the upper bound that reach can approach.

Effect size and confidence intervals

Unduplicated reach
The share of people who like at least one option in a set, and the principal magnitude this method produces. It comes from counting: not a model estimate but the observed headcount divided by the sample size. It has no interpretive benchmark, and only becomes meaningful against the ceiling of total reach and against what one more option would add.
Frequency
The average number of options liked by the people a set reaches, with one as its floor. Where two sets achieve the same reach, the one with the higher frequency engages its audience across more of its options. This too is an average rather than an estimate.
Incremental gain
The number of new people an added option brings into the portfolio. Gains that shrink rapidly say the pool is saturating and that lengthening the list buys less and less. It is read straight off the counts and carries no allowance for sampling variability.

No confidence interval is returned for any quantity here and no significance test is performed. Reach percentages, frequencies and incremental gains are observed counts, and nothing is produced to say whether the gap between two combinations could be sampling noise. When two combinations sit close together, a difference of a few points should not be treated as a settled advantage.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
In an illustrative beverage study, if only three of eight candidate flavours can take shelf space, which three appeal to the most consumers?
Variables
Eight binary columns: 1 for respondents who said they would definitely try the flavour, 0 otherwise · Sample: 500 valid responses
Example result
Individual reach came out as A 212 people (42.4 per cent), B 188 (37.6), D 178 (35.6), C 166 (33.2), E 160 (32.0), F 149 (29.8), G 121 (24.2) and H 103 (20.6). Scanning all sets of three, the strongest was A, D and F with a reach of 341 people (68.2 per cent) and a frequency of 1.58. The greedy portfolio started with A (212), rose to 308 (61.6 per cent) when E added 96 new people, and to 335 (67.0 per cent) when D added 27 more. Offering all eight flavours would have reached 428 people (85.6 per cent).
Interpretation
Flavour B ranks second on its own yet stays out of the leading trio, because the people who like it are largely the same people who like A. Flavour F makes it in despite a lower individual reach, since the segment it covers barely overlaps with the rest. At the same size the exhaustive scan found a set reaching 341 while the portfolio path reached 335, so the two lists disagree: the greedy route takes the best available step each time and cannot promise the best set overall. Set against the ceiling of 85.6 per cent for all eight, the 68.2 per cent from three flavours shows most of the available coverage arriving from a small part of the shelf. The 6-person gap may sit below sampling noise, and nothing in this method tests that. The figures are illustrative.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

Variable examined

Market research survey (synthetic)

Two hundred and fifty respondents: the brand in use, three demographic breaks, six choose all that apply option columns and six purchase intent questions. The option columns overlap, so the extra reach an item brings can genuinely be measured.

Rows
250
Columns
respondent_id, gender, region, age_band, main_brand, opt_price, opt_quality, opt_service, opt_brand, opt_speed, opt_design, buy_a, buy_b, buy_c, buy_d, buy_e, buy_f
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: Which trio of the six products reaches the most people?

Analyzed columns
buy_a,buy_b,buy_c,buy_d,buy_e,buy_f
Largest combination size
3
Combinations to list
10
Valid observations
250
max_items
3
Reach of all items (percent)
92.80

Reach per item

Reach per item
RowPeople reachedReach (percent)
buy_a13855.20
buy_b12048
buy_c10943.60
buy_d9839.20
buy_e9538
buy_f7228.80

Best combinations

Best combinations
ItemsSizePeople reachedReach (percent)Items liked per reached person
-321385.201.67
-320983.601.65
-320883.201.59
-3205821.79
-320280.801.75
-320280.801.69
-320180.401.64
-320180.401.53
-319979.601.60
-319879.201.58

Incremental portfolio (greedy order)

Incremental portfolio (greedy order)
ItemNewly reached peopleNewly reached people (percent)Cumulative reachCumulative reach (percent)
buy_a13855.2013855.20
buy_b4919.6018774.80
buy_d2610.4021385.20

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260916

How to report the result

Across all three-flavour sets, the highest unduplicated reach was obtained for A, D and F: 341 respondents, 68.2 per cent, frequency 1.58 (N = 500; total reach of all eight flavours 85.6 per cent).

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • With no significance test and no interval, closely spaced reach values cannot be separated statistically.
  • Because the scan is exhaustive, set size is capped at six, so wider portfolios cannot be evaluated in a single run.
  • The greedy portfolio and the exhaustive ranking may name different sets at the same size, and the panel does not reconcile them; the choice and its rationale rest with the researcher.
  • Reach only asks whether at least one option appealed, ignoring how strongly, how many units someone would buy, or how they would respond to price.
  • Production cost, shelf cost and cannibalisation between options are not modelled, so a commercial decision needs that information from outside the analysis.
  • The binary cut-off drives the result, and raising or lowering it can reorder everything, which is why it should be fixed once and justified.

What to use when the assumptions are not met

  • Multiple Response AnalysisWhy: When all you need is the share ticking each option, describing a multiple-choice question without the overlap and portfolio questions.
  • MaxDiffWhy: When the options need to be ranked by preference; it yields relative importance and removes the binary cut-off decision.
  • Frequency DistributionWhy: When reporting the distribution of a single list of options suffices, giving counts and percentages without scanning combinations.
  • K-Means Cluster AnalysisWhy: When the question concerns which consumer segments exist rather than which set to offer, grouping respondents by their liking patterns.

Frequently asked questions

Why is taking the three highest-reaching options not the best trio?
Because reach does not add up. If two well-liked options appeal to the same people, putting both on the shelf buys almost nobody new from the second. The scan counts each person in a set once and discards the overlap that way. An option with modest reach but a distinct audience can therefore displace one that sits higher on the individual list.
Why do the best-combination list and the incremental portfolio disagree?
They come from two separate calculations. The combination list tries every possibility up to the size you asked for and orders them by reach. The portfolio advances one step at a time, adding whichever option grows reach most at that moment and never revisiting an earlier choice. That route is fast and readable, since it shows what each item contributed, but it cannot promise the best set at a given size. When the two diverge at the same size, the higher-reaching set is the better set, while the portfolio explains how the ordering came about.
Can I raise the set size above six?
No. Because every combination really is tried, the number of possibilities multiplies with the pool and the size, which is why the requested size is capped at six. To think about a wider portfolio, read the incremental list instead: it shows the gain at each step, so you can see how far the returns have fallen past the sixth item. Narrowing the pool on defensible conceptual grounds is another route.
How do I know whether a difference in reach is real?
This method will not tell you. Neither a p value nor an interval is returned, and every number is an observed count. In practice a gap of a few points between two combinations is small enough to swap places if the study were repeated. Rather than leaning on the reach difference alone, weigh frequency, incremental gain and the cost information that lives outside the analysis; treating two near-equal sets as equivalent is usually the more defensible stance.

References

  • Miaoulis, G., & Michener, R. D. (1976). Market Research Handbook
  • Malhotra, N. K. (2019). Marketing Research: An Applied Orientation
  • Orme, B. K. (2013). Applied Conjoint Analysis
  • numpy documentation

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.