Categorical
Proportion Test
Tests whether an observed proportion on a binary response departs from a hypothesised value, or whether the proportions of two independent groups differ from each other.
Method summary
The proportion test is the direct way to put a percentage from a binary response under scrutiny, and the panel runs it in two modes. In one-sample mode your proportion is set against a value that comes from outside the data, such as the result of a previous wave, a sector average or a target threshold. Three results are reported together in that mode: a z test resting on the normal approximation, an exact probability computed straight from the binomial distribution, and a Wilson confidence interval for the proportion itself. In two-sample mode the same binary response is compared across two independent groups and significance is judged from a pooled z statistic. That mode carries no interval calculation: the z and p values test the hypothesis that the two proportions are equal without estimating the size of the gap between them. Because the subject is a percentage, readability depends on giving the numerator and denominator alongside it: 45 per cent means something entirely different in a sample of 420 than in a sample of 40. YouReply Analyze reports the proportions, the counts and the test results produced for whichever mode you chose.
Which research questions does it answer?
- Has the share who would recommend the brand moved away from the 40 per cent recorded in the previous wave?
- Did survey completion fall short of the 70 per cent threshold we set ourselves?
- Does recall of a campaign differ between the urban and the rural samples?
- Does use of the mobile application vary between participants recruited through two different invitation channels?
When should you use it?
- When the percentage on a single binary response is to be set against a known value originating outside the study.
- When the binary response was measured in two independent groups whose proportions are to be compared.
- When the response really is binary, such as yes and no, completed and not completed, recalled and not recalled.
- When you want to present the finding as a percentage and it helps to show the uncertainty around it through the interval available in one-sample mode.
- When the groups are independent; designs in which the same people are measured twice call for a paired binary test instead.
Required variable types
- One-sample mode: a single binary column plus the hypothesised proportion it is compared against, which you must take from a source independent of this study.
- Two-sample mode: a binary outcome column plus a categorical column with exactly two levels defining the groups.
- If the grouping column carries more than two levels, either narrow the comparison to two or choose a method built on a contingency table.
- Each row has to be a participant; a file of pre-computed percentages cannot feed this test.
- You need to know which level of the binary column counts as a success, because the reported proportion is that level's share.
Key assumptions
- Independence of observations
- Every participant should contribute to the proportion once, and participants should be independent of one another. In samples drawn from within clusters such as households or schools, the true uncertainty is wider than the computation suggests.
- The hypothesised value must not come from the data
- In one-sample mode the comparison value has to originate outside the data being tested. Entering a rounded version of your own sample's proportion empties the test of meaning.
- Counts large enough for the normal approximation
- The z test relies on both the success and the failure counts being reasonably large. A familiar rule asks for at least ten of each, and the approximation deteriorates as the proportion nears zero or one.
- The binary outcome correctly defined
- Which level counts as a success has to be chosen on conceptual grounds, because the direction of the proportion follows from it. Reading missing answers as a no pulls the proportion systematically down.
- Disjoint groups in two-sample mode
- The two groups being compared must not share a participant. If the grouping column places one person in both groups, the independence assumption collapses.
How YouReply checks these assumptions
- Independence of observations: The panel does not check this assumption automatically; the researcher evaluates it.
- The hypothesised value must not come from the data: The panel has no way of knowing where the value you typed came from; naming the source of the hypothesised proportion is your job in the methods section.
- Counts large enough for the normal approximation: No warning or assumption flag is raised for this rule. Since one-sample mode also reports the exact binomial probability, reading the exact value is the safer course in small samples where the approximation is strained.
- The binary outcome correctly defined: Missing values are set aside before the run, so the proportion rests on valid answers only; the missingness itself is not evaluated as a finding.
- Disjoint groups in two-sample mode: The levels in the grouping column are read row by row, but whether individuals recur across groups is not checked.
How the analysis is run
- 1Drag the data file onto the panel's upload area; CSV and XLSX are read and the contents open in the grid on the Data tab.
- 2Inspect the binary outcome column in the grid, along with the grouping column if you will use one, and merge levels that split in two through spelling differences.
- 3On the Variable tab mark the columns as nominal, declare the missing-value codes, and create the derived column that reduces a multi-level outcome to binary where needed.
- 4On the Analysis tab pick the method from the categorical group; methods that do not fit are never removed from the list, only dimmed with the reason beside them.
- 5In the parameter form settle the mode: enter the hypothesised proportion for one-sample mode, or select the grouping column for two-sample mode.
- 6Run the analysis. No chart is produced for this method, so the result tables are listed directly as collapsible sections.
- 7In one-sample mode read the Wilson interval and see whether the hypothesised value falls inside or outside it, then export the tables to Excel.
Statistics and tables produced
- Observed proportion and counts
- The success count, the number of valid observations and the proportion computed from them are given together. In two-sample mode this trio is reported for each group separately.
- The z statistic
- In one-sample mode it expresses how many standard errors separate the observed proportion from the hypothesised value; in two-sample mode the distance between the groups relative to the pooled proportion.
- Significance value
- The p value corresponding to the z statistic. One-sample mode tests the hypothesis that the proportion equals the hypothesised value, two-sample mode that the two proportions are equal.
- Exact binomial probability
- One-sample mode also reports the probability computed directly from the binomial distribution. In small samples, where the normal approximation is strained, this value is more trustworthy.
- Wilson confidence interval for the proportion
- In one-sample mode an interval built by the Wilson method is given for the observed proportion itself. The interval belongs to the proportion; no interval is produced for the gap between two groups.
Effect size and confidence intervals
- No effect size label
- This method returns no standardized measure of magnitude and no strength label to go with one. In practice the most legible measure of magnitude is the proportion itself: the gap in percentage points between the observed proportion and the hypothesised value in one-sample mode, and between the two groups' proportions in two-sample mode. Writing those gaps down with their numerators and denominators produces a more readable report than any separate coefficient.
A confidence interval is returned in one-sample mode only, and it belongs to the observed proportion itself: it is built by the Wilson method, so its bounds stay inside zero and one even when the proportion sits close to an extreme. Two-sample mode computes no interval at all, neither for the gap between the two proportions nor for a relative risk or an odds ratio. Where the difference between two groups has to be reported with an interval, that calculation must happen outside the panel.
Example research question and example result
The numbers below are a representative example, not data from a real study or a real user.
- Research question
- In an illustrative customer study, does the share who would recommend the service depart from the 40 per cent known from the previous wave?
- Variables
- Binary column: would recommend the service (would, would not) · Hypothesised proportion: the value 0.40 carried over from the previous wave
- Example result
- Among 420 valid answers, 189 participants said they would recommend the service, putting the observed proportion at 45.0 per cent. One-sample mode returned z = 2.09 with p = 0.037, and the exact binomial probability came to 0.040. The Wilson interval for the proportion ran from 40.3 to 49.8 per cent.
- Interpretation
- The observed proportion sits 5 points above the hypothesised value and neither test treats that gap as routine sampling movement. The lower bound of the Wilson interval at 40.3 per cent says the same thing from another angle: the hypothesised 40 per cent lies just outside the interval, which places the finding close to the boundary. A result like this is better read with the width of the interval in view than presented as a settled rise, since in a sample of 420 the uncertainty around the proportion spans a band of roughly nine and a half points. Because the mode addresses a single proportion, nothing here speaks to why the share moved, and being sure the sampling procedure held steady between waves is a precondition for any interpretation. The figures are illustrative.
Real output on a sample dataset
The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.
General customer survey (synthetic)
A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.
- Rows
- 300
- Columns
- respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.
Research question: Is the share of buyers different from fifty percent?
- Analyzed column
- purchased
- Value counted as a success
- 1
- Hypothesized value
- 0.5
- column
- purchased
- success_value
- 1
- test_value
- 0.500
- Proportion
- 0.517
- Number of successes
- 155
- Valid observations
- 300
- z statistic
- 0.577
- p value
- 0.564
- Exact binomial p value
- 0.603
- Confidence interval lower
- 0.460
- Confidence interval upper
- 0.573
- Significant
- No
Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914
How to report the result
The share who would recommend the service was 45.0 per cent (189/420), significantly above the 40 per cent observed in the previous wave, z = 2.09, p = .037, 95 per cent Wilson interval for the proportion [40.3, 49.8] per cent.
An example sentence close to APA style; the numbers are representative.
When you should not use it
- Two-sample mode returns no interval of any kind, so a difference between groups can be reported only as a point estimate with its p value.
- Because the condition behind the normal approximation is not checked, the z test can deliver an over-optimistic p value when the success or failure count is very small.
- The method cannot be applied when the grouping column holds more than two levels, so three or more groups' proportions are not compared in one test.
- It does not apply to designs in which the same people are measured twice, since both modes assume independent observations.
- The soundness of the hypothesised value is never examined, so a wrong or outdated comparison figure yields a result that is technically clean and substantively empty.
- Testing proportions across many subgroups within one dataset raises the false positive risk as the comparisons accumulate, and no correction is applied for it.
What to use when the assumptions are not met
- Chi-Square Test of IndependenceWhy: When the outcome or the grouping variable holds more than two categories; it tests independence on a contingency table and supplies a coefficient for the strength of the association.
- Fisher's Exact TestWhy: When two groups' proportions are compared on very small counts; computing the exact probability removes any reliance on the approximation in sparse tables.
- Chi-Square Goodness of FitWhy: When a single categorical variable with more than two categories is compared against an expected distribution.
- Logistic RegressionWhy: When the rate of a binary outcome is to be explained with several variables at once and confounders need to be held constant.
Frequently asked questions
- Can I get a confidence interval for the gap between two groups' proportions?
- You cannot. An interval is produced in one-sample mode only, and it belongs to the observed proportion rather than to a difference between two of them. Two-sample mode gives you the z statistic, the p value, and each group's proportion with its counts. If you need the interval for the difference, running one-sample mode twice and reporting two intervals is not a substitute, because whether two intervals overlap does not map onto whether the difference is significant; the interval for the difference has to be computed outside the panel.
- The z test and the exact binomial probability disagree. Which do I report?
- One-sample mode reports both, and the distance between them widens as the sample shrinks or the proportion approaches an extreme. The working rule is that the z value suffices when both the success and the failure counts are comfortably large, and that the exact probability is the more trustworthy of the two when either count is small. Where the two values straddle your significance threshold, stating both and resting the decision on the exact figure is more defensible than quietly choosing one.
- Why does the Wilson interval differ from the textbook normal-approximation interval?
- The classical approach draws a symmetric band around the proportion, and its bounds can spill into impossible values as the proportion nears zero or one. The Wilson method builds the interval while accounting for the fact that uncertainty depends on the level of the proportion, which makes the interval slightly asymmetric and keeps its bounds between zero and one. In small samples and at extreme proportions the two methods can diverge visibly, so it is worth stating in your report that the interval was built by the Wilson method.
- Can I compare two measurements of the same people with this test?
- No. Two-sample mode assumes the groups are independent of each other, whereas before and after measurements of the same people are paired data, and an independent test applies the wrong standard error to them. For two paired binary measures, the McNemar test built on the participants who switched sides is the right choice, and Cochran's Q test for three or more measures.
References
- Agresti, A. (2013). Categorical Data Analysis
- Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics
- Agresti, A., & Coull, B. A. (1998). Approximate Is Better than Exact for Interval Estimation of Binomial Proportions
- statsmodels.stats.proportion.proportions_ztest documentation
Try it with your own data
The free plan includes 50 analysis runs a month and needs no card.