Descriptive
Weighting
Produces weights that pull the sample's distribution on variables such as gender, age and region towards known population shares, and reports how much information the correction costs.
Method summary
A completed fieldwork almost never matches the composition of the population it was drawn from: one age cohort comes in heavy, one region comes in light. Weighting closes that gap by attaching a number to each respondent, above one where a category is under-represented and below one where it is over-represented. Iterative proportional fitting arrives at those numbers one variable at a time. The shares of the first variable are moved onto their targets, then the shares of the second, which nudges the first slightly off again, and the cycle repeats. Passes continue until the remaining drift on every variable falls under a chosen threshold. What distinguishes the procedure is that it targets the margins only: gender hits its target and age hits its target, while nothing is imposed on the joint distribution of the two, which is why it still works when the census gives you separate totals rather than a full cross-classification. YouReply Analyze returns the weights together with the convergence record, prints the unweighted share next to the weighted and target share for every category, and states the price of the correction as a design effect and an effective sample size.
Which research questions does it answer?
- How do we bring the gender and age composition of a sample onto the shares held in official population registers?
- To what extent can the over-representation of graduates in an online panel be corrected?
- How far does the headline satisfaction figure move once regional shares are pulled onto target?
- What does correcting a field quota overshoot cost in effective sample size?
When should you use it?
- When the population shares are known from a dependable external source such as official statistics, an electoral register or administrative records.
- When the sample composition on one or several nominal variables departs visibly from those shares.
- When every respondent has a recorded value on the variables you intend to weight by.
- When the percentages you publish are meant to generalise to the population and readers should see that the generalisation rests on a correction.
- When only the margins are known, since no target for the joint distribution of the categories is needed here.
Required variable types
- Weighting variables: nominal columns. At least one is required, and in practice gender, age band and region are used together.
- A continuous column cannot be supplied directly. Use the derived column feature on the Variable tab to cut a field such as age into the same bands your targets use.
- Targets: for each weighting variable, a category name and the share that category holds in the population. Shares may be written as percentages or as proportions, since they are normalised to sum to one.
- There is no dependent variable. The method tests nothing; it produces a set of weights for later use together with measures of their quality.
- One row per respondent, because a pre-aggregated summary table leaves nothing to attach a per-person weight to.
Key assumptions
- Accuracy of the targets
- The correction rests on the targets being a true description of the population. Shares taken from an outdated census release or from a source with different coverage drag the sample towards the wrong distribution and conceal error instead of reducing it.
- A target for every observed category
- Each category present in the data needs a target share. A category without one has no share to be fitted to and cannot enter the procedure.
- Complete values on the weighting variables
- A respondent whose gender or region is blank cannot take part, because there is no category to place them in.
- A weight distribution that does not spread too far
- Hauling a very thin cell up to a large target loads a high weight onto a handful of people. Their answers then start to drive every percentage and the estimates become volatile.
- Convergence
- The iterative approach settles within a few passes when the targets are mutually compatible. Targets that contradict each other, or category combinations absent from the data, can leave the passes circling.
How YouReply checks these assumptions
- Accuracy of the targets: Agreement between your targets and an external source cannot be checked; the engine treats the shares arithmetically and rescales them to sum to one. Choosing and citing the source stays with you.
- A target for every observed category: This is enforced during the run: categories lacking a target are named back to you and the analysis stops. Spelling mismatches surface here too, since a category written differently in the data cannot be found in the target list.
- Complete values on the weighting variables: Rows carrying a missing value on any weighting variable are dropped before the fitting begins. Read the remaining count in the result: if many rows fell away, the weighted figures describe a smaller set than the file you uploaded.
- A weight distribution that does not spread too far: The minimum, maximum and mean weight along with the design effect are reported on every run, but no upper bound is imposed and extreme weights are not trimmed. Capping them, if you decide to cap them, has to happen outside the panel.
- Convergence: The number of passes is capped at 50 and the threshold is 0.001, both adjustable in the parameter form. When the cap is reached without convergence the result says so and still returns the weights, so never use the weights without reading the convergence field.
How the analysis is run
- 1Drop the raw file onto the panel and confirm in the grid on the Data tab that the columns you plan to weight by came through.
- 2Merge variant spellings of the same category on the Data tab, since any spelling that does not match your target list will stop the run.
- 3On the Variable tab set those columns to nominal, declare the missing-value codes, and build a derived column that cuts a continuous field such as age into exactly the bands your targets use.
- 4Write down the target shares from your external source alongside the category names, either as percentages or as proportions.
- 5Pick this method from the descriptive group on the Analysis tab and enter the weighting variables with their targets in the parameter form.
- 6The pass limit and the threshold can stay at their defaults; raise the pass limit if your targets are demanding.
- 7Run the analysis, look at the convergence record first, then check in the comparison table whether the weighted shares have landed on their targets.
- 8Note the design effect and the effective sample size and export the cell weights and tables to Excel. This method draws no chart.
Statistics and tables produced
- Convergence record and pass count
- Whether the fitting fell under the threshold and how many passes it used. A run that hit the cap without converging says so, and the weights still come back.
- Cell weights
- One weight per cell of the cross-classification formed by the weighting variables. A respondent's weight is the weight of the cell they belong to.
- Minimum, maximum and mean weight
- Three numbers describing the spread of the weights. The mean equals one because the weights are rescaled at the end of each pass, and the gap between the extremes shows how hard the correction pulled.
- Kish effective sample size
- How large an equally weighted sample your weighted data are worth once the variability of the weights is accounted for. It is always at or below the raw number of observations.
- Design effect
- The number of observations divided by the effective sample size. A value near one means the correction cost little, while a growing value says the estimates have become noticeably more volatile.
- Unweighted, weighted and target share table
- Three shares printed side by side for every category of every weighting variable. This table is where you read what the correction changed and verify the quality of the fit.
- Observations entering the analysis
- How many rows remained after those with missing values on the weighting variables were dropped.
Effect size and confidence intervals
- Design effect
- This takes the place of an effect size here, because nothing is being tested. Read it as the number of observations over the effective sample size: 1.0 describes an equally weighted sample, values around 1.2 are carried comfortably in practice, and 1.5 or above means the weights have made the estimates appreciably noisier. A high design effect usually traces back to a thin cell being hauled up to a large target, and the remedy is normally coarser categories or fewer weighting variables.
- Kish effective sample size
- The figure that tells a reader what information base the weighted percentages rest on. It derives from the sum of the weights and the sum of their squares, and it falls as the weights grow apart. Margins of error should be built on this number rather than on the raw count: a sample of 900 with an effective size of 640 carries the precision of roughly 640 interviews.
- Spread of the weight distribution
- The distance between the smallest and the largest weight shows how much load the correction put on how few people. Since the mean is fixed at one, the reading happens at the extremes: a top weight approaching six means a single answer is counted for six people and can move subgroup percentages on its own.
This method returns no confidence interval of any kind. Weighted shares, the design effect and the effective sample size all arrive as single values, and neither an interval around a weighted percentage nor the uncertainty of the weights themselves is computed. If intervals are needed for weighted estimates, base the arithmetic on the effective sample size and do it outside the panel. No significance test runs either: no p value is produced for whatever gap remains between a weighted share and its target.
Example research question and example result
The numbers below are a representative example, not data from a real study or a real user.
- Research question
- An illustrative city study interviewed 900 residents face to face, and the gender and age composition is to be brought onto official population shares.
- Variables
- Weighting variable: gender (women, men) · Weighting variable: age band (18-34, 35-54, 55 and over) · Targets: 51 and 49 per cent on gender; 36, 38 and 26 per cent across the age bands
- Example result
- Women made up 57.3 per cent of the achieved sample, the youngest band 44.1 per cent and the oldest band 15.2 per cent. The fit converged in 6 passes. Weights ranged from 0.71 to 1.93 with a mean of 1.00. The effective sample size came to 806 and the design effect to 1.12. After weighting, women stood at 51.0 per cent and the age bands at 36.0, 38.0 and 26.0 per cent, so all three landed on target.
- Interpretation
- Younger women were the easiest group to reach in the field, which is the direction in which the achieved sample leaned, and the fit moved the shares onto target without stretching the weights severely. The oldest band was the thinnest category and therefore carries the largest weight: one interview there stands in for close to two. A design effect of 1.12 says the 900 interviews behave like roughly 806 after weighting, so the margin of error is a little wider than the raw size suggests. Bear in mind that the panel does not carry these weights into other methods, so any cross-tabulation or mean comparison you run next works on unweighted rows. The figures are illustrative.
Real output on a sample dataset
The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.
General customer survey (synthetic)
A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.
- Rows
- 300
- Columns
- respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.
Research question: Which weights are needed to pull the sample to the population distribution of gender and region?
- Analyzed variables
- gender,region
- Target population shares
- [object Object]
- Valid observations
- 300
- Estimation converged
- Yes
- Iterations
- 5
- Effective sample size (Kish)
- 281.30
- Design effect
- 1.07
- Smallest weight
- 0.667
- Largest weight
- 1.43
- Mean weight
- 1
Cell weights
| Row | Cell weight | Valid observations |
|---|---|---|
| Female | Aegean | 0.834 | 39 |
| Female | Central Anatolia | 1.04 | 32 |
| Female | Marmara | 1.32 | 44 |
| Female | Mediterranean | 0.667 | 44 |
| Male | Aegean | 0.905 | 37 |
| Male | Central Anatolia | 1.13 | 37 |
| Male | Marmara | 1.43 | 33 |
| Male | Mediterranean | 0.724 | 34 |
Achieved margins
| Row | Male | Female |
|---|---|---|
| gender | - | - |
| region | - | - |
Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914
How to report the result
Data were weighted to the gender and age margins by iterative proportional fitting (convergence in 6 passes, weights from 0.71 to 1.93), giving an effective sample size of 806 and a design effect of 1.12.
An example sentence close to APA style; the numbers are representative.
When you should not use it
- The correction only works through the variables you weight by, so selection bias unrelated to those variables survives the procedure untouched.
- Because only the margins are fitted, the joint distribution of gender and age need not resemble the population, which leaves the approach short where the joint distribution is what matters.
- Nothing is added about the answers of people who were never reached; the composition of non-response is corrected, non-response itself is not.
- Thin cells taking high weights make estimates volatile, and the volatility shows up more sharply in subgroup percentages; since extreme weights are not trimmed, any cap is yours to impose.
- The resulting weights are not applied automatically to later analyses, so a weighted t-test or a weighted regression cannot be run in this flow.
- Revising the targets after seeing the output turns the exercise into a search for a preferred percentage; the target set belongs to the design stage and should be reported with its source.
What to use when the assumptions are not met
- Frequency DistributionWhy: Run first to see how far the achieved distribution sits from the targets, which is what decides whether weighting is needed at all.
- Descriptive StatisticsWhy: When the unweighted and weighted tables need to be set against each other to show which direction the correction moved the averages.
- Banner TableWhy: When the question is the significance of differences between subgroups rather than representativeness, it delivers breakdown reporting with significance letters.
Frequently asked questions
- Will the weights I produce carry into my other analyses?
- They will not. This method reports the weights, their quality measures and the comparison table, while every other method in the panel works on unweighted rows. So read the weighted percentages here and carry them into your report, and present the remaining tests as unweighted results. If a weighted model estimate is unavoidable, the only route is to export the cell weights to Excel and do the arithmetic outside the panel.
- Should my targets be percentages or proportions?
- Either is accepted. Writing 51 gives the same answer as writing 0.51, because targets are normalised to sum to one once read, and a total that lands slightly off 100 through rounding is no problem either. What does demand care is the category names: the name in the target list has to match the spelling in the data, otherwise the run stops and names the category left without a target.
- How many variables should I weight by?
- Adding variables improves representativeness but multiplies the cells, leaves some of them very thin, pulls the weight extremes apart and raises the design effect. Gender, age band and region are usually enough in practice. If you decide to extend the list, inspect the design effect after each addition: when the loss in effective sample size outweighs the representativeness you gained, the addition is better reversed.
- The result says it did not converge. Can I still use the weights?
- Proceed carefully. The weights do come back, but since the shares never fell inside the threshold one or more categories may be sitting off target, so start by measuring the remaining gap in the comparison table. A small residual gap is acceptable for most reporting. Where the gap is substantial, try raising the pass limit above 50, defining the categories more coarsely, or merging combinations that are nearly empty in the data, since failure to converge usually signals incompatible targets or cells with almost nobody in them.
References
- Kish, L. (1965). Survey Sampling
- Valliant, R., & Dever, J. A. (2018). Survey Weights: A Step-by-Step Guide to Calculation
- Ireland, C. T., & Kullback, S. (1968). An Iterative Procedure for Estimation in Contingency Tables
- numpy documentation
Try it with your own data
The free plan includes 50 analysis runs a month and needs no card.