Categorical

Frequency Distribution

Describes how many people fall in each category of one or several categorical columns, together with the percentage and the cumulative percentage.

Method summary

A frequency distribution is the plainest table for opening up a categorical column: every category becomes a row carrying the number of people who chose it, that number's share of the valid answers, and the share accumulated from the top of the table down to that row. The cumulative column only bears interpretation where the categories are ordered, and it then gives the portion up to a given level at a glance, such as the combined share of everyone at secondary level or below. Because the method accepts several columns in one run, a whole demographic block or a group of related questions can be tabulated at once. This is a descriptive instrument: it makes the distribution visible without stating or testing a hypothesis, so the result contains no test statistic, no p value and no effect size. Testing whether the distribution matches an external expectation means moving to the chi-square goodness-of-fit test, and breaking it down by a second variable means moving to cross-tabulation. Even so, this is where most analyses begin, because sparse categories, unexpected labels and concentrations of missing data all surface here first.

Which research questions does it answer?

  • What does the educational composition of the sample look like, and at which level does it concentrate?
  • How many participants sit in each subscription tier, and what share does each hold?
  • How do the options in the brand awareness question rank by share?
  • Which questions in the demographic block carry a noticeably high rate of missing answers?

When should you use it?

  • When you are meeting the data for the first time and want a quick view of how full the categories are before settling on an analysis plan.
  • When the descriptive section of the report has to present the composition of the sample to the reader.
  • When an ordinal measure invites you to state the portion below a given level, so the cumulative column has a real use.
  • When several columns are to be tabulated in one run and collected into a single Excel file.
  • When deciding whether to merge categories and you first need to see which labels hold very few people and which spelling variants have multiplied a category.

Required variable types

  • One or several categorical columns, held either as text labels or as integer category codes.
  • Columns at the ordinal level suit the method particularly well, since the cumulative percentage only makes sense when the category order does.
  • Binary variables tabulate too, in which case the table shrinks to two rows and the percentages read directly as proportions.
  • Each row should stand for one participant, because a person with several records is counted several times in the percentages.
  • Handing over a continuous measure produces hundreds of one-person rows, so bin it first with a derived column on the Variable tab.

Key assumptions

At least nominal measurement
The values in the column need to fall into countable categories, with each participant landing on one label. Once the number of distinct values approaches the number of participants, the table stops describing anything.
Meaningful, non-overlapping categories
Each label should name one interpretable group and the labels must not overlap. Upper and lower case spellings of the same category open two rows, split the share between them and shift the cumulative column too.
Missing values and the percentage base
Which base the percentages rest on is decisive for the accuracy of the report. Non-respondents left as a row of their own put the whole sample in the base, while filtering them out bases the percentages on valid answers alone.
Interpretability of the cumulative percentage
A cumulative percentage only says something where a meaningful order runs through the categories. On unorderable labels the column is still computed, but it cannot be interpreted.
Enough people per row
In categories built from a handful of people, the percentage moves noticeably once a single answer is added, so the rows have to be reasonably full for the description to hold up.

How YouReply checks these assumptions

  • At least nominal measurement: The measurement level you assign on the Variable tab drives the method picker, and an unsuitable selection dims the method with the reason stated. Nothing checks whether the number of categories is still interpretable.
  • Meaningful, non-overlapping categories: The panel counts the labels from the cleaned data exactly as they appear and never merges spelling variants; merging happens through value labels or a derived column on the Variable tab.
  • Missing values and the percentage base: Missing-value codes declared on the Variable tab are excluded from the count, so the percentages rest on valid answers. Skip that declaration and those values are counted as an ordinary category, with no separate warning raised.
  • Interpretability of the cumulative percentage: The panel does not drop the column according to whether the variable is truly ordered; the cumulative percentage is printed in every case. Whether to interpret it is left to the researcher.
  • Enough people per row: No minimum category size is required; however few people a label holds, the row enters the table and its percentage is printed. Which percentages are solid enough to report is for you to settle from the frequency column.

How the analysis is run

  1. 1Drag your CSV or XLSX file onto the upload area and confirm in the editable grid on the Data tab that the columns were parsed correctly.
  2. 2Look over the labels in the grid, since an unexpected label usually traces back to a spelling variant or to an undeclared missing-value code.
  3. 3On the Variable tab set each column's measurement level as nominal or ordinal, enter value labels so the table rows read well, and declare the missing-value codes.
  4. 4Search for the method on the Analysis tab and pick one or several columns to tabulate in the parameter form; the data requirements box lists the eligible columns.
  5. 5Run the analysis. Each column arrives as a collapsible section holding a table with frequency, percentage and cumulative percentage columns, and no method-specific chart is drawn for this method.
  6. 6Read how full the rows are from the frequency column, the shares from the percentage column, and, on ordinal variables, the portion below a threshold from the cumulative column.
  7. 7If you have spotted sparse categories, return to the Variable tab, build a derived column that merges them and run the table again.
  8. 8Export the tables to Excel; the computation credits card names the library call and its version, and the run is written to the history tab.

Statistics and tables produced

Frequency column
The number of people in each category. This is the first column to consult when judging how robust the percentages are.
Percentage column
Each category's share of the valid answers. The base shifts depending on whether you declared the missing-value codes.
Cumulative percentage column
The share accumulated from the top of the table to the row in question. On ordered categories it gives the portion below a threshold directly; on unorderable labels it is not interpreted.
Category list
Every label that actually appears in the cleaned data. An unexpected label or a category split in two is noticed here.
Separate tables for several columns
Select several columns in one run and each gets its own table, with all of them collected into a single Excel file.
Valid observation count
The base on which each table's percentages were computed, that is, the number of valid observations entering the count.

Effect size and confidence intervals

Neither an effect size nor significance is produced
Since the method is descriptive, no effect size is computed and no significance test is performed, so the result holds neither a test statistic nor a p value. To test whether a concentration you see departs from an expected distribution, move to the chi-square goodness-of-fit test.
Percentages and the cumulative distribution as descriptive quantities
The quantities that carry the numerical reading here are the percentage and the cumulative percentage. The percentage gives a category's share of the valid answers, and the gap in points between the fullest and the sparsest category conveys how uneven the distribution is. The cumulative percentage reduces the portion below a threshold on an ordinal variable to a single number, such as the combined share of the levels at secondary and below. Neither is a standardised measure of magnitude, neither can be used to compare strength across samples, and neither carries any verdict on significance; both describe the composition of the sample in front of you and belong next to the frequency column.

No confidence interval is returned for this method. Percentages and cumulative percentages come as point estimates with no band of uncertainty computed around them. If generalising a category's share to a population calls for an interval estimate, move to the proportion test or work the figure out outside the panel. With no interval printed, how much a sparse category's percentage might move can only be sensed from the frequency column.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
In an illustrative community study, how is the highest completed level of education distributed, and up to which level does the majority fall?
Variables
Variable tabulated: highest completed level of education (primary, secondary, upper secondary, bachelor's, postgraduate) · Missing values: 12 people who did not answer, excluded from the count because a missing-value code was declared
Example result
Of 640 records, 628 entered the count as valid answers. Frequencies were 86 at primary, 154 at secondary, 198 at upper secondary, 152 at bachelor's and 38 at postgraduate level. The percentages came to 13.7, 24.5, 31.5, 24.2 and 6.1, and the cumulative percentages to 13.7, 38.2, 69.7, 93.9 and 100.0.
Interpretation
The fullest level is upper secondary, holding 31.5 per cent of the sample, while postgraduate forms the sparsest row at 6.1 per cent, and the 25-point spread between those two marks an uneven distribution. Because education is an ordinal measure, the cumulative column is directly interpretable here: everyone at upper secondary level or below makes up a majority at 69.7 per cent, and every level short of postgraduate covers 93.9 per cent. The table describes that composition and does not weigh whether it departs from the composition of the population, since the method performs no test at all; such a comparison requires a goodness-of-fit run with official proportions entered as the expected distribution. The report should also state that the 12 non-respondents were taken out of the percentage base. The figures are illustrative.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

Variable examined

General customer survey (synthetic)

A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.

Rows
300
Columns
respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: How is the sample distributed across gender, education and region?

Analyzed columns
gender,education,region
Number of analyzed columns
3

Frequency tables

Frequency tables
RowCountsPercentagesTotal valid responsesDistinct valuesMissing responsesFrequency of the modeCumulative countsCumulative percentages
gender--30020159--
education--30030104--
region--3004078--

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914

How to report the result

The distribution of highest completed education was described as follows: primary 13.7 per cent (86), secondary 24.5 per cent (154), upper secondary 31.5 per cent (198), bachelor's 24.2 per cent (152) and postgraduate 6.1 per cent (38), with levels up to and including upper secondary accounting for a cumulative 69.7 per cent (valid N = 628, excluding 12 non-respondents).

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • It performs no test at all, so the question of whether a distribution departs from an expected one calls for a goodness-of-fit run.
  • Looking at one variable at a time, it cannot show whether a share varies between subgroups, and that question needs a two-variable table.
  • As the count of distinct values rises the table stretches to hundreds of rows holding one or two people each, and the description stops carrying information.
  • The cumulative column is printed even for unorderable labels, and interpreting it on a variable with no order invites a false reading.
  • On select-all questions the frequencies add up beyond the number of participants and the base of the percentages becomes ambiguous.
  • When missing-value codes go undeclared they are counted as ordinary categories and the percentages rest silently on the wrong base.

What to use when the assumptions are not met

  • Chi-Square Goodness of FitWhy: When you move from describing to testing and need to weigh the distribution against an expectation brought in from outside.
  • Cross-TabulationWhy: When the distribution has to be broken down by a second categorical variable so that subgroup shares can be compared.
  • Descriptive StatisticsWhy: When the column is a continuous measure and the mean, median and measures of spread say more than a frequency table would.
  • Multiple Response AnalysisWhy: When the question permits several options to be ticked, since the shares then need a different logic relative to the participant count.

Frequently asked questions

What base are the percentages computed on?
The base is the number of valid observations entering the count, and you are the one who sets it. Declare the missing-value codes on the Variable tab and those answers stay out of the count, so the percentages rest on respondents alone. Leave the codes undeclared and missing answers are counted as an ordinary category, putting the whole sample in the base. Both readings are defensible, but the one you used has to be stated in the report, because the panel raises no warning about it.
Can I always use the cumulative percentage column?
It is printed in every table, yet it only bears interpretation where a meaningful order runs through the categories. On ordinal variables such as education or a five-point agreement scale, the column reduces the portion below a threshold to one number. On unorderable labels such as region or marital status, the same column produces a meaningless running total that depends on the order the rows happen to take, and it should be left alone.
There is a category in the table I did not expect. What is it?
Usually one of two things. The first is a spelling variant: the same category written with and without an initial capital opens two rows and splits the share between them. The second is an undeclared missing-value code, where a marker such as a dash or a nine-nine sequence gets counted as an ordinary category. In either case, return to the Variable tab, correct the value labels or the missing-value codes, and run the table again.
Can I tabulate several columns in one run?
Yes, the parameter form accepts more than one column and produces a separate table for each. That makes it easier to report a survey's demographic block or a group of questions sharing one structure in a single pass, since all the tables land in the same Excel file. Given that the free plan allows fifty analysis runs a month, passing the columns together rather than running them one by one is also the practical choice.

References

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.