Categorical

Cross-Tabulation

Brings two categorical variables into one table and describes the count in every cell alongside row, column and overall percentages.

Method summary

A cross-tabulation is the basic device for making the joint behaviour of two categorical variables visible: the categories of one variable take the rows, those of the other take the columns, and each cell states how many people carry both categories at once. Because raw counts on their own can mislead, the same table is also returned with three separate sets of percentages. Row percentages sum to one hundred across each row and expose the internal composition of the group in that row; column percentages sum to one hundred down each column and show which groups the people in a given outcome category came from; overall percentages relate every cell to the whole sample. Which set you read follows from the direction of your research question, and reading the wrong one can invert the comparison between groups. The method is descriptive: it makes a pattern visible without addressing whether chance could account for it. When a test is wanted, the same two columns go to the chi-square test of independence, while the cross-tabulation serves to get to know the table beforehand and, afterwards, to explain in percentages where the difference comes from.

Which research questions does it answer?

  • How does the preferred device for completing a survey break down by type of settlement?
  • What share does each age band hold within the subscription tiers?
  • Which complaint topics stand out in which customer segments, and how large are their shares?
  • How should the joint distribution of education and news source be presented in the descriptive section of the report?

When should you use it?

  • When the descriptive section of a report needs to show the reader how two classifications sit together.
  • When you want to get to know the shape of the table before any test, seeing how full the cells are and how balanced the categories.
  • When an association already found significant has to be explained in percentages, group by group and direction by direction.
  • When the groups differ markedly in size, since raw counts then distort the comparison and row percentages become necessary.
  • When deciding whether to merge or recode categories and you first need to find out which categories hold very few people.

Required variable types

  • First variable: the categorical column that will form the rows of the table.
  • Second variable: a second categorical column, which will form the columns.
  • Each variable needs at least two categories; a single-category column reduces the table to one strip and leaves nothing to compare.
  • Ordinal variables work as well and the category order is preserved in the table, which eases reading, although the ordering is never turned into a measure of trend.
  • Each row should stand for one participant, since the cells of a pre-aggregated frequency table are not read as people.

Key assumptions

At least nominal measurement
Both variables have to be divisible into categories, meaning each participant lands on exactly one label of each. Feeding a continuous measure straight into the table produces hundreds of one-person cells and nothing readable.
Meaningful, non-overlapping categories
Categories must not overlap, and each should name a group that can be interpreted. Two spellings of the same category open two separate strips in the table and split the percentages between them.
Deliberate handling of missing values
If non-respondents remain as a category of their own, the percentages describe the composition of the whole sample rather than of those who answered. Deciding which reading you want matters for the accuracy of the report.
Cells full enough to interpret
For the description to mean anything, the cells should not be made up of individual people. A row percentage computed from a handful of cases swings by tens of points when a single answer is added.

How YouReply checks these assumptions

  • At least nominal measurement: The measurement level you assign on the Variable tab feeds the method picker: an unsuitable column selection dims the method and states why in plain language. Nothing checks whether a column's number of categories is still interpretable.
  • Meaningful, non-overlapping categories: The panel takes the category labels from the cleaned data exactly as they stand and never merges spelling variants. Merging is yours to do, through value labels or a derived column on the Variable tab.
  • Deliberate handling of missing values: Missing-value codes declared on the Variable tab are filtered out before the table is built. Skip that declaration and those values enter the table as an ordinary category, with no warning raised about it.
  • Cells full enough to interpret: No minimum cell size is required: however few people a cell holds, the table is built and the percentages printed. Which percentages are robust enough to report is for you to settle from the counts in the table.

How the analysis is run

  1. 1Drop your data file onto the panel and confirm in the grid on the Data tab that the columns and the category labels were read correctly.
  2. 2On the Variable tab set the measurement level of the two columns, enter value labels for readable category names, and declare the missing-value codes.
  3. 3Where many sparse categories appear, build a derived column that merges them on conceptual grounds, since the readability of the table depends on it.
  4. 4Find the method through the search box on the Analysis tab; the data requirements box shows which columns can enter the table.
  5. 5In the parameter form choose which variable goes into the rows and which into the columns, and say whether you want the marginal totals included.
  6. 6Run the analysis. The count table and the row, column and overall percentage tables arrive as separate collapsible sections, and no method-specific chart is drawn for this method.
  7. 7Pick the percentage set that matches the direction of your question and read it, following the table headings so the row and column sets do not get confused.
  8. 8Export the tables to Excel, and if you want to test whether chance could account for the association, run the chi-square test of independence on the same two columns.

Statistics and tables produced

Count table
The number of people in each cell who carry that row and that column category together. This is the table to consult when judging how robust the percentages are.
Row percentages
Percentages that sum to one hundred along each row, showing how the group in that row divides across the column categories; comparisons between groups usually come from here.
Column percentages
Percentages that sum to one hundred down each column, showing which row groups the people in a given outcome category came from.
Overall percentages
Percentages relating each cell to the total number of observations, with the whole table summing to one hundred so the weight of each cell in the sample is visible.
Marginal totals
On request, the row and column totals and the grand total are added to the table, which puts the standalone distribution of both variables into the same view.
Observations entering the analysis
The number of valid observations forming the table once missing values have been filtered out, reported so the base of the percentages is clear.

Effect size and confidence intervals

No effect size is produced
Because the method is descriptive, no effect size is computed and no significance test is performed; neither Cramer's V nor an odds ratio nor anything comparable appears in the result. If you need the strength of an association summarised in a single number, run the chi-square test of independence on the same columns.
Percentages as descriptive quantities
What stands in place of an effect size here is the comparison between percentages. For a given column category, the gap in points between two rows conveys the size of the pattern directly and in terms a reader grasps at once: where row percentages run from 55.1 to 33.9 per cent, the 21 points between them make a divergence worth describing. That gap is not a standardised measure, so it cannot be compared across different tables and carries no verdict on significance. With few people per cell the same gap is far more fragile, which is why percentages belong in a report next to the count table.

No confidence interval is returned for this method. The cell percentages and the row and column shares are all point estimates, and no band of uncertainty is computed around any of them. Reporting that calls for an interval around a proportion has to obtain it from the proportion test or from a calculation outside the panel. With no interval printed, how much a percentage from a small cell might move can only be judged from the count table.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
In an illustrative user study, how does the channel people prefer for answering a survey invitation break down by type of settlement?
Variables
Row variable: type of settlement (metropolitan, city, town) · Column variable: preferred channel (mobile app, electronic mail link, text message link)
Example result
The table was built from 512 valid responses. Among the 214 metropolitan respondents, 118 chose the mobile app, 64 the electronic mail link and 32 the text message link; for the 186 city respondents the counts were 82, 66 and 38, and for the 112 town respondents 38, 34 and 40. Row percentages printed as 55.1 / 29.9 / 15.0 in metropolitan areas, 44.1 / 35.5 / 20.4 in cities and 33.9 / 30.4 / 35.7 in towns. Column totals were 238, 164 and 110, giving overall channel shares of 46.5, 32.0 and 21.5 per cent.
Interpretation
Read across the rows, preference for the mobile app falls steadily from metropolitan areas to towns while the text message link rises the other way, and the gap between the two extremes exceeds 21 points for the app. This table is a description and says nothing about whether chance could account for the pattern, since the method yields neither a p value nor an effect size. Tying the finding to a test means running the chi-square test of independence on the same two columns. Switching to the column percentages changes the reading: it becomes clear that most people choosing the text message link are not from towns, and that all three groups contribute comparable counts to that column because the town sample is the smallest. Stating in the report which set you used keeps the reader from making the wrong comparison. The figures are illustrative.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

Variable examined

General customer survey (synthetic)

A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.

Rows
300
Columns
respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: How does the five point satisfaction level spread across the education groups?

Row variable
education
Column variable
satisfaction_level
Include the margins
true
Valid observations
300
row_column
education
col_column
satisfaction_level
Grand total
300

Contingency table

Contingency table
Row12345Total
Bachelor113314114100
High school012363818104
Postgraduate0832401696
All1339911948300

Row percentages

Row percentages
Row12345
Bachelor113314114.00
High school011.5434.6236.5417.31
Postgraduate08.3333.3341.6716.67

Column percentages

Column percentages
Row12345
Bachelor10039.3931.3134.4529.17
High school036.3636.3631.9337.50
Postgraduate024.2432.3233.6133.33

Percentages of the total

Percentages of the total
Row12345
Bachelor0.3334.3310.3313.674.67
High school041212.676
Postgraduate02.6710.6713.335.33

Row totals

Bachelor
100
High school
104
Postgraduate
96

Column totals

1
1
2
33
3
99
4
119
5
48

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914

How to report the result

Preferred channel was described by type of settlement: the mobile app was chosen by 55.1 per cent of metropolitan respondents (118 of 214), 44.1 per cent of city respondents (82 of 186) and 33.9 per cent of town respondents (38 of 112), N = 512; percentages are row percentages and no significance test was performed.

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • Since it performs no test, it cannot on its own support a claim about whether chance explains a pattern, and that question needs a separate run.
  • With no effect size computed, the patterns in two different tables cannot be compared directly, because point gaps are not standardised.
  • As the number of categories grows the table grows quickly and the cells thin out, which lends the percentages a misleading air of precision.
  • Even when the categories are ordered, the ordering is not converted into a measure of trend and the table only reports shares.
  • The joint distribution of three or more variables cannot be built in a single run; nested readings need separate runs or a banner layout.
  • Comparing groups on a single percentage turns into a strong claim when the cells hold few people, and that fragility stays invisible in the table.

What to use when the assumptions are not met

  • Chi-Square Test of IndependenceWhy: When you move from description to testing and need to judge whether chance can account for the association, working from the same two columns.
  • Banner TableWhy: When many questions have to be reported in one layout against the same breakdown variables, with group comparisons marked by letters.
  • Frequency DistributionWhy: When the interest is the distribution of one variable on its own rather than a breakdown by a second variable.
  • Multiple Response AnalysisWhy: When the question being broken down allows several answers, so the cell totals exceed the participant count and a different counting logic applies.

Frequently asked questions

Should I report the row percentage or the column percentage?
It depends on the direction of your question. If you are comparing groups with one another, that is, showing the internal make-up of each group, put the groups in the rows and read the row percentages. If you are profiling the people who landed in a particular outcome category, the column percentages are what you want. The two sets describe the same table yet produce different sentences, so naming the one you used keeps your reader from drawing the wrong comparison.
Why is there no p value in this table?
The method is descriptive by design: it shows the joint distribution in counts and percentages without stating or testing a hypothesis. That is why the result holds no test statistic, no p value and no effect size. To judge whether the divergence you see could be explained by chance, hand the same two columns to the chi-square test of independence, which prints expected counts next to the observed table.
Should I switch the marginal totals on?
Marginal totals let you see each variable's standalone distribution in the same table and make it easier to follow which base a percentage was computed on, so they are usually worth having in tables destined for a report. If you are embedding the table in running text, or the category count is high, a total row and column can add unnecessary width. The option changes nothing in the numbers, only the shape of the table.
One cell holds only three people. Can I report its percentage?
The panel requires no minimum cell size and will compute and print that percentage, but a percentage resting on three people moves by tens of points with a single extra answer. Common practice in such cells is to give the count instead of the percentage, or to merge the category with another one on defensible conceptual grounds. Whichever you choose, reporting the count table alongside the percentages lets the reader see the fragility.

References

  • Agresti, A. (2018). An Introduction to Categorical Data Analysis
  • Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics
  • pandas.crosstab documentation

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.