Reliability
Inter-Rater Agreement
Measures how far people coding the same material reach the same decision, using kappa coefficients corrected for chance and the intraclass correlation.
Method summary
Inter-rater agreement asks whether a measurement stands independently of who produced it. Sorting open-ended answers into themes, placing interview passages into categories or scoring a product image against criteria are all decisions made by human judgement, and two people assigning the same code to the same material is not something to take for granted. Percentage agreement misleads here, because even people assigning codes at random land on the same answer a certain share of the time; kappa coefficients subtract that chance share and report what agreement remains. With numeric scoring the question is less about matching categories than about scores tracking each other, and the intraclass correlation is what captures that. YouReply Analyze gathers all three measures in one method. With exactly two raters it computes Cohen's kappa, a one-sided p value for it and an agreement table setting the two sets of decisions against each other; with three or more raters Fleiss' kappa takes over. The intraclass correlation is computed for numeric data and comes back as two figures, one for a single measurement and one for the average. Which measures to request is a parameter, and by default all are asked for, in which case a measure that does not fit the number of raters is skipped silently rather than raising an error. Kappa values are labelled with the Landis and Koch bands and the intraclass correlation with those of Koo and Li.
Which research questions does it answer?
- Do two coders sorting open-ended answers into themes reach the same decisions?
- How much agreement beyond chance is there among three experts categorising interview passages?
- Can the numeric marks given by assessors rating the same presentations stand in for one another?
- Did inter-coder agreement improve after the coding manual was revised?
When should you use it?
- When a measurement is produced by human judgement and its independence from the individual coder has to be demonstrated.
- When categorising qualitative material and the clarity of the coding manual needs to be evidenced.
- When at least two coders have independently coded all the units or a shared subset of them.
- When percentage agreement is not enough on its own, which chance correction makes unavoidable as soon as one category dominates.
Required variable types
- Rater columns: one column per coder, each row holding the code that coder gave the same unit. At least two columns are required.
- Rows are the coded units: an open-ended answer, an interview passage, an image. The order of units has to be identical across all columns.
- Kappa suits categorical codes and the intraclass correlation suits numeric ratings, and both can be requested for the same set of columns.
- Code labels have to be spelled identically across columns, since the same category written two ways counts as a disagreement.
Key assumptions
- Coders working independently
- Each coder must decide without seeing the other's code. Material coded by discussion around a table shows artificially high agreement, and the coefficient then reflects a process of reaching consensus rather than the reliability of the measurement.
- A fixed set of categories
- Kappa assumes every coder chose from the same list of categories. When one uses a code absent from the list, or merges categories on their own initiative, the basis for comparison shifts.
- Reasonably balanced categories
- When coders place nearly every unit in one category, kappa comes out unexpectedly low, because the chance agreement is computed as very high. This is the kappa paradox: percentage agreement above ninety can sit alongside a merely moderate kappa.
- All disagreements weighted alike
- Unweighted kappa treats every disagreement as equally serious. With ordered categories that is a problem: a gap between two adjacent steps is penalised exactly as heavily as one between the two extremes.
- The choice of intraclass correlation form
- The literature distinguishes several forms of the intraclass correlation, separated by whether the raters were randomly chosen, whether consistency or absolute agreement is at stake, and whether the result concerns a single rater or an average.
How YouReply checks these assumptions
- Coders working independently: Whether the coding was independent cannot be verified; the engine sees only the codes in the columns and has no way of knowing how they came about. This rests entirely on how the research was run.
- A fixed set of categories: Code labels are compared across columns but spelling differences are never flagged: the same category written differently in two columns silently counts as a disagreement. Review the labels on the Data tab and make them match exactly.
- Reasonably balanced categories: The balance of the category distribution is never assessed and no warning accompanies a paradoxical result. With two raters the agreement table is the practical way to see it, since the decisions piling into a single cell of the diagonal are visible there.
- All disagreements weighted alike: Only unweighted kappa is computed; there is no linear or quadratic weighting option for ordered categories. If your codes are ordered, the coefficient understates the agreement, and that has to be stated as a limitation.
- The choice of intraclass correlation form: The engine computes one form of the intraclass correlation and offers no choice among them; what it separates is only the single-measure and average-measure figures. Where a publication asks for a specific form, it has to be computed outside the panel.
How the analysis is run
- 1Lay your coding data out with one column per coder and one row per coded unit, then upload it.
- 2Compare the code labels on the Data tab and make sure each category is spelled identically in every column, including capitalisation and stray spaces.
- 3Set the measurement level of the rater columns on the Variable tab: nominal for categorical codes, scale for numeric ratings.
- 4Choose the method from the reliability category on the Analysis tab and tick the rater columns in the parameter form.
- 5The measure setting can stay at its default, in which case whatever fits your number of raters is computed and whatever does not is skipped silently and never appears.
- 6Run it. With two raters, read the agreement table before interpreting the coefficient, because where the disagreements concentrate tells you which part of the coding manual is ambiguous. No chart is produced here, so export the tables to Excel.
Statistics and tables produced
- Cohen's kappa
- Computed when there are exactly two raters, giving agreement net of chance. With more than two raters this measure is not calculated and does not appear in the output.
- One-sided p value for Cohen's kappa
- The significance of agreement exceeding zero, which is to say exceeding chance level. The test says nothing about the agreement being adequate, only about it being better than chance.
- Agreement table
- A cross-tabulation setting the two coders' decisions against each other. Cells on the diagonal hold the agreed decisions and the rest hold disagreements, which is where you see which pair of categories gets confused.
- Fleiss' kappa
- The multi-rater coefficient computed when three or more raters are present. Individual pairs of coders are not broken out, and a single value covers all of them.
- Intraclass correlation
- Computed for numeric coding, it expresses how far the ratings track each other across coders. The single-measure figure concerns relying on one coder's score, the average-measure figure on the mean across coders.
- Interpretation label for kappa
- A label from the Landis and Koch bands: slight below 0.21, fair below 0.41, moderate below 0.61, substantial below 0.81 and almost perfect above that. The bands are a convention in the literature rather than firm boundaries.
- Interpretation label for the intraclass correlation
- A label from the Koo and Li bands: poor below 0.5, moderate below 0.75, good below 0.9 and excellent above. The single-measure and average-measure figures can fall in different bands.
- Measures computed
- Which measures come back depends on the number of raters and the type of data. Under the default setting an unsuitable measure is skipped without an error, so a missing coefficient is a prompt to re-examine your rater count and data type.
Effect size and confidence intervals
- Kappa
- Kappa is the agreement left once the share attributable to chance has been deducted, and it reads directly as a magnitude. Zero means agreement at chance level and one means perfect agreement, while negative values indicate coders doing worse than chance, usually because the manual was read in opposite ways. The common publication expectation is above 0.60, rising to 0.80 for an established coding manual. Remember that the bands are sensitive to distributions: with unbalanced categories the same level of real agreement produces a lower kappa.
- Intraclass correlation
- It summarises how much of the variation in ratings comes from genuine differences between the coded units and how much from differences between coders. The single-measure figure gives the reliability of relying on one coder's rating and the average-measure figure that of using the mean across coders, with the latter always the larger of the two. Which to report follows from your design: if the final dataset uses one coder's score, the single-measure figure is the relevant one, and if it uses the mean, the average-measure figure is.
No interval is returned for kappa or for the intraclass correlation; the output holds point estimates, interpretation labels and, with two raters, the one-sided p value for kappa. The omission matters most with small coding sets: a kappa computed on fifty units carries wide uncertainty, and which side of the 0.61 mark it falls on can turn on the sample. Since most journals expect 95% bounds for kappa and for the intraclass correlation, those have to be computed outside the panel. It is also worth stressing that the p value is no substitute for an interval: that test says agreement exceeds chance, not that agreement is acceptable.
Example research question and example result
The numbers below are a representative example, not data from a real study or a real user.
- Research question
- In an illustrative qualitative study, is the agreement between two coders sorting open-ended answers into five themes good enough to work with the coding manual?
- Variables
- Rater 1: the theme code the first researcher gave each answer · Rater 2: the theme code the second researcher gave the same answers · Units: 240 open-ended answers coded independently
- Example result
- Cohen's kappa was 0.68 (one-sided p < .001, label: substantial agreement) with observed percentage agreement of 76. Of the 34 disagreements across the 240 units, 21 fell between the themes 'speed of service' and 'staff attitude'. With only two raters, Fleiss' kappa was not computed, and because the codes were categorical no intraclass correlation was returned either.
- Interpretation
- At 0.68 kappa lands in the substantial band of Landis and Koch and clears the 0.60 mark often used as a publication threshold, so the coding manual is workable. The distance between 76% observed agreement and a kappa of 0.68 shows how much the chance correction absorbs and why the percentage alone would not do. More instructive still is the agreement table: two thirds of the disagreements concentrate between just two themes, meaning the trouble lies at one boundary rather than throughout the manual. The practical step is to sharpen the definitions of those two themes with distinguishing criteria and recode the contested answers. Because no interval accompanies the coefficient, how safely this estimate from 240 units sits above the 0.61 mark remains unknown.
Real output on a sample dataset
The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.
General customer survey (synthetic)
A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.
- Rows
- 300
- Columns
- respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.
Research question: Do the three raters place the same answers in the same categories, and is agreement among three raters measured with Fleiss kappa?
- Rater columns
- rater_1,rater_2,rater_3
- Valid observations
- 300
- Number of raters
- 3
- Fleiss' kappa
- 0.582
- Intraclass correlation, single rater ICC(2,1)
- 0.609
- Intraclass correlation, average of raters ICC(2,k)
- 0.824
Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914
How to report the result
Agreement between the two coders was substantial (Cohen's kappa = .68, p < .001; observed agreement 76%; 240 units), with most disagreements concentrated between two adjacent themes.
An example sentence close to APA style; the numbers are representative.
When you should not use it
- No confidence interval is returned for kappa or the intraclass correlation, so with a small coding set it stays unclear which side of a threshold the coefficient really falls on.
- Only unweighted kappa is computed, which means that with ordered categories a disagreement between adjacent steps is penalised as heavily as one between the extremes.
- A single form of the intraclass correlation is computed and no choice among forms is offered, so publication requirements naming a specific form cannot be met inside the panel.
- Kappa responds to the category distribution: with one dominant category a high percentage agreement can accompany a low kappa, and nothing warns you about it.
- With three or more raters only one Fleiss value is returned; which pair of coders diverges cannot be seen and no per-coder agreement is reported.
- Because an unsuitable measure is skipped silently under the default setting, a missing coefficient announces itself through careful reading of the output rather than through an error message.
What to use when the assumptions are not met
- Cronbach's AlphaWhy: When coders give numeric ratings and each can be treated as a source of measurement, the item-deletion table also exposes the weaker coder.
- Split-Half ReliabilityWhy: When the reliability question concerns sets of items within one instrument rather than the people applying it.
- Chi-Square Test of IndependenceWhy: When the question is not how far coders agree but whether their category distributions differ from one another.
- Spearman Rank CorrelationWhy: When coders assign ordered scores and what matters is that the rankings coincide rather than that the scores match exactly.
Frequently asked questions
- Why should I use kappa rather than percentage agreement?
- Percentage agreement counts coincidence as agreement, and that share is larger than intuition suggests. With two categories, even wholly random decisions put two people on the same code for about half the units, so anything near 50% proves nothing at all. Kappa works that share out and deducts it, leaving the agreement above chance level. The discrepancy grows when the category distribution is skewed: if nine units in ten belong to one category, high percentage agreement is nearly guaranteed and reporting it alone misleads the reader. The practical recommendation is to give both, because the percentage puts the size of kappa in context and lets the reader see what happened in a paradoxical case.
- My kappa is low but my percentage agreement is high, how can that be?
- This is the kappa paradox, and an unbalanced category distribution is almost always behind it. When your coders put the bulk of the units in one category, chance agreement is computed as very high, and kappa stays small despite high observed agreement. What to do depends on the question. Where the sparse categories are theoretically close, merging them can rebalance the distribution. Where a sparse category is central to the research, the honest route is to report the low kappa while supplying the agreement table, since the cell counts let the reader see why the coefficient is small. Coding more units also yields a steadier estimate in the thin category.
- Should I use kappa or the intraclass correlation?
- The distinction turns on what the code is. When coders pick one of several categories that cannot substitute for each other, with no magnitude relation between codes, kappa is the right measure. When they assign a numeric score, marking a presentation out of ten for instance, what matters is not exact matching but the scores moving together, and the intraclass correlation captures that. The awkward middle case is ordered categories, where a five-point severity grade carries both category and rank. The literature recommends weighted kappa there, but the panel offers only the unweighted version, which penalises adjacent-step disagreements more heavily than it should. For such data the soundest course is to request both measures, weigh the results together, and write up the absence of weighting as a limitation.
- How many units do I need to double-code?
- Having two coders work through all the material gives the strongest evidence, but on large datasets that is expensive and common practice is to double-code a subset. The criterion often cited is between 10% and 20% of the units, with a floor of around fifty units suggested, since below that kappa becomes so volatile that the reported figure misleads. The subset has to be drawn at random and to represent the material as a whole, because double-coding only the easy answers inflates agreement artificially. As the panel returns no interval, you have to infer the precision of the estimate from the number of units yourself, which is a further reason to keep the double-coded subset generous.
References
- Landis, J. R., & Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data
- Koo, T. K., & Li, M. Y. (2016). A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research
- Agresti, A. (2013). Categorical Data Analysis
- statsmodels inter_rater documentation
Try it with your own data
The free plan includes 50 analysis runs a month and needs no card.