Upload a CSV, map one column per rater, and get Fleiss' kappa with a confidence interval, Krippendorff's alpha, a per-category breakdown, and the one rater whose labels are dragging the panel down. Free.
Free analyses run on up to 10,000 rows. Larger files are randomly sampled to that size — sign up to analyze your full dataset.
Building the item-by-category matrix and computing Fleiss' kappa...
Sent to — Fleiss' kappa with its bootstrap interval and benchmark band, Krippendorff's alpha in nominal and ordinal form, the per-category kappa breakdown, each rater's agreement with the consensus plus leave-one-out kappas, the item-level consensus distribution, R code, and AI insights.
Analyze another fileThe analysis reduces either input shape to an item-by-category matrix of how many raters put each item in each category, then works on that. Raw pairwise agreement is the mean over items of the share of rater pairs choosing the same category; expected agreement is the sum of squared overall category shares; Fleiss' kappa is the observed excess over that expectation as a fraction of the amount available. Fleiss' classical null-hypothesis variance gives a z-test against chance-level agreement, and is reported as belonging to that test only — the 95% intervals come from a seeded nonparametric bootstrap that resamples items, which is the correct sampling unit and needs no complete-grid assumption. Krippendorff's alpha is computed from the coincidence matrix as one minus observed over expected disagreement, on every item carrying at least two judgements, in nominal form and — when the labels are detected to be ordered — in ordinal form as well. Category-specific kappas, per-rater agreement with the consensus, leave-one-rater-out kappas and the item-level consensus distribution are all derived from the same matrix in closed form.
Use it whenever three or more people, systems, or passes assign a category to the same items and you need to know how much of their agreement is real, or which of them is out of step: annotation and labelling QA, multi-reader diagnostic studies, content-moderation panels, peer review, survey coding, or grading against a rubric.
Not for continuous or near-continuous ratings — scores, measurements, times — where an intraclass correlation is the right tool and any kappa would throw the scale away. Not for exactly two raters, where Cohen's kappa is the standard and this generalisation buys nothing. Not for comparing raters against a known-correct answer, which is an accuracy question rather than an agreement one.
Built for: Annotation and ML data teams, clinical and research coders, trust-and-safety and QA leads, editors running multi-reviewer panels, and anyone reporting inter-rater reliability for a categorical rubric across more than two raters
Typical data source: A spreadsheet with either one column per rater and one row per item, or one row per rating with item, rater and judgement columns
One row per item, one column per rater. For example, four reviewers grading the same submissions:
Minimum 20 rows · Best with 50-10,000 items, 3-8 raters and 2-8 categories
Standard-library analysis: three or more raters, one categorical judgement per item — how much do they really agree, and who is out of step? Map one column per rater (or, if your file has one row per rating, map the item, rater and judgement columns — the shape is detected from the data) and get Fleiss' kappa with a bootstrap confidence interval and a test against chance, Krippendorff's alpha in nominal and — when the labels turn out to be ordered — ordinal form, a per-category kappa showing exactly which categories the panel cannot pin down, a per-rater agreement with the consensus plus a leave-one-rater-out kappa that names the outlier, the item-level distribution of how strong the consensus actually was, and an explicit account of what each coefficient did with any missing ratings instead of quietly dropping them.
Pairwise agreement, chance agreement, Fleiss' kappa with its bootstrap interval, and Krippendorff's alpha in nominal and — where the labels are ordered — ordinal form.
Each rater's agreement with the rest of the panel, and what the panel's kappa would be without them — the fastest route from a reliability number to a concrete conversation.
Category-specific kappa — which parts of the rubric the panel reads the same way and which it does not.
How many items drew a unanimous verdict, a majority, or no majority at all; the no-majority items are the work list.
How often each rater reaches for each category; the distributions the chance correction is built from, and the evidence behind any kappa paradox or outlier.
The formulas in full, which input shape and which orderedness rule fired, and the three things these coefficients cannot tell you.
Plain-English interpretation — what the numbers mean, what's significant, and what to do next.
Do our five annotators actually agree?
Map each annotator's column. You get Fleiss' kappa with a bootstrap confidence interval and a benchmark reading, Krippendorff's alpha alongside it, a per-category breakdown showing which parts of the rubric the panel reads differently, and how many items drew a unanimous verdict versus no majority at all.
See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.
Run any analysis on your own data — validated R analyses, interactive reports, AI insights, and PDF export.
Try Free — No Credit CardTell us what went wrong, in your own words. We capture the page you're on automatically, so no need to describe where you are.