Standard Agreement Multi
Executive Summary

Executive Summary

Fleiss' kappa 0.484 (moderate) across 4 raters on 200 items

Items Rated
200
Raters
4
Categories
4
Raw Agreement
63.2%
Fleiss Kappa
0.484
Fleiss 95% CI
0.426 to 0.535
Krippendorff Alpha
0.485
Agreement Strength
moderate
Least Aligned Rater
Reviewer D Rating
Kappa vs Chance p
< 0.001
4 raters judged 200 items and agreed on 63.2% of rater pairs. Fleiss' kappa is 0.484 (95% CI 0.426 to 0.535), which the Landis & Koch convention calls MODERATE agreement, and which is clearly better than chance (z = 24.45, p < 0.001). That interval rests on 200 item(s). Krippendorff's alpha, which reaches the same question from observed against expected disagreement, gives 0.485. Chance alone would have produced 28.7% pairwise agreement given how often each category is used, leaving 71.3% of the scale available above chance, of which 48.4% was achieved. Reviewer D Rating is the least aligned member of the panel (rater kappa 0.273 against 0.584 for Reviewer A Rating); with that rater excluded the panel's kappa would be 0.690, 0.205 higher than it is now. 75 of 200 items (37.5%) drew a unanimous verdict, and 53 (26.5%) failed to produce a majority at all. One caution that applies whatever the number: these coefficients measure whether the panel makes the same call, not whether the call is right — a panel can agree unanimously and be unanimously wrong, and only a gold standard can tell you which.
Suggested Interpretation

The short answer

The four reviewers agreed on 63.2% of rater pairs across 200 items, with Fleiss' kappa of 0.484 (95% CI 0.426 to 0.535)—moderate agreement and clearly better than chance (p < 0.001). However, Reviewer D is notably out of step; without that rater, the panel's kappa would rise to 0.690.

The detail

Fleiss' kappa is 0.484, which the Landis & Koch convention classifies as moderate. Krippendorff's alpha gives 0.485, confirming the result. Chance alone would have produced 28.7% agreement given the panel's category habits, leaving 71.3% of the scale available; the panel achieved 48.4% of that headroom. The z-statistic is 24.45 with p < 0.001, far better than random. Reviewer A is most aligned (rater kappa 0.584); Reviewer D is least aligned (rater kappa 0.273), a gap of 0.311. Of the 200 items, 75 (37.5%) drew unanimous agreement, while 53 (26.5%) produced no majority at all. Critical caveat: these coefficients measure whether the panel agrees, not whether the panel is right.

What this can't tell you

Agreement does not imply correctness. A unanimous panel can be unanimously wrong; only a gold standard or external validation settles whether the judgements are accurate.

Overview

Analysis Overview

Fleiss' kappa and Krippendorff's alpha on 200 items judged by 4 raters across 4 categories.

N Items200
N Raters4
N Categories4
Pairwise Agreement Pct63.2
Fleiss Kappa0.484
Suggested Interpretation

The short answer

The four reviewers show moderate agreement at 0.484 (Fleiss' kappa), well above what random chance would produce. However, raw agreement of 63.2% looks higher than it is because the panel leans heavily on certain categories—a chance baseline of 28.7% means they'd collide on those categories even without reading carefully. Krippendorff's alpha reaches the same answer (0.485) from a different angle, asking about disagreement instead.

The detail

Fleiss' kappa of 0.484 (95% CI 0.426 to 0.535) captures the 48.4% of available headroom above chance that the panel actually won. Raw pairwise agreement was 63.2%, but the panel's own category habits would have produced 28.7% agreement by luck alone, leaving 71.3% of the scale available. Kappa divides the achieved agreement by that available headroom. Krippendorff's alpha (0.485) computes the same idea through observed versus expected disagreement and tolerates missing ratings. With exactly 4 raters and 4 categories, both coefficients apply; with 2 raters you would use Cohen's kappa; with continuous scores, an intraclass correlation.

What this can't tell you

These coefficients measure consistency of calls, not correctness. A panel can agree unanimously and be wrong; only comparison to a gold standard settles that.

Data Preparation

Data Quality

200 items and 800 judgements from 200 rows loaded.

Initial Rows200
Items200
Raters4
Judgements Found800
Missing Ratings0
Items On Basis200
Judgements On Basis800
Suggested Interpretation

The short answer

All 200 items were complete: every one of the 4 raters judged every item, producing 800 judgements with no gaps. The data came in wide form (one row per item, one column per rater) and was correctly recognized because each rater column held only 4 distinct labels rather than unique values. Both Fleiss' kappa and Krippendorff's alpha were computed on the same full dataset.

The detail

200 rows resolved to 200 items, 4 raters, and 800 judgements across 4 categories. No ratings were missing: every rater judged every item, so both coefficients used exactly the same 200 items and 800 judgements. The 95% confidence interval around kappa runs from 0.426 to 0.535, a width of 0.109, based on these 200 items. Labels were trimmed of surrounding spaces before comparison, and blanks or placeholders were treated as missing judgements rather than as a category.

What this can't tell you

The interval width (0.109) reflects the sample size of 200 items; a smaller item pool would widen it. The completeness of the data means both kappa and alpha use identical item sets, so any visible gap between them would point to different methods rather than different data.

Data Table

Agreement Statistics

Pairwise agreement, chance agreement, Fleiss' kappa and Krippendorff's alpha with bootstrap intervals.

StatisticEstimateCI LowCI HighInterpretation
Raw pairwise agreement0.6325Across every pair of raters who judged the same item, 63.2% of those pairs chose the same category.
Agreement expected by chance0.2871What a panel using the categories this often would have agreed on by luck alone (28.7%).
Fleiss' kappa0.48450.42580.534648.4% of the agreement left available above chance was achieved, on the 200 item(s) every one of the 4 raters judged — moderate on the Landis & Koch convention.
Krippendorff's alpha (nominal)0.48510.42450.5448The disagreement-based coefficient, computed on all 200 items with two or more judgements — it does not require every rater to have judged every item.
Krippendorff's alpha (ordinal)0.77190.72290.815The same coefficient with an ordinal distance between categories, so a near-miss between neighbouring levels counts as a smaller disagreement than a jump across the scale.
Suggested Interpretation

The short answer

The panel agreed on 63.2% of rater pairs, but chance would have delivered 28.7% on its own given how often each category is used. Fleiss' kappa reports that 48.4% of the remaining headroom was actually won. The 95% confidence interval of 0.426 to 0.535 comes from resampling items, not from a formula, because the classical Fleiss standard error describes a null hypothesis and would misstate the interval around a non-zero estimate.

The detail

Raw pairwise agreement is 0.6325 (63.2%). Expected agreement by chance is 0.2871 (28.7%). Fleiss' kappa is 0.4845 (48.4% of available headroom), with 95% CI from 0.4258 to 0.5346 on all 200 items. Krippendorff's alpha (nominal) is 0.4851 with CI 0.4245 to 0.5448. Ordinal alpha is 0.7719 (CI 0.7229 to 0.815) because most disagreements here fall between neighbouring categories rather than across the full scale. The Fleiss and nominal alpha estimates differ by 0.001, confirming they are different arithmetic routes to the same complete dataset. Ordinal alpha will always be more flattering when categories carry an order.

What this can't tell you

The Landis & Koch bands are convention, not a statistical standard. Whether 0.484 is fit for purpose depends on what the judgements decide—that is your call, not the data's.

Visualization

Which Rater Is Out of Step

Each rater's agreement with the rest of the panel, and the panel's kappa without them.

Suggested Interpretation

Reviewer A Rating leads at 70.3% raw agreement with the rest of the panel and a rater kappa of 0.584. Reviewer D Rating trails at 48.17% raw agreement and a rater kappa of 0.273—a gap of 0.311 of kappa, large enough that particular people, not the rubric as a whole, are holding the headline number down. Removing Reviewer D Rating would lift the panel's kappa from 0.484 to 0.690, a gain of 0.205. Treat that as diagnostic: dropping the rater who disagrees most always raises measured agreement, even when that rater is the one who is right. A low rater-level score is associated with, not proof of, misreading the rubric; the same pattern appears when one rater applies it correctly and the others do not.

Visualization

Which Categories the Panel Fights Over

Category-specific kappa across 4 categories.

Suggested Interpretation

Disagreement is concentrated rather than spread evenly: "Excellent" is the most reliable at 0.644, while "Fair" is the least at 0.368—a spread of 0.276. Fixing "Fair" (which absorbs 25.75% of all judgements) would move the headline number more than a general calibration session. "Good" (39.25% of judgements) scores 0.4059, and "Poor" (12.25% of judgements) scores 0.6046. A category that one rater reaches for freely and others barely touch scores low here even when overall kappa looks healthy, which is exactly the detail a single summary coefficient hides. Rubric work aimed at "Fair" and "Good" would be the highest-leverage intervention.

Visualization

How Often the Panel Reached a Verdict

Consensus strength across 200 items on the analysis basis.

Suggested Interpretation

Disagreement is not spread evenly: 75 items (37.5%) drew unanimous verdicts, 72 (36.0%) produced a majority without unanimity, and 53 (26.5%) produced no majority at all. Those 53 no-majority items are the concrete work list—they are where the rubric is silent or the item is genuinely ambiguous. Re-reading a sample of them usually explains the headline number faster than the coefficient itself does. Consensus bands are set from each item's largest block of agreeing raters, so an evenly split item needs no arbitrary winner. With only 4 raters the possible values are coarse; empty bands mean the arithmetic cannot land there rather than that no item did.

Visualization

How Each Rater Uses the Scale

Category shares per rater, side by side.

Suggested Interpretation

The busiest category, "Good", takes 39.2% of all judgements—a reasonably even spread across raters (41.5%, 41%, 40.5%, 34%) that keeps the chance-agreement floor at 28.7% and lets kappa stay closer to the raw agreement of 63.2%. Reviewer D Rating stands visibly apart: it reaches for "Poor" (22%) and "Fair" (34%) much more often than the others (9–11% and 18.5–28.5% respectively), and "Excellent" much less often (10% vs 24–29%). This is a different threshold, not random error, and will depress agreement even when each individual judgement looks defensible. Reviewers A, B, and C are much more aligned in their use of the scale.

Data Table

Methods & Disclosure

Every formula behind the numbers, and what they cannot decide.

ItemDetail
Design200 item(s) judged by 4 rater(s) across 4 categories, 800 judgement(s) in total.
Input shapeDetected from the data, not from the mapping: the 4 mapped columns were read as one column per rater, because each holds only a handful of repeated labels (4, 4, 4, 4 distinct) rather than one value per row.
Raw pairwise agreementMean over items of the share of rater PAIRS on that item choosing the same category: 63.2%.
Chance agreementSum of the squared overall category shares (Poor 12.2%, Fair 25.8%, Good 39.2%, Excellent 22.8%): 28.7%.
Fleiss' kappa(observed - expected) / (1 - expected) = (0.632 - 0.287) / (1 - 0.287) = 0.484, computed on the 200 item(s) every one of the 4 raters judged.
Significance testFleiss' null-hypothesis standard error is 0.020, giving z = 24.45 and p < 0.001 against the hypothesis that agreement is no better than chance. That variance describes kappa when kappa is truly zero, so it belongs to this test and NOT to the confidence interval.
Confidence intervalsA seeded nonparametric bootstrap resampling ITEMS with replacement (500 resamples, fixed seed), taking the 2.5th and 97.5th percentiles. The bootstrap standard error of kappa is 0.030. This is used rather than a closed form because the null variance above is the wrong quantity for an interval around a non-zero estimate, and alpha has no simple closed-form variance.
Missing ratingsNo ratings are missing: every rater judged every item, so Fleiss' kappa and Krippendorff's alpha are computed on exactly the same data.
Krippendorff's alphaBuilt from the coincidence matrix: alpha = 1 - observed disagreement / expected disagreement, with (n-1) weighting so it is defined for small samples. Nominal alpha treats every disagreement as equal. Ordinal alpha weights a disagreement by the distance between the categories along the detected order.
Category orderOrderedness was detected, not assumed: the labels are all points on a quality scale, so they were ordered along it.
Per-category kappaFleiss' category-specific coefficient, generalized to a variable number of raters per item: one minus the observed splits on that category over the splits expected from its overall share.
Per-rater figuresA rater's agreement figure is the share of the OTHER raters' judgements on the same items that matched this rater's own, chance-corrected against the same 28.7% floor. The leave-one-out column recomputes Fleiss' kappa for the panel with that rater removed.
Benchmark labelsLandis & Koch (1977) labels (slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.
What agreement is notAgreement is not correctness: a panel can agree unanimously and be unanimously wrong. These figures describe this panel on these items and do not generalise to other raters or other items.
Suggested Interpretation

The short answer

Kappa is computed from a simple formula on the 200-by-4 matrix of category counts: (observed agreement – expected agreement) / (1 – expected agreement). The confidence intervals come from resampling items 500 times, not from a closed-form formula, because the null-hypothesis standard error (0.020, giving z = 24.45) describes what kappa looks like when it is truly zero and does not apply to an interval around a non-zero estimate.

The detail

Raw pairwise agreement: mean share of rater pairs per item choosing the same category = 63.2%. Chance agreement: sum of squared category shares (Poor 12.2%, Fair 25.8%, Good 39.2%, Excellent 22.8%) = 28.7%. Fleiss' kappa = (0.632 − 0.287) / (1 − 0.287) = 0.484 on all 200 items. Significance: null standard error 0.020, z = 24.45, p < 0.001. Bootstrap intervals: 500 resamples of items with replacement, fixed seed, 2.5th and 97.5th percentiles; bootstrap SE = 0.030. Krippendorff's alpha uses the coincidence matrix: alpha = 1 − (observed disagreement / expected disagreement), with (n−1) weighting. Ordinal alpha applies category distance weights. Category order was detected from labels, not assumed.

What this can't tell you

Agreement is not correctness; no coefficient here validates whether the panel is right. These figures describe this panel on these items only—different raters or a different item pool with different category prevalence will produce different kappa values. The confidence intervals are approximate and widen sharply with small item counts.

Methodology

Methodology

Statistical methodology and diagnostics for Multi-Rater Agreement — Fleiss Kappa

Statistical Method

Multi-Rater Agreement — Fleiss Kappa

Standard-library analysis: three or more raters, one categorical judgement per item — how much do they really agree, and who is out of step? Map one column per rater (or, if your file has one row per rating, map the item, rater and judgement columns — the shape is detected from the data) and get Fleiss' kappa with a bootstrap confidence interval and a test against chance, Krippendorff's alpha in nominal and — when the labels turn out to be ordered — ordinal form, a per-category kappa showing exactly which categories the panel cannot pin down, a per-rater agreement with the consensus plus a leave-one-rater-out kappa that names the outlier, the item-level distribution of how strong the consensus actually was, and an explicit account of what each coefficient did with any missing ratings instead of quietly dropping them.

Data
N = 200 observations
Assumptions
  • Each item is judged once by each rater, from the same list of categories
  • The judgements are independent — no rater saw another's answer before deciding
  • Items are independent of each other; the same item is not counted twice
  • Fleiss' kappa additionally assumes every item carries the same number of raters; where that fails the analysis says so and reports Krippendorff's alpha alongside
Limitations
  • Agreement is not correctness — a panel can agree unanimously and be unanimously wrong; only a gold standard can settle that
  • Kappa is depressed when one category dominates or when raters use the categories at different rates; this is the kappa paradox, and the analysis reports it rather than hiding it
  • The confidence intervals come from resampling items and are an approximation; with few items they are wide, and the analysis reports the item count next to the verdict
  • Fleiss' kappa is computed on items every rater judged; when ratings are missing that is a different item set from the one Krippendorff's alpha uses, and both are shown rather than blended
Software & Citation
MCP Analytics · mcpanalytics.ai
Code Appendix

Analysis Code

Complete R source code for this analysis

Multi-Rater Agreement — Fleiss Kappa

Three or more raters, one categorical judgement per item: how much do they agree, how much of that is more than chance would have produced, and which rater is pulling the group apart? The analysis computes Fleiss' kappa with a standard error and confidence interval, Krippendorff's alpha (nominal, and ordinal when the categories turn out to be ordered), a per-category kappa showing which categories the raters fight over, a per-rater agreement with the consensus plus a leave-one-rater-out kappa that names the outlier, and the item-level distribution of how strong the consensus actually was.

Why This Method?

Raw agreement across a panel is not interpretable on its own: a panel that answers "Yes" to almost everything will agree constantly without reading a single item. Fleiss' kappa subtracts the agreement chance alone would deliver given how often each category is used overall. Krippendorff's alpha answers the same question through a different route and, unlike Fleiss', keeps working when not every rater judged every item.

What This Analysis Covers

  • Raw pairwise agreement, chance agreement, and Fleiss' kappa with SE + CI
  • Krippendorff's alpha, nominal and (when the labels are ordered) ordinal
  • Explicit accounting for missing ratings — what each statistic used
  • Per-category kappa: which categories the panel cannot pin down
  • Per-rater agreement with the consensus and leave-one-rater-out kappa
  • The item-level consensus distribution, and the kappa paradox when it fires

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {rater_1 .. rater_N}. Both real-world layouts are accepted and the shape is detected from the data, never assumed. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

Rule 1 — the labels are numbers (1, 2, 3 / 0-10 scales).

num <- suppressWarnings(as.numeric(clean))
  if (!anyNA(num) && length(unique(num)) == length(num)) {
    return(list(ordered = TRUE, order = levs[order(num)],
                rule = "the labels are numbers, so they were ordered by value"))
  }

Rule 2 — the labels start with a number ("1 - Poor", "2 = Fair").

lead <- suppressWarnings(as.numeric(sub("^\\s*([-+]?[0-9]+(\\.[0-9]+)?)\\s*[-=.):|].*$",
                                          "\\1", clean)))
  has_lead <- grepl("^\\s*[-+]?[0-9]+(\\.[0-9]+)?\\s*[-=.):|]", clean)
  if (all(has_lead) && !anyNA(lead) && length(unique(lead)) == length(lead)) {
    return(list(ordered = TRUE, order = levs[order(lead)],
                rule = "every label starts with a number, so they were ordered by that number"))
  }

Rule 3 — the labels are all drawn from one recognised ordinal vocabulary.

scales <- list(
    "an agreement scale" = c("strongly disagree", "disagree", "somewhat disagree",
                             "slightly disagree", "neutral", "neither agree nor disagree",
                             "slightly agree", "somewhat agree", "agree", "strongly agree"),
    "a quality scale" = c("very poor", "poor", "below average", "fair", "average",
                          "satisfactory", "good", "very good", "excellent", "outstanding"),
    "a frequency scale" = c("never", "rarely", "seldom", "sometimes", "occasionally",
                            "often", "frequently", "usually", "always"),
    "a severity scale" = c("none", "minimal", "mild", "moderate", "severe",
                           "very severe", "extreme", "critical"),
    "a magnitude scale" = c("very low", "low", "medium", "moderate", "high", "very high"),
    "a satisfaction scale" = c("very dissatisfied", "dissatisfied", "neutral",
                               "satisfied", "very satisfied"),
    "a likelihood scale" = c("very unlikely", "unlikely", "possible", "likely",
                             "very likely", "certain"),
    "a decision scale" = c("reject", "major revision", "revise", "minor revision",
                           "accept with changes", "accept"),
    "a priority scale" = c("trivial", "minor", "moderate", "major", "critical", "blocker")
  )
  low <- tolower(clean)
  if (length(levs) >= 3 && !any(duplicated(low))) {
    for (nm in names(scales)) {
      sc <- scales[[nm]]
      if (all(low %in% sc)) {
        return(list(ordered = TRUE, order = levs[order(match(low, sc))],
                    rule = paste0("the labels are all points on ", nm,
                                  ", so they were ordered along it")))
      }
    }
  }

  list(ordered = FALSE, order = levs,
       rule = "no numeric values, numeric prefixes, or recognised ordinal wording were found in the labels")
}

Step 1: Collect the mapped judgement columns, in mapped order

rc <- grep("^rater_[0-9]+$", names(df), value = TRUE)
  rc <- rc[order(as.numeric(sub("^rater_", "", rc)))]
  ch <- humanize_semantic(rc, col_map)

  if (length(rc) < 3) {
    stop(sprintf("Only %d judgement column(s) were mapped(%s). This analysis needs three or more raters. For exactly two raters the right tool is Cohen&#x27;s kappa, not Fleiss' — or, if your file has one row per rating, map the three columns holding the item, the rater and the judgement.",
                 length(rc),
                 if (length(ch) > 0) paste(ch, collapse = ", ") else "none"))
  }

  cols <- lapply(rc, function(cn) as_label(df[[cn]]))
  names(cols) <- rc
  d_distinct <- vapply(cols, function(v) length(unique(v[v != ""])), integer(1))

Step 2: Detect the input SHAPE from the data, not from the mapping.

Wide means one column per rater and one row per item. Long means one row per rating, with an item column, a rater column and a judgement column. A long file is unmistakable in the counts: the item column carries many distinct values, each repeating once per rater, while the other two carry only a handful.

shape <- "wide"
  shape_rule <- ""
  li <- ri <- ji <- NA_integer_
  if (length(rc) == 3) {
    ii <- which.max(d_distinct)
    others <- setdiff(seq_len(3), ii)
    avg_rep <- if (d_distinct[ii] > 0) initial_rows / d_distinct[ii] else 0
    if (d_distinct[ii] >= 10 && avg_rep >= 2.5 && min(d_distinct[others]) >= 2 &&
        d_distinct[ii] > 3 * max(d_distinct[others])) {
      item_v <- cols[[ii]]

Which of the remaining two is the RATER? The one for which each item appears at most once — a rater judges an item once, whereas several raters give an item the same judgement all the time.

dupr <- vapply(others, function(o)
        sum(duplicated(paste0(item_v, "\r", cols[[o]]))), numeric(1))
      ri <- others[which.min(dupr)]
      ji <- setdiff(others, ri)
      li <- ii
      shape <- "long"
      shape_rule <- sprintf(
        "%s holds %d distinct values each repeating about %s times, while %s and %s hold only %d and %d distinct values — that is one row per rating, not one column per rater",
        ch[li], d_distinct[li], r2(avg_rep), ch[ri], ch[ji],
        d_distinct[ri], d_distinct[ji])
    }
  }

Step 3: Flatten either shape into one table of (item, rater, label)

n_dup_ratings <- 0L
  if (shape == "long") {
    item_v <- cols[[li]]; rater_v <- cols[[ri]]; lab_v <- cols[[ji]]
    ok <- item_v != "" & rater_v != ""
    n_blank_rows <- sum(!ok)
    item_v <- item_v[ok]; rater_v <- rater_v[ok]; lab_v <- lab_v[ok]

The same rater judging the same item twice is a data error, not a second opinion; the first judgement is kept and the count is reported.

key <- paste0(item_v, "\r", rater_v)
    dupd <- duplicated(key)
    n_dup_ratings <- sum(dupd)
    item_v <- item_v[!dupd]; rater_v <- rater_v[!dupd]; lab_v <- lab_v[!dupd]
    if (length(unique(rater_v)) > 20) {
      stop(sprintf("%s holds %d distinct raters. Above 20 raters this looks like an identifier column rather than a panel; map the column naming who gave each judgement.",
                   ch[ri], length(unique(rater_v))))
    }
    row_unit <- "rating"
    shape_rule <- paste0("the three mapped columns were read as a long file because ", shape_rule)
  } else {

Wide: a column holding a distinct value on nearly every row is an identifier or free text, not a category, and is refused by name. Both an absolute cap and a proportional test are needed — a key column in a short file has few distinct values in absolute terms but repeats nothing.

n_filled <- vapply(cols, function(v) sum(v != ""), integer(1))
    bad <- which(d_distinct > 25 | (d_distinct >= 10 & d_distinct > 0.5 * n_filled))
    if (length(bad) > 0) {
      b <- bad[1]
      why <- if (d_distinct[b] > 25)
        "which is too many to be a set of categories"
      else
        sprintf("across %d judgement(s) — nearly one per row, so almost nothing repeats",
                n_filled[b])
      stop(sprintf("%s holds %d distinct values, %s. That makes it an identifier or free text rather than a rater&#x27;s category choice. Map one column per rater, each holding that rater's category. If your file instead has one row per rating, map exactly three columns: the item, the rater, and the judgement.",
                   ch[b], d_distinct[b], why))
    }
    n_blank_rows <- 0L
    item_v <- rep(as.character(seq_len(initial_rows)), times = length(rc))
    rater_v <- rep(ch, each = initial_rows)
    lab_v <- unlist(cols, use.names = FALSE)
    keep <- lab_v != ""
    item_v <- item_v[keep]; rater_v <- rater_v[keep]; lab_v <- lab_v[keep]
    row_unit <- "item"
    shape_rule <- sprintf(
      "the %d mapped columns were read as one column per rater, because each holds only a handful of repeated labels(%s distinct) rather than one value per row",
      length(rc), paste(d_distinct, collapse = ", "))
  }

Step 4: Merge labels that differ only in capitalisation, and say so.

by_case <- split(lab_v, tolower(lab_v))
  canon <- list(); n_case_merged <- 0L; case_example <- NULL
  for (keyc in names(by_case)) {
    spellings <- by_case[[keyc]]
    uq <- unique(spellings)
    tab <- sort(table(spellings), decreasing = TRUE)
    winner <- names(tab)[1]
    canon[[keyc]] <- winner
    if (length(uq) > 1) {
      n_case_merged <- n_case_merged + sum(spellings != winner)
      if (is.null(case_example)) case_example <- c(setdiff(uq, winner)[1], winner)
    }
  }
  lab_v <- unlist(canon[tolower(lab_v)], use.names = FALSE)

Step 5: Refuse free text; lump a long tail of rare labels into "Other".

freq <- sort(table(lab_v), decreasing = TRUE)
  n_lumped <- 0L
  if (length(freq) > 25) {
    stop(sprintf("The judgement columns(%s) together hold %d distinct values — that looks like free text rather than a set of categories. Map the columns holding each rater&#x27;s category choice.",
                 paste(ch, collapse = ", "), length(freq)))
  }
  if (length(freq) > 12) {
    keep_lab <- names(freq)[1:11]
    n_lumped <- length(freq) - 11L
    lab_v[!(lab_v %in% keep_lab)] <- "Other"
  }

  rater_names <- sort(unique(rater_v))
  n_raters <- length(rater_names)
  if (n_raters < 3) {
    stop(sprintf("Only %d rater(s) were found in %s. Fleiss&#x27; kappa needs three or more; with exactly two raters use Cohen's kappa instead.",
                 n_raters, if (shape == "long") ch[ri] else paste(ch, collapse = ", ")))
  }

  levs_all <- sort(unique(lab_v))
  if (length(levs_all) < 2) {
    stop(sprintf("Every judgement in %s is \"%s\", so there is only one category and agreement beyond chance is undefined — nothing varies for kappa to explain.",
                 paste(ch, collapse = ", "), levs_all[1]))
  }

Step 6: Decide the category order (detected, not assumed)

det <- detect_category_order(levs_all)
  if (det$ordered) {
    levs <- det$order
  } else {
    tot <- vapply(levs_all, function(l) sum(lab_v == l), numeric(1))
    levs <- levs_all[order(-tot, levs_all)]
  }
  k <- length(levs)

Step 7: Build the item-by-category count matrix

items <- unique(item_v)
  Nmat <- matrix(0, nrow = length(items), ncol = k,
                 dimnames = list(NULL, levs))
  im <- match(item_v, items)
  cm <- match(lab_v, levs)
  tabm <- table(factor(im, levels = seq_along(items)),
                factor(cm, levels = seq_len(k)))
  Nmat[] <- as.numeric(tabm)
  m_i <- rowSums(Nmat)

  n_ratings <- sum(m_i)
  all_idx <- which(m_i >= 2)
  n_items_dropped <- length(items) - length(all_idx)
  complete_idx <- which(m_i == n_raters)
  n_items_partial <- sum(m_i >= 2 & m_i < n_raters)
  n_missing <- n_raters * length(items) - n_ratings

  if (length(all_idx) < 20) {
    stop(sprintf("Only %d item(s) carry judgements from at least two of %s. Multi-rater agreement needs at least 20 such items before the numbers mean anything.",
                 length(all_idx), paste(ch, collapse = ", ")))
  }

Step 8: Fleiss' kappa. It is a complete-cases statistic: the classical

coefficient assumes every item was judged by the same number of raters. When ratings are missing the module reports BOTH the complete-case value (the one other software produces) and a generalized value over every item with two or more ratings — and says which items each one used.

fleiss_calc <- function(ii) {
    m <- m_i[ii]
    sq <- rowSums(Nmat[ii, , drop = FALSE]^2)
    P_i <- (sq - m) / (m * (m - 1))
    P_bar <- mean(P_i)
    tot <- sum(m)
    p <- colSums(Nmat[ii, , drop = FALSE]) / tot
    P_e <- sum(p^2)
    kap <- if ((1 - P_e) > 1e-12) (P_bar - P_e) / (1 - P_e) else NA_real_
    list(P_bar = P_bar, P_e = P_e, kappa = kap, p = p, tot = tot,
         n = length(ii), m_const = if (length(unique(m)) == 1) m[1] else NA_real_)
  }

  basis <- if (length(complete_idx) >= 20) "complete" else "all"
  basis_idx <- if (basis == "complete") complete_idx else all_idx
  fb <- fleiss_calc(basis_idx)
  fa <- fleiss_calc(all_idx)
  p_bar <- fb$P_bar; p_e <- fb$P_e; kappa <- fb$kappa
  p_j <- fb$p
  kappa_all <- fa$kappa
  n_basis <- fb$n

Fleiss' null-hypothesis standard error (Fleiss 1971), which exists only when every item on the basis carries the same number of raters. It is the variance of kappa when kappa is truly zero, so it belongs to the significance test and NOT to the confidence interval.

se_null <- NA_real_; z_stat <- NA_real_; p_value <- NA_real_
  if (!is.na(fb$m_const) && fb$m_const >= 2) {
    R <- fb$m_const
    S2 <- sum(p_j^2); S3 <- sum(p_j^3)
    inner <- S2 - (2 * R - 3) * S2^2 + 2 * (R - 2) * S3
    if (is.finite(inner) && inner > 0 && (1 - S2) > 1e-12) {
      se_null <- sqrt(2 / (n_basis * R * (R - 1)) * inner) / (1 - S2)
      z_stat <- kappa / se_null
      p_value <- 2 * pnorm(-abs(z_stat))
    }
  }

Step 9: Krippendorff's alpha, via the coincidence matrix. Alpha is the

statistic that tolerates missing ratings, so it uses EVERY item carrying at least two judgements — including the ones Fleiss had to set aside.

Ocon <- t(apply(Nmat, 1, function(row) {
    m <- sum(row)
    if (m < 2) return(rep(0, k * k))
    as.vector((outer(row, row) - diag(row, nrow = k)) / (m - 1))
  }))
  if (k == 1) Ocon <- matrix(Ocon, ncol = 1)

  LO <- outer(seq_len(k), seq_len(k), pmin)
  HI <- outer(seq_len(k), seq_len(k), pmax)
  ordinal_ok <- det$ordered && k >= 3

  alpha_calc <- function(ii) {
    o <- matrix(colSums(Ocon[ii, , drop = FALSE]), nrow = k, ncol = k)
    nc <- rowSums(o)
    ntot <- sum(nc)
    a_nom <- NA_real_; a_ord <- NA_real_
    De_n <- ntot^2 - sum(nc^2)
    if (is.finite(De_n) && De_n > 1e-12 && ntot > 1) {
      Do_n <- ntot - sum(diag(o))
      a_nom <- 1 - (ntot - 1) * Do_n / De_n
    }
    if (ordinal_ok && ntot > 1) {
      cum0 <- c(0, cumsum(nc))
      S <- matrix(cum0[HI + 1] - cum0[LO], nrow = k, ncol = k)
      Dm <- (S - outer(nc, nc, function(x, y) (x + y) / 2))^2
      De_o <- sum(outer(nc, nc) * Dm)
      if (is.finite(De_o) && De_o > 1e-12) {
        a_ord <- 1 - (ntot - 1) * sum(o * Dm) / De_o
      }
    }
    list(nominal = a_nom, ordinal = a_ord)
  }

  aa <- alpha_calc(all_idx)
  alpha_nom <- aa$nominal; alpha_ord <- aa$ordinal
  alpha_nom_basis <- alpha_calc(basis_idx)$nominal

Step 10: Confidence intervals by resampling ITEMS. The closed-form

variance above is a null-hypothesis quantity and is wrong for an interval around a non-zero estimate; alpha has no simple closed form at all. A seeded nonparametric bootstrap over items answers both, and treats items as the sampling unit, which is what they are.

B <- 500L
  set.seed(42)
  nb <- length(basis_idx); na_ <- length(all_idx)
  bk <- numeric(B); ban <- numeric(B); bao <- numeric(B)
  for (b in seq_len(B)) {
    ib <- basis_idx[sample.int(nb, nb, replace = TRUE)]
    bk[b] <- fleiss_calc(ib)$kappa
    ia <- all_idx[sample.int(na_, na_, replace = TRUE)]
    ab <- alpha_calc(ia)
    ban[b] <- ab$nominal; bao[b] <- ab$ordinal
  }
  qci <- function(v) {
    v <- v[is.finite(v)]
    if (length(v) < 50) return(c(NA_real_, NA_real_, NA_real_))
    c(as.numeric(quantile(v, 0.025, names = FALSE)),
      as.numeric(quantile(v, 0.975, names = FALSE)), sd(v))
  }
  qk <- qci(bk); qan <- qci(ban); qao <- qci(bao)
  ci_low <- qk[1]; ci_high <- qk[2]; se_boot <- qk[3]
  alpha_nom_lo <- qan[1]; alpha_nom_hi <- qan[2]
  alpha_ord_lo <- qao[1]; alpha_ord_hi <- qao[2]

  band <- landis_koch(kappa)

Step 11: Per-category kappa — which categories the panel fights over.

Fleiss' category-specific coefficient, generalized to a variable number of raters per item.

mb <- m_i[basis_idx]
  Nb <- Nmat[basis_idx, , drop = FALSE]
  tot_b <- sum(mb)
  cat_kappa <- vapply(seq_len(k), function(j) {
    pj <- sum(Nb[, j]) / tot_b
    den <- tot_b * pj * (1 - pj)
    if (!is.finite(den) || den <= 1e-12) return(NA_real_)
    obs <- sum(Nb[, j] * (mb - Nb[, j]) / (mb - 1))
    1 - obs / den
  }, numeric(1))
  category_df <- data.frame(
    category = levs,
    category_kappa = round(cat_kappa, 4),
    share_pct = round(100 * colSums(Nb) / tot_b, 2),
    n_uses = as.numeric(colSums(Nb)),
    stringsAsFactors = FALSE
  )
  ok_cat <- which(!is.na(cat_kappa))
  worst_cat <- if (length(ok_cat) > 0) levs[ok_cat[which.min(cat_kappa[ok_cat])]] else NA_character_
  worst_val <- if (length(ok_cat) > 0) min(cat_kappa[ok_cat]) else NA_real_
  best_cat <- if (length(ok_cat) > 0) levs[ok_cat[which.max(cat_kappa[ok_cat])]] else NA_character_
  best_val <- if (length(ok_cat) > 0) max(cat_kappa[ok_cat]) else NA_real_

With exactly two categories every disagreement is the same disagreement, so both category-specific kappas equal each other and the overall coefficient. Naming a "worst" category there would be an artefact of tie-breaking.

cat_degenerate <- k == 2 ||
    (length(ok_cat) > 1 && (best_val - worst_val) < 1e-9)

Step 12: Per-rater agreement with the rest of the panel, and the

leave-one-rater-out kappa. The second is the actionable one: it says what the panel's agreement would be if this rater were not in it.

basis_items <- items[basis_idx]
  in_basis <- item_v %in% basis_items
  bi_item <- item_v[in_basis]; bi_rater <- rater_v[in_basis]; bi_lab <- lab_v[in_basis]
  bi_row <- match(bi_item, basis_items)
  bi_col <- match(bi_lab, levs)
  m_of_row <- mb[bi_row]
  n_same <- Nb[cbind(bi_row, bi_col)]

  rater_obs <- numeric(n_raters); rater_kap <- numeric(n_raters)
  rater_wo <- numeric(n_raters); rater_n <- numeric(n_raters)
  for (r in seq_len(n_raters)) {
    sel <- bi_rater == rater_names[r]
    den <- sum(m_of_row[sel] - 1)
    rater_n[r] <- sum(sel)
    rater_obs[r] <- if (den > 0) sum(n_same[sel] - 1) / den else NA_real_
    rater_kap[r] <- if (is.na(rater_obs[r]) || (1 - p_e) <= 1e-12) NA_real_
                    else (rater_obs[r] - p_e) / (1 - p_e)

Leave-one-rater-out: strip this rater's ratings and recompute Fleiss on the items that still carry two or more judgements.

Nw <- Nb
    idxw <- which(sel)
    if (length(idxw) > 0) {
      dec <- table(factor(bi_row[idxw], levels = seq_len(nrow(Nb))),
                   factor(bi_col[idxw], levels = seq_len(k)))
      Nw <- Nw - matrix(as.numeric(dec), nrow = nrow(Nb), ncol = k)
    }
    mw <- rowSums(Nw)
    keepw <- which(mw >= 2)
    if (length(keepw) >= 5) {
      sqw <- rowSums(Nw[keepw, , drop = FALSE]^2)
      Pw <- (sqw - mw[keepw]) / (mw[keepw] * (mw[keepw] - 1))
      totw <- sum(mw[keepw])
      pw <- colSums(Nw[keepw, , drop = FALSE]) / totw
      Pew <- sum(pw^2)
      rater_wo[r] <- if ((1 - Pew) > 1e-12) (mean(Pw) - Pew) / (1 - Pew) else NA_real_
    } else {
      rater_wo[r] <- NA_real_
    }
  }

  rater_df <- data.frame(
    rater = rater_names,
    agreement_pct = round(100 * rater_obs, 2),
    rater_kappa = round(rater_kap, 4),
    kappa_without_rater = round(rater_wo, 4),
    items_rated = rater_n,
    stringsAsFactors = FALSE
  )
  ok_r <- which(!is.na(rater_kap))
  outlier_rater <- if (length(ok_r) > 0) rater_names[ok_r[which.min(rater_kap[ok_r])]] else NA_character_
  outlier_kappa <- if (length(ok_r) > 0) min(rater_kap[ok_r]) else NA_real_
  best_rater <- if (length(ok_r) > 0) rater_names[ok_r[which.max(rater_kap[ok_r])]] else NA_character_
  best_rater_kappa <- if (length(ok_r) > 0) max(rater_kap[ok_r]) else NA_real_
  outlier_gain <- if (!is.na(outlier_rater) && !is.na(rater_wo[match(outlier_rater, rater_names)]))
    rater_wo[match(outlier_rater, rater_names)] - kappa else NA_real_
  rater_spread <- if (length(ok_r) > 1) max(rater_kap[ok_r]) - min(rater_kap[ok_r]) else NA_real_

Step 13: Item-level consensus. Banded by each item's MODAL share —

the largest number of raters landing on one category, over the raters who judged it. Only the maximum COUNT is used, never its position, so an evenly split item needs no arbitrary winner.

modal_share <- apply(Nb, 1, max) / mb
  band_labels <- c("No majority — the raters split",
                   "Simple majority — more than half",
                   "Strong majority — more than two thirds",
                   "Unanimous — every rater agreed")
  band_idx <- ifelse(modal_share >= 1, 4L,
              ifelse(modal_share > 2 / 3, 3L,
              ifelse(modal_share > 0.5, 2L, 1L)))
  item_df <- data.frame(
    consensus_band = band_labels,
    n_items = as.numeric(tabulate(band_idx, nbins = 4)),
    stringsAsFactors = FALSE
  )
  item_df$share_pct <- round(100 * item_df$n_items / n_basis, 2)
  n_unanimous <- item_df$n_items[4]
  n_split <- item_df$n_items[1]

Step 14: Per-rater marginals — how each rater uses the scale. These

are the habits behind both the chance floor and any outlier.

marg <- matrix(0, nrow = n_raters, ncol = k, dimnames = list(rater_names, levs))
  tabr <- table(factor(bi_rater, levels = rater_names),
                factor(bi_lab, levels = levs))
  marg[] <- as.numeric(tabr)
  marg_share <- marg / pmax(rowSums(marg), 1)
  overall_share <- colSums(marg) / sum(marg)
  divergence <- rowSums(abs(sweep(marg_share, 2, overall_share)))
  show_r <- rater_names
  n_raters_hidden <- 0L
  if (n_raters > 8) {
    show_r <- rater_names[order(-divergence)][1:8]
    n_raters_hidden <- n_raters - 8L
  }
  marginal_df <- data.frame(
    category = rep(levs, times = length(show_r)),
    rater = rep(show_r, each = k),
    share_pct = round(100 * as.vector(t(marg_share[show_r, , drop = FALSE])), 2),
    stringsAsFactors = FALSE
  )

  prev_i <- which.max(p_j)
  max_prev <- p_j[prev_i]
  prev_cat <- levs[prev_i]
  paradox <- isTRUE(p_bar >= 0.70 && !is.na(kappa) && kappa < 0.60 && max_prev >= 0.60)

Step 15: The results table

basis_phrase <- if (basis == "complete")
    sprintf("the %s item(s) every one of the %d raters judged", fmt_n(n_basis), n_raters)
  else
    sprintf("all %s item(s) carrying at least two judgements", fmt_n(n_basis))

  rows <- list(
    list("Raw pairwise agreement", p_bar, NA_real_, NA_real_,
         sprintf("Across every pair of raters who judged the same item, %s of those pairs chose the same category.",
                 pct1(p_bar))),
    list("Agreement expected by chance", p_e, NA_real_, NA_real_,
         sprintf("What a panel using the categories this often would have agreed on by luck alone(%s).",
                 pct1(p_e))),
    list("Fleiss&#x27; kappa", kappa, ci_low, ci_high,
         sprintf("%s of the agreement left available above chance was achieved, on %s — %s on the Landis & Koch convention.",
                 pct1(kappa), basis_phrase, band))
  )
  if (n_missing > 0 && basis == "complete") {
    rows[[length(rows) + 1]] <- list(
      "Fleiss&#x27; kappa (all items, generalized)", kappa_all, NA_real_, NA_real_,
      sprintf("The same coefficient generalized to a variable number of raters per item, so it uses all %s items with two or more judgements instead of only the %s complete ones.",
              fmt_n(length(all_idx)), fmt_n(n_basis)))
  }
  rows[[length(rows) + 1]] <- list(
    "Krippendorff&#x27;s alpha (nominal)", alpha_nom, alpha_nom_lo, alpha_nom_hi,
    sprintf("The disagreement-based coefficient, computed on all %s items with two or more judgements — it does not require every rater to have judged every item.",
            fmt_n(length(all_idx))))
  if (!is.na(alpha_ord)) {
    rows[[length(rows) + 1]] <- list(
      "Krippendorff&#x27;s alpha (ordinal)", alpha_ord, alpha_ord_lo, alpha_ord_hi,
      "The same coefficient with an ordinal distance between categories, so a near-miss between neighbouring levels counts as a smaller disagreement than a jump across the scale.")
  }
  if (n_missing > 0 && basis == "complete" && !is.na(alpha_nom_basis)) {
    rows[[length(rows) + 1]] <- list(
      "Krippendorff&#x27;s alpha (complete items only)", alpha_nom_basis, NA_real_, NA_real_,
      "Alpha restricted to the same complete items Fleiss&#x27; kappa used, so the two headline coefficients can be compared on identical data.")
  }

  agreement_df <- data.frame(
    statistic = vapply(rows, function(r) r[[1]], character(1)),
    estimate = round(vapply(rows, function(r) as.numeric(r[[2]]), numeric(1)), 4),
    ci_low = round(vapply(rows, function(r) as.numeric(r[[3]]), numeric(1)), 4),
    ci_high = round(vapply(rows, function(r) as.numeric(r[[4]]), numeric(1)), 4),
    interpretation = vapply(rows, function(r) r[[5]], character(1)),
    stringsAsFactors = FALSE
  )

Step 16: Methods disclosure

se_line <- if (is.na(se_null)) {
    sprintf("The classical null-hypothesis standard error is not available here, because the items on the analysis basis do not all carry the same number of raters. The significance test is therefore not reported; read the bootstrap interval instead, which does not need that assumption.")
  } else {
    sprintf("Fleiss&#x27; null-hypothesis standard error is %s, giving z = %s and %s against the hypothesis that agreement is no better than chance. That variance describes kappa when kappa is truly zero, so it belongs to this test and NOT to the confidence interval.",
            r3(se_null), r2(z_stat), fmt_pp(p_value))
  }
  methods_df <- data.frame(
    item = c(
      "Design",
      "Input shape",
      "Raw pairwise agreement",
      "Chance agreement",
      "Fleiss&#x27; kappa",
      "Significance test",
      "Confidence intervals",
      "Missing ratings",
      "Krippendorff&#x27;s alpha",
      "Category order",
      "Per-category kappa",
      "Per-rater figures",
      "Benchmark labels",
      "What agreement is not"
    ),
    detail = c(
      sprintf("%s item(s) judged by %d rater(s) across %d categories, %s judgement(s) in total.",
              fmt_n(length(items)), n_raters, k, fmt_n(n_ratings)),
      sprintf("Detected from the data, not from the mapping: %s.", shape_rule),
      sprintf("Mean over items of the share of rater PAIRS on that item choosing the same category: %s.",
              pct1(p_bar)),
      sprintf("Sum of the squared overall category shares(%s): %s.",
              paste(sprintf("%s %s", levs, vapply(p_j, pct1, character(1))), collapse = ", "),
              pct1(p_e)),
      sprintf("(observed - expected) / (1 - expected) = (%s - %s) / (1 - %s) = %s, computed on %s.",
              r3(p_bar), r3(p_e), r3(p_e), r3(kappa), basis_phrase),
      se_line,
      sprintf("A seeded nonparametric bootstrap resampling ITEMS with replacement(%d resamples, fixed seed), taking the 2.5th and 97.5th percentiles. The bootstrap standard error of kappa is %s. This is used rather than a closed form because the null variance above is the wrong quantity for an interval around a non-zero estimate, and alpha has no simple closed-form variance.",
              B, r3(se_boot)),
      if (n_missing > 0)
        sprintf("%s of the %s possible rater-by-item judgements are absent. Fleiss&#x27; kappa is a complete-cases statistic and used %s; Krippendorff's alpha used all %s items with at least two judgements; %s item(s) carrying a single judgement contribute to neither, because agreement needs at least two opinions.",
                fmt_n(n_missing), fmt_n(n_raters * length(items)), basis_phrase,
                fmt_n(length(all_idx)), fmt_n(n_items_dropped))
      else
        "No ratings are missing: every rater judged every item, so Fleiss&#x27; kappa and Krippendorff's alpha are computed on exactly the same data.",
      sprintf("Built from the coincidence matrix: alpha = 1 - observed disagreement / expected disagreement, with(n-1) weighting so it is defined for small samples. Nominal alpha treats every disagreement as equal. %s",
              if (!is.na(alpha_ord))
                "Ordinal alpha weights a disagreement by the distance between the categories along the detected order."
              else
                "Ordinal alpha is not reported for these categories."),
      sprintf("Orderedness was detected, not assumed: %s.", det$rule),
      "Fleiss&#x27; category-specific coefficient, generalized to a variable number of raters per item: one minus the observed splits on that category over the splits expected from its overall share.",
      sprintf("A rater&#x27;s agreement figure is the share of the OTHER raters' judgements on the same items that matched this rater's own, chance-corrected against the same %s floor. The leave-one-out column recomputes Fleiss' kappa for the panel with that rater removed.",
              pct1(p_e)),
      "Landis & Koch(1977) labels(slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.",
      "Agreement is not correctness: a panel can agree unanimously and be unanimously wrong. These figures describe this panel on these items and do not generalise to other raters or other items."
    ),
    stringsAsFactors = FALSE
  )

Step 17: Headline metrics + the computed one-paragraph answer

metrics <- list(
    `Items Rated`        = as.numeric(length(items)),
    `Raters`             = as.numeric(n_raters),
    `Categories`         = as.numeric(k),
    `Raw Agreement`      = pct1(p_bar),
    `Fleiss Kappa`       = round(kappa, 3),
    `Fleiss 95% CI`      = paste0(r3(ci_low), " to ", r3(ci_high)),
    `Krippendorff Alpha` = round(alpha_nom, 3),
    `Agreement Strength` = band,
    `Least Aligned Rater` = if (is.na(outlier_rater)) "not identifiable" else outlier_rater,
    `Kappa vs Chance p`  = fmt_p(p_value)
  )

  paradox_sentence <- if (paradox) {
    paste0(
      " Read the two agreement numbers together before quoting either: raw pairwise agreement is high(",
      pct1(p_bar), ") while kappa is only ", r3(kappa),
      ", which is the well-known kappa paradox rather than a contradiction. ",
      pct1(max_prev), " of all judgements fell into the single category \"", prev_cat,
      "\", so a panel with these habits would already have agreed on ", pct1(p_e),
      " of rater pairs by chance; only ", pct1(1 - p_e),
      " of the scale was left for skill to win, and ", pct1(kappa), " of that remainder was won.")
  } else {
    paste0(
      " Chance alone would have produced ", pct1(p_e),
      " pairwise agreement given how often each category is used, leaving ",
      pct1(1 - p_e), " of the scale available above chance, of which ",
      pct1(kappa), " was achieved.")
  }

  missing_sentence <- if (n_missing > 0) {
    paste0(" ", fmt_n(n_missing), " of the ", fmt_n(n_raters * length(items)),
           " possible judgements are missing, which the two coefficients handle differently: Fleiss&#x27; kappa is a complete-cases statistic and used ",
           basis_phrase, ", while Krippendorff&#x27;s alpha used all ", fmt_n(length(all_idx)),
           " items carrying at least two judgements",
           if (n_items_dropped > 0)
             paste0(" and ", fmt_n(n_items_dropped),
                    " item(s) with a single judgement were used by neither")
           else "",
           ". Where ratings are missing, alpha is the coefficient to quote.")
  } else ""

  ordinal_sentence <- if (!is.na(alpha_ord)) {
    paste0(" The categories are ordered(", det$rule,
           "), so ordinal alpha(", r3(alpha_ord),
           ") is also reported; it counts a disagreement between neighbouring levels as smaller than a jump across the scale, which is why it sits above the nominal value of ",
           r3(alpha_nom), ".")
  } else {
    paste0(" Ordinal alpha is not reported because ", det$rule,
           "; treating unordered labels as if some disagreements were milder than others would invent structure the data does not carry.")
  }

  outlier_sentence <- if (!is.na(outlier_rater) && !is.na(outlier_gain)) {
    paste0(" ", outlier_rater, " is the least aligned member of the panel(rater kappa ",
           r3(outlier_kappa), " against ", r3(best_rater_kappa), " for ", best_rater,
           "); with that rater excluded the panel&#x27;s kappa would be ",
           r3(kappa + outlier_gain),
           if (outlier_gain > 0) paste0(", ", r3(outlier_gain), " higher than it is now")
           else paste0(", ", r3(abs(outlier_gain)), " lower than it is now"),
           ".")
  } else ""

  json_output <- list(
    answer = paste0(
      fmt_n(n_raters), " raters judged ", fmt_n(length(items)), " items across ", k,
      " categories, agreeing on ", pct1(p_bar),
      " of rater pairs. Fleiss&#x27; kappa is ", r3(kappa),
      " (95% CI ", r3(ci_low), " to ", r3(ci_high), "), ", band,
      " agreement on the Landis & Koch convention, and ",
      if (!is.na(p_value) && p_value < 0.05)
        paste0("clearly better than chance(", fmt_pp(p_value), ")")
      else if (!is.na(p_value))
        paste0("not distinguishable from chance(", fmt_pp(p_value), ")")
      else
        "could not be tested against chance because the raters did not all judge the same items",
      "; Krippendorff&#x27;s alpha is ", r3(alpha_nom), ".",
      paradox_sentence, missing_sentence, ordinal_sentence, outlier_sentence,
      if (cat_degenerate)
        paste0(" With only ", k,
               " categories every disagreement is the same disagreement, so the category-specific kappas are necessarily equal to each other and to the overall coefficient — there is no per-category story to tell here.")
      else
        paste0(" The panel is least consistent on the category \"", worst_cat,
               "\" (category kappa ", r3(worst_val), ") and most consistent on \"",
               best_cat, "\" (", r3(best_val), ")."),
      " Agreement is not correctness: a panel can agree unanimously and be unanimously wrong."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "agreement_results", "rater_consensus",
        "category_kappa", "item_consensus", "rater_marginals", "methods"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  final_rows <- if (row_unit == "item") length(all_idx) else n_ratings

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = max(0, initial_rows - final_rows),
    row_unit = row_unit, shape = shape, shape_rule = shape_rule,
    n_blank_rows = n_blank_rows, n_dup_ratings = n_dup_ratings,
    n_case_merged = n_case_merged, case_example = case_example, n_lumped = n_lumped,
    col_names = ch, rater_names = rater_names, n_raters = n_raters,
    levs = levs, k_cats = k, n_items = length(items), n_ratings = n_ratings,
    n_missing = n_missing, n_items_partial = n_items_partial,
    n_items_dropped = n_items_dropped, n_all = length(all_idx),
    basis = basis, n_basis = n_basis, n_basis_ratings = tot_b,
    basis_phrase = basis_phrase,
    p_bar = p_bar, p_e = p_e, p_j = p_j, kappa = kappa, kappa_all = kappa_all,
    se_null = se_null, se_boot = se_boot, z_stat = z_stat, p_value = p_value,
    ci_low = ci_low, ci_high = ci_high,
    alpha_nom = alpha_nom, alpha_nom_lo = alpha_nom_lo, alpha_nom_hi = alpha_nom_hi,
    alpha_ord = alpha_ord, alpha_nom_basis = alpha_nom_basis,
    ordered_flag = det$ordered, order_rule = det$rule, band = band,
    paradox = paradox, paradox_sentence = paradox_sentence,
    missing_sentence = missing_sentence, ordinal_sentence = ordinal_sentence,
    outlier_sentence = outlier_sentence,
    max_prev = max_prev, prev_cat = prev_cat,
    outlier_rater = outlier_rater, outlier_kappa = outlier_kappa,
    outlier_gain = outlier_gain, best_rater = best_rater,
    best_rater_kappa = best_rater_kappa, rater_spread = rater_spread,
    n_raters_hidden = n_raters_hidden,
    worst_cat = worst_cat, worst_val = worst_val,
    best_cat = best_cat, best_val = best_val, cat_degenerate = cat_degenerate,
    n_unanimous = n_unanimous, n_split = n_split,
    agreement_df = agreement_df, rater_df = rater_df, category_df = category_df,
    item_df = item_df, marginal_df = marginal_df, methods_df = methods_df,
    metrics = metrics, json_output = json_output, boot_B = B
  )
}
Your data has more stories to tell. Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No Signup Sign Up Free

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing