Standard Agreement Multi
Executive Summary

Executive Summary

Fleiss' kappa 0.670 (substantial) across 5 raters on 9,986 items

Items Rated
9986
Raters
5
Categories
3
Raw Agreement
78.0%
Fleiss Kappa
0.67
Fleiss 95% CI
0.664 to 0.677
Krippendorff Alpha
0.67
Agreement Strength
substantial
Least Aligned Rater
annotator 2
Kappa vs Chance p
< 0.001
5 raters judged 9,986 items and agreed on 78.0% of rater pairs. Fleiss' kappa is 0.670 (95% CI 0.664 to 0.677), which the Landis & Koch convention calls SUBSTANTIAL agreement, and which is clearly better than chance (z = 299.29, p < 0.001). That interval rests on 9,986 item(s). Krippendorff's alpha, which reaches the same question from observed against expected disagreement, gives 0.670. Chance alone would have produced 33.3% pairwise agreement given how often each category is used, leaving 66.7% of the scale available above chance, of which 67.0% was achieved. annotator 2 is the least aligned member of the panel (rater kappa 0.651 against 0.718 for annotator 1); with that rater excluded the panel's kappa would be 0.683, 0.013 higher than it is now. 5,479 of 9,986 items (54.9%) drew a unanimous verdict, and 155 (1.6%) failed to produce a majority at all. One caution that applies whatever the number: these coefficients measure whether the panel makes the same call, not whether the call is right — a panel can agree unanimously and be unanimously wrong, and only a gold standard can tell you which.
What this means

Five raters achieved Fleiss' kappa of 0.67 (95% CI 0.664–0.677) on 9,986 items, substantially better than the 33.3% chance agreement their category frequencies would produce. Krippendorff's alpha confirms 0.67. Annotator 1 is most aligned (rater kappa 0.718); annotator 2 is least aligned (rater kappa 0.651), but removing annotator 2 raises panel kappa only to 0.683, indicating the disagreement is systemic rather than person-driven. Of 9,986 items, 5,479 (54.9%) drew unanimous verdicts; 155 (1.6%) produced no majority. Agreement measures consistency, not correctness—a panel can agree unanimously and be wrong.

Overview

Analysis Overview

Fleiss' kappa and Krippendorff's alpha on 9,986 items judged by 5 raters across 3 categories.

N Items9986
N Raters5
N Categories3
Pairwise Agreement Pct78
Fleiss Kappa0.67
What this means

Five raters judged 9,986 sentence pairs across three categories (entailment, neutral, contradiction). Raw pairwise agreement reached 78.0%, but this inflates when raters lean on the same category by chance alone. Fleiss' kappa (0.67) removes that floor by computing what the panel's category habits would produce randomly (33.3%) and reporting the share of remaining headroom won. Krippendorff's alpha (0.67) asks the same question through disagreement instead and tolerates missing ratings. The two coefficients nearly coincide here because data is complete. With exactly two raters, Cohen's kappa applies; with continuous measurements, an intraclass correlation replaces kappa entirely because kappa discards the scale.

Data Preparation

Data Quality

9,986 items and 49,930 judgements from 9,986 rows loaded.

Initial Rows9986
Items9986
Raters5
Judgements Found49930
Missing Ratings0
Items On Basis9986
Judgements On Basis49930
What this means

The short answer

The input file was in wide form (one row per item, one column per rater) and was correctly detected from the data: each of the 5 rater columns held only 3 distinct labels instead of one value per row. No ratings were missing—all 5 raters judged all 9,986 items, yielding 49,930 judgements. Both Fleiss' kappa and Krippendorff's alpha were computed on identical complete data, so they should (and do) converge.

The detail

9,986 rows loaded, resolving to 9,986 items, 5 raters, and 49,930 judgements across 3 categories. The file was recognised as wide form because the 5 mapped columns showed 3, 3, 3, 3, 3 distinct values respectively—the hallmark of one column per rater. The panel resolved to annotator 1, annotator 2, annotator 3, annotator 4, annotator 5. No ratings are missing; every rater judged every item. The 95% confidence interval around kappa runs 0.664 to 0.677, a width of 0.014, anchored to 9,986 items on the analysis basis. Labels were trimmed of surrounding spaces, and blanks and standard placeholders were treated as no judgement rather than as a category.

What this can't tell you

Complete data eliminates the need to choose between coefficients on the basis of missing patterns. If ratings had been sparse, Krippendorff's alpha would have used every item with two or more judgements while Fleiss' kappa would have dropped incomplete items—a distinction that matters when missingness is not random.

Data Table

Agreement Statistics

Pairwise agreement, chance agreement, Fleiss' kappa and Krippendorff's alpha with bootstrap intervals.

StatisticEstimateCI LowCI HighInterpretation
Raw pairwise agreement0.78Across every pair of raters who judged the same item, 78.0% of those pairs chose the same category.
Agreement expected by chance0.3334What a panel using the categories this often would have agreed on by luck alone (33.3%).
Fleiss' kappa0.670.66350.677467.0% of the agreement left available above chance was achieved, on the 9,986 item(s) every one of the 5 raters judged — substantial on the Landis & Koch convention.
Krippendorff's alpha (nominal)0.670.66250.6778The disagreement-based coefficient, computed on all 9,986 items with two or more judgements — it does not require every rater to have judged every item.
What this means

Pairwise agreement of 78.0% against chance agreement of 33.3% leaves 66.7% headroom, of which kappa captures 67.0%. The 95% bootstrap interval (0.6635 to 0.6774) derives from item resampling, not formula—the classical Fleiss standard error applies only to the null hypothesis, not to intervals around non-zero estimates. Fleiss' kappa and nominal Krippendorff's alpha match within 0.000, which is correct when data are complete and identical; they are two arithmetic routes to one question. Ordinal alpha was not computed because category labels carry no numeric order, numeric prefix, or ordinal wording. If categories are ordered, rename them (e.g., "1 - entailment") and re-run. Whether 0.67 meets your standard depends on the decision it serves.

Visualization

Which Rater Is Out of Step

Each rater's agreement with the rest of the panel, and the panel's kappa without them.

What this means

Annotator 1 leads at rater kappa 0.7178 (81.19% raw agreement with the panel); annotator 2 trails at rater kappa 0.6507 (76.72%). The spread between highest and lowest is 0.067, so no single rater dominates the disagreement. Removing annotator 2 would raise panel kappa only to 0.6829, a gain of 0.013—negligible. Removing annotator 1 would lower it to 0.6382, showing that alignment is distributed across the panel, not concentrated in one outlier. A low rater-level score reflects threshold drift (applying the rubric differently) rather than careless error, and can occur whether that rater or the rest of the panel is correct.

Visualization

Which Categories the Panel Fights Over

Category-specific kappa across 3 categories.

What this means

The short answer

Contradiction was the most reliably coded category (kappa 0.759), neutral the least (kappa 0.568), a gap of 0.191. No single category is carrying the disagreement; the spread is modest enough that improvement must come from the rubric as a whole, not from fixing one category definition.

The detail

Category-specific kappa: entailment 0.6826 (33.97% of judgements, 16,959 uses); neutral 0.5684 (33.21%, 16,581 uses); contradiction 0.7594 (32.83%, 16,390 uses). All three categories appear in roughly equal proportions, so none is a rare label thinly evidenced. The 0.191-point spread between contradiction and neutral is the largest single-category gap. A category that one rater reaches for freely and others barely touch scores low here even when overall kappa looks healthy—this detail a single summary coefficient hides.

What this can't tell you

Category-specific kappa does not isolate whether disagreement stems from the category's definition or from how individual raters apply it. To locate the rubric's weakest point, sample items coded as neutral where raters split, and compare their wording or context to items where neutral was unanimous. The equal distribution of judgements across categories means the chance-agreement floor is not inflated by any single dominant class.

Visualization

How Often the Panel Reached a Verdict

Consensus strength across 9,986 items on the analysis basis.

What this means

The short answer

More than half the items (5,479 of 9,986, or 54.9%) drew unanimous agreement. A further 4,352 (43.6%) produced a majority without unanimity. Only 155 items (1.6%) produced no majority at all—the concrete work list where the rubric is silent or the item is genuinely ambiguous.

The detail

Consensus bands: unanimous (every rater agreed) 5,479 items (54.87%); strong majority (more than two-thirds) 2,845 items (28.49%); simple majority (more than half) 1,507 items (15.09%); no majority (raters split) 155 items (1.55%). Bands are set from each item's largest block of agreeing raters, so an evenly split item needs no arbitrary winner. With 5 raters the possible consensus values are coarse (unanimous, 4-of-5, 3-of-5, tie); empty bands in a visualization mean arithmetic simply cannot land there, not that no item did.

What this can't tell you

Concentration of disagreement in a small minority of items does not mean the rubric is clear; it may mean most items are unambiguous while a hard core are genuinely borderline. Re-reading a sample of the 155 no-majority items usually explains the headline kappa faster than the headline number itself does. Consider exporting those 155 items separately to identify patterns in their wording, context, or source.

Visualization

How Each Rater Uses the Scale

Category shares per rater, side by side.

What this means

All five raters distribute their judgements across the three categories with no dramatic skew. Entailment ranges from 33.35% (annotator 1) to 34.34% (annotator 5); neutral from 32.71% (annotator 4) to 33.73% (annotator 2); contradiction from 32.30% (annotator 2) to 33.35% (annotator 1). The even spread is the mechanical reason the chance-agreement floor stays at 33.3%, allowing the 78.0% raw agreement to translate into a kappa of 0.67 rather than being inflated by category bias. A rater whose bars deviate visibly from the group applies a different threshold, not random noise, and depresses agreement even when individual judgements are defensible.

Data Table

Methods & Disclosure

Every formula behind the numbers, and what they cannot decide.

ItemDetail
Design9,986 item(s) judged by 5 rater(s) across 3 categories, 49,930 judgement(s) in total.
Input shapeDetected from the data, not from the mapping: the 5 mapped columns were read as one column per rater, because each holds only a handful of repeated labels (3, 3, 3, 3, 3 distinct) rather than one value per row.
Raw pairwise agreementMean over items of the share of rater PAIRS on that item choosing the same category: 78.0%.
Chance agreementSum of the squared overall category shares (entailment 34.0%, neutral 33.2%, contradiction 32.8%): 33.3%.
Fleiss' kappa(observed - expected) / (1 - expected) = (0.780 - 0.333) / (1 - 0.333) = 0.670, computed on the 9,986 item(s) every one of the 5 raters judged.
Significance testFleiss' null-hypothesis standard error is 0.002, giving z = 299.29 and p < 0.001 against the hypothesis that agreement is no better than chance. That variance describes kappa when kappa is truly zero, so it belongs to this test and NOT to the confidence interval.
Confidence intervalsA seeded nonparametric bootstrap resampling ITEMS with replacement (500 resamples, fixed seed), taking the 2.5th and 97.5th percentiles. The bootstrap standard error of kappa is 0.004. This is used rather than a closed form because the null variance above is the wrong quantity for an interval around a non-zero estimate, and alpha has no simple closed-form variance.
Missing ratingsNo ratings are missing: every rater judged every item, so Fleiss' kappa and Krippendorff's alpha are computed on exactly the same data.
Krippendorff's alphaBuilt from the coincidence matrix: alpha = 1 - observed disagreement / expected disagreement, with (n-1) weighting so it is defined for small samples. Nominal alpha treats every disagreement as equal. Ordinal alpha is not reported for these categories.
Category orderOrderedness was detected, not assumed: no numeric values, numeric prefixes, or recognised ordinal wording were found in the labels.
Per-category kappaFleiss' category-specific coefficient, generalized to a variable number of raters per item: one minus the observed splits on that category over the splits expected from its overall share.
Per-rater figuresA rater's agreement figure is the share of the OTHER raters' judgements on the same items that matched this rater's own, chance-corrected against the same 33.3% floor. The leave-one-out column recomputes Fleiss' kappa for the panel with that rater removed.
Benchmark labelsLandis & Koch (1977) labels (slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.
What agreement is notAgreement is not correctness: a panel can agree unanimously and be unanimously wrong. These figures describe this panel on these items and do not generalise to other raters or other items.
What this means

The short answer

All figures except the confidence intervals are closed-form arithmetic on the 9,986-by-3 matrix of rater counts per item per category. The intervals come from resampling items 500 times with a fixed seed, producing reproducible but approximate bounds that widen when item counts are small. Two judgement calls were made: input shape (detected as wide form from the data) and category order (detected as nominal—no numeric values, prefixes, or ordinal wording found in the labels).

The detail

Design: 9,986 items, 5 raters, 3 categories, 49,930 judgements. Raw pairwise agreement: mean share of rater pairs per item choosing the same category = 78.0%. Chance agreement: sum of squared overall category shares (0.340² + 0.332² + 0.328²) = 0.333. Fleiss' kappa: (0.780 − 0.333) / (1 − 0.333) = 0.670, on 9,986 items where all 5 raters judged. Significance: null standard error 0.002, z = 299.29, p < 0.001. Confidence intervals: nonparametric bootstrap resampling items with replacement (500 resamples, seeded), taking 2.5th and 97.5th percentiles; bootstrap standard error 0.004. Krippendorff's alpha: 1 − (observed disagreement / expected disagreement), using (n−1) weighting. Category order: not detected; no ordinal alpha reported.

What this can't tell you

Agreement is not correctness. No coefficient on this page establishes whether the panel is right. The figures describe this panel on these items only; a different rater set or item pool with a different category distribution will produce a different kappa from the same rubric. The Landis & Koch bands are convention without theoretical foundation. Consider whether agreement is uniform across item subgroups (e.g., by source, difficulty, or domain) by requesting a stratified re-run or a finer-grained export.

Rate this report Was this the answer you needed?
The exact source that produced this report — yours to keep, read, and re-run.
Download PDF
How this was computed method · R source · citation
The code that did it

Multi-Rater Agreement — Fleiss Kappa

Three or more raters, one categorical judgement per item: how much do they agree, how much of that is more than chance would have produced, and which rater is pulling the group apart? The analysis computes Fleiss' kappa with a standard error and confidence interval, Krippendorff's alpha (nominal, and ordinal when the categories turn out to be ordered), a per-category kappa showing which categories the raters fight over, a per-rater agreement with the consensus plus a leave-one-rater-out kappa that names the outlier, and the item-level distribution of how strong the consensus actually was.

Why This Method?

Raw agreement across a panel is not interpretable on its own: a panel that answers "Yes" to almost everything will agree constantly without reading a single item. Fleiss' kappa subtracts the agreement chance alone would deliver given how often each category is used overall. Krippendorff's alpha answers the same question through a different route and, unlike Fleiss', keeps working when not every rater judged every item.

What This Analysis Covers

  • Raw pairwise agreement, chance agreement, and Fleiss' kappa with SE + CI
  • Krippendorff's alpha, nominal and (when the labels are ordered) ordinal
  • Explicit accounting for missing ratings — what each statistic used
  • Per-category kappa: which categories the panel cannot pin down
  • Per-rater agreement with the consensus and leave-one-rater-out kappa
  • The item-level consensus distribution, and the kappa paradox when it fires

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {rater_1 .. rater_N}. Both real-world layouts are accepted and the shape is detected from the data, never assumed. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

Rule 1 — the labels are numbers (1, 2, 3 / 0-10 scales).

num <- suppressWarnings(as.numeric(clean))
  if (!anyNA(num) && length(unique(num)) == length(num)) {
    return(list(ordered = TRUE, order = levs[order(num)],
                rule = "the labels are numbers, so they were ordered by value"))
  }

Rule 2 — the labels start with a number ("1 - Poor", "2 = Fair").

lead <- suppressWarnings(as.numeric(sub("^\\s*([-+]?[0-9]+(\\.[0-9]+)?)\\s*[-=.):|].*$",
                                          "\\1", clean)))
  has_lead <- grepl("^\\s*[-+]?[0-9]+(\\.[0-9]+)?\\s*[-=.):|]", clean)
  if (all(has_lead) && !anyNA(lead) && length(unique(lead)) == length(lead)) {
    return(list(ordered = TRUE, order = levs[order(lead)],
                rule = "every label starts with a number, so they were ordered by that number"))
  }

Rule 3 — the labels are all drawn from one recognised ordinal vocabulary.

scales <- list(
    "an agreement scale" = c("strongly disagree", "disagree", "somewhat disagree",
                             "slightly disagree", "neutral", "neither agree nor disagree",
                             "slightly agree", "somewhat agree", "agree", "strongly agree"),
    "a quality scale" = c("very poor", "poor", "below average", "fair", "average",
                          "satisfactory", "good", "very good", "excellent", "outstanding"),
    "a frequency scale" = c("never", "rarely", "seldom", "sometimes", "occasionally",
                            "often", "frequently", "usually", "always"),
    "a severity scale" = c("none", "minimal", "mild", "moderate", "severe",
                           "very severe", "extreme", "critical"),
    "a magnitude scale" = c("very low", "low", "medium", "moderate", "high", "very high"),
    "a satisfaction scale" = c("very dissatisfied", "dissatisfied", "neutral",
                               "satisfied", "very satisfied"),
    "a likelihood scale" = c("very unlikely", "unlikely", "possible", "likely",
                             "very likely", "certain"),
    "a decision scale" = c("reject", "major revision", "revise", "minor revision",
                           "accept with changes", "accept"),
    "a priority scale" = c("trivial", "minor", "moderate", "major", "critical", "blocker")
  )
  low <- tolower(clean)
  if (length(levs) >= 3 && !any(duplicated(low))) {
    for (nm in names(scales)) {
      sc <- scales[[nm]]
      if (all(low %in% sc)) {
        return(list(ordered = TRUE, order = levs[order(match(low, sc))],
                    rule = paste0("the labels are all points on ", nm,
                                  ", so they were ordered along it")))
      }
    }
  }

  list(ordered = FALSE, order = levs,
       rule = "no numeric values, numeric prefixes, or recognised ordinal wording were found in the labels")
}

Step 1: Collect the mapped judgement columns, in mapped order

rc <- grep("^rater_[0-9]+$", names(df), value = TRUE)
  rc <- rc[order(as.numeric(sub("^rater_", "", rc)))]
  ch <- humanize_semantic(rc, col_map)

  if (length(rc) < 3) {
    stop(sprintf("Only %d judgement column(s) were mapped(%s). This analysis needs three or more raters. For exactly two raters the right tool is Cohen&#x27;s kappa, not Fleiss' — or, if your file has one row per rating, map the three columns holding the item, the rater and the judgement.",
                 length(rc),
                 if (length(ch) > 0) paste(ch, collapse = ", ") else "none"))
  }

  cols <- lapply(rc, function(cn) as_label(df[[cn]]))
  names(cols) <- rc
  d_distinct <- vapply(cols, function(v) length(unique(v[v != ""])), integer(1))

Step 2: Detect the input SHAPE from the data, not from the mapping.

Wide means one column per rater and one row per item. Long means one row per rating, with an item column, a rater column and a judgement column. A long file is unmistakable in the counts: the item column carries many distinct values, each repeating once per rater, while the other two carry only a handful.

shape <- "wide"
  shape_rule <- ""
  li <- ri <- ji <- NA_integer_
  if (length(rc) == 3) {
    ii <- which.max(d_distinct)
    others <- setdiff(seq_len(3), ii)
    avg_rep <- if (d_distinct[ii] > 0) initial_rows / d_distinct[ii] else 0
    if (d_distinct[ii] >= 10 && avg_rep >= 2.5 && min(d_distinct[others]) >= 2 &&
        d_distinct[ii] > 3 * max(d_distinct[others])) {
      item_v <- cols[[ii]]

Which of the remaining two is the RATER? The one for which each item appears at most once — a rater judges an item once, whereas several raters give an item the same judgement all the time.

dupr <- vapply(others, function(o)
        sum(duplicated(paste0(item_v, "\r", cols[[o]]))), numeric(1))
      ri <- others[which.min(dupr)]
      ji <- setdiff(others, ri)
      li <- ii
      shape <- "long"
      shape_rule <- sprintf(
        "%s holds %d distinct values each repeating about %s times, while %s and %s hold only %d and %d distinct values — that is one row per rating, not one column per rater",
        ch[li], d_distinct[li], r2(avg_rep), ch[ri], ch[ji],
        d_distinct[ri], d_distinct[ji])
    }
  }

Step 3: Flatten either shape into one table of (item, rater, label)

n_dup_ratings <- 0L
  if (shape == "long") {
    item_v <- cols[[li]]; rater_v <- cols[[ri]]; lab_v <- cols[[ji]]
    ok <- item_v != "" & rater_v != ""
    n_blank_rows <- sum(!ok)
    item_v <- item_v[ok]; rater_v <- rater_v[ok]; lab_v <- lab_v[ok]

The same rater judging the same item twice is a data error, not a second opinion; the first judgement is kept and the count is reported.

key <- paste0(item_v, "\r", rater_v)
    dupd <- duplicated(key)
    n_dup_ratings <- sum(dupd)
    item_v <- item_v[!dupd]; rater_v <- rater_v[!dupd]; lab_v <- lab_v[!dupd]
    if (length(unique(rater_v)) > 20) {
      stop(sprintf("%s holds %d distinct raters. Above 20 raters this looks like an identifier column rather than a panel; map the column naming who gave each judgement.",
                   ch[ri], length(unique(rater_v))))
    }
    row_unit <- "rating"
    shape_rule <- paste0("the three mapped columns were read as a long file because ", shape_rule)
  } else {

Wide: a column holding a distinct value on nearly every row is an identifier or free text, not a category, and is refused by name. Both an absolute cap and a proportional test are needed — a key column in a short file has few distinct values in absolute terms but repeats nothing.

n_filled <- vapply(cols, function(v) sum(v != ""), integer(1))
    bad <- which(d_distinct > 25 | (d_distinct >= 10 & d_distinct > 0.5 * n_filled))
    if (length(bad) > 0) {
      b <- bad[1]
      why <- if (d_distinct[b] > 25)
        "which is too many to be a set of categories"
      else
        sprintf("across %d judgement(s) — nearly one per row, so almost nothing repeats",
                n_filled[b])
      stop(sprintf("%s holds %d distinct values, %s. That makes it an identifier or free text rather than a rater&#x27;s category choice. Map one column per rater, each holding that rater's category. If your file instead has one row per rating, map exactly three columns: the item, the rater, and the judgement.",
                   ch[b], d_distinct[b], why))
    }
    n_blank_rows <- 0L
    item_v <- rep(as.character(seq_len(initial_rows)), times = length(rc))
    rater_v <- rep(ch, each = initial_rows)
    lab_v <- unlist(cols, use.names = FALSE)
    keep <- lab_v != ""
    item_v <- item_v[keep]; rater_v <- rater_v[keep]; lab_v <- lab_v[keep]
    row_unit <- "item"
    shape_rule <- sprintf(
      "the %d mapped columns were read as one column per rater, because each holds only a handful of repeated labels(%s distinct) rather than one value per row",
      length(rc), paste(d_distinct, collapse = ", "))
  }

Step 4: Merge labels that differ only in capitalisation, and say so.

by_case <- split(lab_v, tolower(lab_v))
  canon <- list(); n_case_merged <- 0L; case_example <- NULL
  for (keyc in names(by_case)) {
    spellings <- by_case[[keyc]]
    uq <- unique(spellings)
    tab <- sort(table(spellings), decreasing = TRUE)
    winner <- names(tab)[1]
    canon[[keyc]] <- winner
    if (length(uq) > 1) {
      n_case_merged <- n_case_merged + sum(spellings != winner)
      if (is.null(case_example)) case_example <- c(setdiff(uq, winner)[1], winner)
    }
  }
  lab_v <- unlist(canon[tolower(lab_v)], use.names = FALSE)

Step 5: Refuse free text; lump a long tail of rare labels into "Other".

freq <- sort(table(lab_v), decreasing = TRUE)
  n_lumped <- 0L
  if (length(freq) > 25) {
    stop(sprintf("The judgement columns(%s) together hold %d distinct values — that looks like free text rather than a set of categories. Map the columns holding each rater&#x27;s category choice.",
                 paste(ch, collapse = ", "), length(freq)))
  }
  if (length(freq) > 12) {
    keep_lab <- names(freq)[1:11]
    n_lumped <- length(freq) - 11L
    lab_v[!(lab_v %in% keep_lab)] <- "Other"
  }

  rater_names <- sort(unique(rater_v))
  n_raters <- length(rater_names)
  if (n_raters < 3) {
    stop(sprintf("Only %d rater(s) were found in %s. Fleiss&#x27; kappa needs three or more; with exactly two raters use Cohen's kappa instead.",
                 n_raters, if (shape == "long") ch[ri] else paste(ch, collapse = ", ")))
  }

  levs_all <- sort(unique(lab_v))
  if (length(levs_all) < 2) {
    stop(sprintf("Every judgement in %s is \"%s\", so there is only one category and agreement beyond chance is undefined — nothing varies for kappa to explain.",
                 paste(ch, collapse = ", "), levs_all[1]))
  }

Step 6: Decide the category order (detected, not assumed)

det <- detect_category_order(levs_all)
  if (det$ordered) {
    levs <- det$order
  } else {
    tot <- vapply(levs_all, function(l) sum(lab_v == l), numeric(1))
    levs <- levs_all[order(-tot, levs_all)]
  }
  k <- length(levs)

Step 7: Build the item-by-category count matrix

items <- unique(item_v)
  Nmat <- matrix(0, nrow = length(items), ncol = k,
                 dimnames = list(NULL, levs))
  im <- match(item_v, items)
  cm <- match(lab_v, levs)
  tabm <- table(factor(im, levels = seq_along(items)),
                factor(cm, levels = seq_len(k)))
  Nmat[] <- as.numeric(tabm)
  m_i <- rowSums(Nmat)

  n_ratings <- sum(m_i)
  all_idx <- which(m_i >= 2)
  n_items_dropped <- length(items) - length(all_idx)
  complete_idx <- which(m_i == n_raters)
  n_items_partial <- sum(m_i >= 2 & m_i < n_raters)
  n_missing <- n_raters * length(items) - n_ratings

  if (length(all_idx) < 20) {
    stop(sprintf("Only %d item(s) carry judgements from at least two of %s. Multi-rater agreement needs at least 20 such items before the numbers mean anything.",
                 length(all_idx), paste(ch, collapse = ", ")))
  }

Step 8: Fleiss' kappa. It is a complete-cases statistic: the classical

coefficient assumes every item was judged by the same number of raters. When ratings are missing the module reports BOTH the complete-case value (the one other software produces) and a generalized value over every item with two or more ratings — and says which items each one used.

fleiss_calc <- function(ii) {
    m <- m_i[ii]
    sq <- rowSums(Nmat[ii, , drop = FALSE]^2)
    P_i <- (sq - m) / (m * (m - 1))
    P_bar <- mean(P_i)
    tot <- sum(m)
    p <- colSums(Nmat[ii, , drop = FALSE]) / tot
    P_e <- sum(p^2)
    kap <- if ((1 - P_e) > 1e-12) (P_bar - P_e) / (1 - P_e) else NA_real_
    list(P_bar = P_bar, P_e = P_e, kappa = kap, p = p, tot = tot,
         n = length(ii), m_const = if (length(unique(m)) == 1) m[1] else NA_real_)
  }

  basis <- if (length(complete_idx) >= 20) "complete" else "all"
  basis_idx <- if (basis == "complete") complete_idx else all_idx
  fb <- fleiss_calc(basis_idx)
  fa <- fleiss_calc(all_idx)
  p_bar <- fb$P_bar; p_e <- fb$P_e; kappa <- fb$kappa
  p_j <- fb$p
  kappa_all <- fa$kappa
  n_basis <- fb$n

Fleiss' null-hypothesis standard error (Fleiss 1971), which exists only when every item on the basis carries the same number of raters. It is the variance of kappa when kappa is truly zero, so it belongs to the significance test and NOT to the confidence interval.

se_null <- NA_real_; z_stat <- NA_real_; p_value <- NA_real_
  if (!is.na(fb$m_const) && fb$m_const >= 2) {
    R <- fb$m_const
    S2 <- sum(p_j^2); S3 <- sum(p_j^3)
    inner <- S2 - (2 * R - 3) * S2^2 + 2 * (R - 2) * S3
    if (is.finite(inner) && inner > 0 && (1 - S2) > 1e-12) {
      se_null <- sqrt(2 / (n_basis * R * (R - 1)) * inner) / (1 - S2)
      z_stat <- kappa / se_null
      p_value <- 2 * pnorm(-abs(z_stat))
    }
  }

Step 9: Krippendorff's alpha, via the coincidence matrix. Alpha is the

statistic that tolerates missing ratings, so it uses EVERY item carrying at least two judgements — including the ones Fleiss had to set aside.

Ocon <- t(apply(Nmat, 1, function(row) {
    m <- sum(row)
    if (m < 2) return(rep(0, k * k))
    as.vector((outer(row, row) - diag(row, nrow = k)) / (m - 1))
  }))
  if (k == 1) Ocon <- matrix(Ocon, ncol = 1)

  LO <- outer(seq_len(k), seq_len(k), pmin)
  HI <- outer(seq_len(k), seq_len(k), pmax)
  ordinal_ok <- det$ordered && k >= 3

  alpha_calc <- function(ii) {
    o <- matrix(colSums(Ocon[ii, , drop = FALSE]), nrow = k, ncol = k)
    nc <- rowSums(o)
    ntot <- sum(nc)
    a_nom <- NA_real_; a_ord <- NA_real_
    De_n <- ntot^2 - sum(nc^2)
    if (is.finite(De_n) && De_n > 1e-12 && ntot > 1) {
      Do_n <- ntot - sum(diag(o))
      a_nom <- 1 - (ntot - 1) * Do_n / De_n
    }
    if (ordinal_ok && ntot > 1) {
      cum0 <- c(0, cumsum(nc))
      S <- matrix(cum0[HI + 1] - cum0[LO], nrow = k, ncol = k)
      Dm <- (S - outer(nc, nc, function(x, y) (x + y) / 2))^2
      De_o <- sum(outer(nc, nc) * Dm)
      if (is.finite(De_o) && De_o > 1e-12) {
        a_ord <- 1 - (ntot - 1) * sum(o * Dm) / De_o
      }
    }
    list(nominal = a_nom, ordinal = a_ord)
  }

  aa <- alpha_calc(all_idx)
  alpha_nom <- aa$nominal; alpha_ord <- aa$ordinal
  alpha_nom_basis <- alpha_calc(basis_idx)$nominal

Step 10: Confidence intervals by resampling ITEMS. The closed-form

variance above is a null-hypothesis quantity and is wrong for an interval around a non-zero estimate; alpha has no simple closed form at all. A seeded nonparametric bootstrap over items answers both, and treats items as the sampling unit, which is what they are.

B <- 500L
  set.seed(42)
  nb <- length(basis_idx); na_ <- length(all_idx)
  bk <- numeric(B); ban <- numeric(B); bao <- numeric(B)
  for (b in seq_len(B)) {
    ib <- basis_idx[sample.int(nb, nb, replace = TRUE)]
    bk[b] <- fleiss_calc(ib)$kappa
    ia <- all_idx[sample.int(na_, na_, replace = TRUE)]
    ab <- alpha_calc(ia)
    ban[b] <- ab$nominal; bao[b] <- ab$ordinal
  }
  qci <- function(v) {
    v <- v[is.finite(v)]
    if (length(v) < 50) return(c(NA_real_, NA_real_, NA_real_))
    c(as.numeric(quantile(v, 0.025, names = FALSE)),
      as.numeric(quantile(v, 0.975, names = FALSE)), sd(v))
  }
  qk <- qci(bk); qan <- qci(ban); qao <- qci(bao)
  ci_low <- qk[1]; ci_high <- qk[2]; se_boot <- qk[3]
  alpha_nom_lo <- qan[1]; alpha_nom_hi <- qan[2]
  alpha_ord_lo <- qao[1]; alpha_ord_hi <- qao[2]

  band <- landis_koch(kappa)

Step 11: Per-category kappa — which categories the panel fights over.

Fleiss' category-specific coefficient, generalized to a variable number of raters per item.

mb <- m_i[basis_idx]
  Nb <- Nmat[basis_idx, , drop = FALSE]
  tot_b <- sum(mb)
  cat_kappa <- vapply(seq_len(k), function(j) {
    pj <- sum(Nb[, j]) / tot_b
    den <- tot_b * pj * (1 - pj)
    if (!is.finite(den) || den <= 1e-12) return(NA_real_)
    obs <- sum(Nb[, j] * (mb - Nb[, j]) / (mb - 1))
    1 - obs / den
  }, numeric(1))
  category_df <- data.frame(
    category = levs,
    category_kappa = round(cat_kappa, 4),
    share_pct = round(100 * colSums(Nb) / tot_b, 2),
    n_uses = as.numeric(colSums(Nb)),
    stringsAsFactors = FALSE
  )
  ok_cat <- which(!is.na(cat_kappa))
  worst_cat <- if (length(ok_cat) > 0) levs[ok_cat[which.min(cat_kappa[ok_cat])]] else NA_character_
  worst_val <- if (length(ok_cat) > 0) min(cat_kappa[ok_cat]) else NA_real_
  best_cat <- if (length(ok_cat) > 0) levs[ok_cat[which.max(cat_kappa[ok_cat])]] else NA_character_
  best_val <- if (length(ok_cat) > 0) max(cat_kappa[ok_cat]) else NA_real_

With exactly two categories every disagreement is the same disagreement, so both category-specific kappas equal each other and the overall coefficient. Naming a "worst" category there would be an artefact of tie-breaking.

cat_degenerate <- k == 2 ||
    (length(ok_cat) > 1 && (best_val - worst_val) < 1e-9)

Step 12: Per-rater agreement with the rest of the panel, and the

leave-one-rater-out kappa. The second is the actionable one: it says what the panel's agreement would be if this rater were not in it.

basis_items <- items[basis_idx]
  in_basis <- item_v %in% basis_items
  bi_item <- item_v[in_basis]; bi_rater <- rater_v[in_basis]; bi_lab <- lab_v[in_basis]
  bi_row <- match(bi_item, basis_items)
  bi_col <- match(bi_lab, levs)
  m_of_row <- mb[bi_row]
  n_same <- Nb[cbind(bi_row, bi_col)]

  rater_obs <- numeric(n_raters); rater_kap <- numeric(n_raters)
  rater_wo <- numeric(n_raters); rater_n <- numeric(n_raters)
  for (r in seq_len(n_raters)) {
    sel <- bi_rater == rater_names[r]
    den <- sum(m_of_row[sel] - 1)
    rater_n[r] <- sum(sel)
    rater_obs[r] <- if (den > 0) sum(n_same[sel] - 1) / den else NA_real_
    rater_kap[r] <- if (is.na(rater_obs[r]) || (1 - p_e) <= 1e-12) NA_real_
                    else (rater_obs[r] - p_e) / (1 - p_e)

Leave-one-rater-out: strip this rater's ratings and recompute Fleiss on the items that still carry two or more judgements.

Nw <- Nb
    idxw <- which(sel)
    if (length(idxw) > 0) {
      dec <- table(factor(bi_row[idxw], levels = seq_len(nrow(Nb))),
                   factor(bi_col[idxw], levels = seq_len(k)))
      Nw <- Nw - matrix(as.numeric(dec), nrow = nrow(Nb), ncol = k)
    }
    mw <- rowSums(Nw)
    keepw <- which(mw >= 2)
    if (length(keepw) >= 5) {
      sqw <- rowSums(Nw[keepw, , drop = FALSE]^2)
      Pw <- (sqw - mw[keepw]) / (mw[keepw] * (mw[keepw] - 1))
      totw <- sum(mw[keepw])
      pw <- colSums(Nw[keepw, , drop = FALSE]) / totw
      Pew <- sum(pw^2)
      rater_wo[r] <- if ((1 - Pew) > 1e-12) (mean(Pw) - Pew) / (1 - Pew) else NA_real_
    } else {
      rater_wo[r] <- NA_real_
    }
  }

  rater_df <- data.frame(
    rater = rater_names,
    agreement_pct = round(100 * rater_obs, 2),
    rater_kappa = round(rater_kap, 4),
    kappa_without_rater = round(rater_wo, 4),
    items_rated = rater_n,
    stringsAsFactors = FALSE
  )
  ok_r <- which(!is.na(rater_kap))
  outlier_rater <- if (length(ok_r) > 0) rater_names[ok_r[which.min(rater_kap[ok_r])]] else NA_character_
  outlier_kappa <- if (length(ok_r) > 0) min(rater_kap[ok_r]) else NA_real_
  best_rater <- if (length(ok_r) > 0) rater_names[ok_r[which.max(rater_kap[ok_r])]] else NA_character_
  best_rater_kappa <- if (length(ok_r) > 0) max(rater_kap[ok_r]) else NA_real_
  outlier_gain <- if (!is.na(outlier_rater) && !is.na(rater_wo[match(outlier_rater, rater_names)]))
    rater_wo[match(outlier_rater, rater_names)] - kappa else NA_real_
  rater_spread <- if (length(ok_r) > 1) max(rater_kap[ok_r]) - min(rater_kap[ok_r]) else NA_real_

Step 13: Item-level consensus. Banded by each item's MODAL share —

the largest number of raters landing on one category, over the raters who judged it. Only the maximum COUNT is used, never its position, so an evenly split item needs no arbitrary winner.

modal_share <- apply(Nb, 1, max) / mb
  band_labels <- c("No majority — the raters split",
                   "Simple majority — more than half",
                   "Strong majority — more than two thirds",
                   "Unanimous — every rater agreed")
  band_idx <- ifelse(modal_share >= 1, 4L,
              ifelse(modal_share > 2 / 3, 3L,
              ifelse(modal_share > 0.5, 2L, 1L)))
  item_df <- data.frame(
    consensus_band = band_labels,
    n_items = as.numeric(tabulate(band_idx, nbins = 4)),
    stringsAsFactors = FALSE
  )
  item_df$share_pct <- round(100 * item_df$n_items / n_basis, 2)
  n_unanimous <- item_df$n_items[4]
  n_split <- item_df$n_items[1]

Step 14: Per-rater marginals — how each rater uses the scale. These

are the habits behind both the chance floor and any outlier.

marg <- matrix(0, nrow = n_raters, ncol = k, dimnames = list(rater_names, levs))
  tabr <- table(factor(bi_rater, levels = rater_names),
                factor(bi_lab, levels = levs))
  marg[] <- as.numeric(tabr)
  marg_share <- marg / pmax(rowSums(marg), 1)
  overall_share <- colSums(marg) / sum(marg)
  divergence <- rowSums(abs(sweep(marg_share, 2, overall_share)))
  show_r <- rater_names
  n_raters_hidden <- 0L
  if (n_raters > 8) {
    show_r <- rater_names[order(-divergence)][1:8]
    n_raters_hidden <- n_raters - 8L
  }
  marginal_df <- data.frame(
    category = rep(levs, times = length(show_r)),
    rater = rep(show_r, each = k),
    share_pct = round(100 * as.vector(t(marg_share[show_r, , drop = FALSE])), 2),
    stringsAsFactors = FALSE
  )

  prev_i <- which.max(p_j)
  max_prev <- p_j[prev_i]
  prev_cat <- levs[prev_i]
  paradox <- isTRUE(p_bar >= 0.70 && !is.na(kappa) && kappa < 0.60 && max_prev >= 0.60)

Step 15: The results table

basis_phrase <- if (basis == "complete")
    sprintf("the %s item(s) every one of the %d raters judged", fmt_n(n_basis), n_raters)
  else
    sprintf("all %s item(s) carrying at least two judgements", fmt_n(n_basis))

  rows <- list(
    list("Raw pairwise agreement", p_bar, NA_real_, NA_real_,
         sprintf("Across every pair of raters who judged the same item, %s of those pairs chose the same category.",
                 pct1(p_bar))),
    list("Agreement expected by chance", p_e, NA_real_, NA_real_,
         sprintf("What a panel using the categories this often would have agreed on by luck alone(%s).",
                 pct1(p_e))),
    list("Fleiss&#x27; kappa", kappa, ci_low, ci_high,
         sprintf("%s of the agreement left available above chance was achieved, on %s — %s on the Landis & Koch convention.",
                 pct1(kappa), basis_phrase, band))
  )
  if (n_missing > 0 && basis == "complete") {
    rows[[length(rows) + 1]] <- list(
      "Fleiss&#x27; kappa (all items, generalized)", kappa_all, NA_real_, NA_real_,
      sprintf("The same coefficient generalized to a variable number of raters per item, so it uses all %s items with two or more judgements instead of only the %s complete ones.",
              fmt_n(length(all_idx)), fmt_n(n_basis)))
  }
  rows[[length(rows) + 1]] <- list(
    "Krippendorff&#x27;s alpha (nominal)", alpha_nom, alpha_nom_lo, alpha_nom_hi,
    sprintf("The disagreement-based coefficient, computed on all %s items with two or more judgements — it does not require every rater to have judged every item.",
            fmt_n(length(all_idx))))
  if (!is.na(alpha_ord)) {
    rows[[length(rows) + 1]] <- list(
      "Krippendorff&#x27;s alpha (ordinal)", alpha_ord, alpha_ord_lo, alpha_ord_hi,
      "The same coefficient with an ordinal distance between categories, so a near-miss between neighbouring levels counts as a smaller disagreement than a jump across the scale.")
  }
  if (n_missing > 0 && basis == "complete" && !is.na(alpha_nom_basis)) {
    rows[[length(rows) + 1]] <- list(
      "Krippendorff&#x27;s alpha (complete items only)", alpha_nom_basis, NA_real_, NA_real_,
      "Alpha restricted to the same complete items Fleiss&#x27; kappa used, so the two headline coefficients can be compared on identical data.")
  }

  agreement_df <- data.frame(
    statistic = vapply(rows, function(r) r[[1]], character(1)),
    estimate = round(vapply(rows, function(r) as.numeric(r[[2]]), numeric(1)), 4),
    ci_low = round(vapply(rows, function(r) as.numeric(r[[3]]), numeric(1)), 4),
    ci_high = round(vapply(rows, function(r) as.numeric(r[[4]]), numeric(1)), 4),
    interpretation = vapply(rows, function(r) r[[5]], character(1)),
    stringsAsFactors = FALSE
  )

Step 16: Methods disclosure

se_line <- if (is.na(se_null)) {
    sprintf("The classical null-hypothesis standard error is not available here, because the items on the analysis basis do not all carry the same number of raters. The significance test is therefore not reported; read the bootstrap interval instead, which does not need that assumption.")
  } else {
    sprintf("Fleiss&#x27; null-hypothesis standard error is %s, giving z = %s and %s against the hypothesis that agreement is no better than chance. That variance describes kappa when kappa is truly zero, so it belongs to this test and NOT to the confidence interval.",
            r3(se_null), r2(z_stat), fmt_pp(p_value))
  }
  methods_df <- data.frame(
    item = c(
      "Design",
      "Input shape",
      "Raw pairwise agreement",
      "Chance agreement",
      "Fleiss&#x27; kappa",
      "Significance test",
      "Confidence intervals",
      "Missing ratings",
      "Krippendorff&#x27;s alpha",
      "Category order",
      "Per-category kappa",
      "Per-rater figures",
      "Benchmark labels",
      "What agreement is not"
    ),
    detail = c(
      sprintf("%s item(s) judged by %d rater(s) across %d categories, %s judgement(s) in total.",
              fmt_n(length(items)), n_raters, k, fmt_n(n_ratings)),
      sprintf("Detected from the data, not from the mapping: %s.", shape_rule),
      sprintf("Mean over items of the share of rater PAIRS on that item choosing the same category: %s.",
              pct1(p_bar)),
      sprintf("Sum of the squared overall category shares(%s): %s.",
              paste(sprintf("%s %s", levs, vapply(p_j, pct1, character(1))), collapse = ", "),
              pct1(p_e)),
      sprintf("(observed - expected) / (1 - expected) = (%s - %s) / (1 - %s) = %s, computed on %s.",
              r3(p_bar), r3(p_e), r3(p_e), r3(kappa), basis_phrase),
      se_line,
      sprintf("A seeded nonparametric bootstrap resampling ITEMS with replacement(%d resamples, fixed seed), taking the 2.5th and 97.5th percentiles. The bootstrap standard error of kappa is %s. This is used rather than a closed form because the null variance above is the wrong quantity for an interval around a non-zero estimate, and alpha has no simple closed-form variance.",
              B, r3(se_boot)),
      if (n_missing > 0)
        sprintf("%s of the %s possible rater-by-item judgements are absent. Fleiss&#x27; kappa is a complete-cases statistic and used %s; Krippendorff's alpha used all %s items with at least two judgements; %s item(s) carrying a single judgement contribute to neither, because agreement needs at least two opinions.",
                fmt_n(n_missing), fmt_n(n_raters * length(items)), basis_phrase,
                fmt_n(length(all_idx)), fmt_n(n_items_dropped))
      else
        "No ratings are missing: every rater judged every item, so Fleiss&#x27; kappa and Krippendorff's alpha are computed on exactly the same data.",
      sprintf("Built from the coincidence matrix: alpha = 1 - observed disagreement / expected disagreement, with(n-1) weighting so it is defined for small samples. Nominal alpha treats every disagreement as equal. %s",
              if (!is.na(alpha_ord))
                "Ordinal alpha weights a disagreement by the distance between the categories along the detected order."
              else
                "Ordinal alpha is not reported for these categories."),
      sprintf("Orderedness was detected, not assumed: %s.", det$rule),
      "Fleiss&#x27; category-specific coefficient, generalized to a variable number of raters per item: one minus the observed splits on that category over the splits expected from its overall share.",
      sprintf("A rater&#x27;s agreement figure is the share of the OTHER raters' judgements on the same items that matched this rater's own, chance-corrected against the same %s floor. The leave-one-out column recomputes Fleiss' kappa for the panel with that rater removed.",
              pct1(p_e)),
      "Landis & Koch(1977) labels(slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.",
      "Agreement is not correctness: a panel can agree unanimously and be unanimously wrong. These figures describe this panel on these items and do not generalise to other raters or other items."
    ),
    stringsAsFactors = FALSE
  )

Step 17: Headline metrics + the computed one-paragraph answer

metrics <- list(
    `Items Rated`        = as.numeric(length(items)),
    `Raters`             = as.numeric(n_raters),
    `Categories`         = as.numeric(k),
    `Raw Agreement`      = pct1(p_bar),
    `Fleiss Kappa`       = round(kappa, 3),
    `Fleiss 95% CI`      = paste0(r3(ci_low), " to ", r3(ci_high)),
    `Krippendorff Alpha` = round(alpha_nom, 3),
    `Agreement Strength` = band,
    `Least Aligned Rater` = if (is.na(outlier_rater)) "not identifiable" else outlier_rater,
    `Kappa vs Chance p`  = fmt_p(p_value)
  )

  paradox_sentence <- if (paradox) {
    paste0(
      " Read the two agreement numbers together before quoting either: raw pairwise agreement is high(",
      pct1(p_bar), ") while kappa is only ", r3(kappa),
      ", which is the well-known kappa paradox rather than a contradiction. ",
      pct1(max_prev), " of all judgements fell into the single category \"", prev_cat,
      "\", so a panel with these habits would already have agreed on ", pct1(p_e),
      " of rater pairs by chance; only ", pct1(1 - p_e),
      " of the scale was left for skill to win, and ", pct1(kappa), " of that remainder was won.")
  } else {
    paste0(
      " Chance alone would have produced ", pct1(p_e),
      " pairwise agreement given how often each category is used, leaving ",
      pct1(1 - p_e), " of the scale available above chance, of which ",
      pct1(kappa), " was achieved.")
  }

  missing_sentence <- if (n_missing > 0) {
    paste0(" ", fmt_n(n_missing), " of the ", fmt_n(n_raters * length(items)),
           " possible judgements are missing, which the two coefficients handle differently: Fleiss&#x27; kappa is a complete-cases statistic and used ",
           basis_phrase, ", while Krippendorff&#x27;s alpha used all ", fmt_n(length(all_idx)),
           " items carrying at least two judgements",
           if (n_items_dropped > 0)
             paste0(" and ", fmt_n(n_items_dropped),
                    " item(s) with a single judgement were used by neither")
           else "",
           ". Where ratings are missing, alpha is the coefficient to quote.")
  } else ""

  ordinal_sentence <- if (!is.na(alpha_ord)) {
    paste0(" The categories are ordered(", det$rule,
           "), so ordinal alpha(", r3(alpha_ord),
           ") is also reported; it counts a disagreement between neighbouring levels as smaller than a jump across the scale, which is why it sits above the nominal value of ",
           r3(alpha_nom), ".")
  } else {
    paste0(" Ordinal alpha is not reported because ", det$rule,
           "; treating unordered labels as if some disagreements were milder than others would invent structure the data does not carry.")
  }

  outlier_sentence <- if (!is.na(outlier_rater) && !is.na(outlier_gain)) {
    paste0(" ", outlier_rater, " is the least aligned member of the panel(rater kappa ",
           r3(outlier_kappa), " against ", r3(best_rater_kappa), " for ", best_rater,
           "); with that rater excluded the panel&#x27;s kappa would be ",
           r3(kappa + outlier_gain),
           if (outlier_gain > 0) paste0(", ", r3(outlier_gain), " higher than it is now")
           else paste0(", ", r3(abs(outlier_gain)), " lower than it is now"),
           ".")
  } else ""

  json_output <- list(
    answer = paste0(
      fmt_n(n_raters), " raters judged ", fmt_n(length(items)), " items across ", k,
      " categories, agreeing on ", pct1(p_bar),
      " of rater pairs. Fleiss&#x27; kappa is ", r3(kappa),
      " (95% CI ", r3(ci_low), " to ", r3(ci_high), "), ", band,
      " agreement on the Landis & Koch convention, and ",
      if (!is.na(p_value) && p_value < 0.05)
        paste0("clearly better than chance(", fmt_pp(p_value), ")")
      else if (!is.na(p_value))
        paste0("not distinguishable from chance(", fmt_pp(p_value), ")")
      else
        "could not be tested against chance because the raters did not all judge the same items",
      "; Krippendorff&#x27;s alpha is ", r3(alpha_nom), ".",
      paradox_sentence, missing_sentence, ordinal_sentence, outlier_sentence,
      if (cat_degenerate)
        paste0(" With only ", k,
               " categories every disagreement is the same disagreement, so the category-specific kappas are necessarily equal to each other and to the overall coefficient — there is no per-category story to tell here.")
      else
        paste0(" The panel is least consistent on the category \"", worst_cat,
               "\" (category kappa ", r3(worst_val), ") and most consistent on \"",
               best_cat, "\" (", r3(best_val), ")."),
      " Agreement is not correctness: a panel can agree unanimously and be unanimously wrong."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "agreement_results", "rater_consensus",
        "category_kappa", "item_consensus", "rater_marginals", "methods"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  final_rows <- if (row_unit == "item") length(all_idx) else n_ratings

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = max(0, initial_rows - final_rows),
    row_unit = row_unit, shape = shape, shape_rule = shape_rule,
    n_blank_rows = n_blank_rows, n_dup_ratings = n_dup_ratings,
    n_case_merged = n_case_merged, case_example = case_example, n_lumped = n_lumped,
    col_names = ch, rater_names = rater_names, n_raters = n_raters,
    levs = levs, k_cats = k, n_items = length(items), n_ratings = n_ratings,
    n_missing = n_missing, n_items_partial = n_items_partial,
    n_items_dropped = n_items_dropped, n_all = length(all_idx),
    basis = basis, n_basis = n_basis, n_basis_ratings = tot_b,
    basis_phrase = basis_phrase,
    p_bar = p_bar, p_e = p_e, p_j = p_j, kappa = kappa, kappa_all = kappa_all,
    se_null = se_null, se_boot = se_boot, z_stat = z_stat, p_value = p_value,
    ci_low = ci_low, ci_high = ci_high,
    alpha_nom = alpha_nom, alpha_nom_lo = alpha_nom_lo, alpha_nom_hi = alpha_nom_hi,
    alpha_ord = alpha_ord, alpha_nom_basis = alpha_nom_basis,
    ordered_flag = det$ordered, order_rule = det$rule, band = band,
    paradox = paradox, paradox_sentence = paradox_sentence,
    missing_sentence = missing_sentence, ordinal_sentence = ordinal_sentence,
    outlier_sentence = outlier_sentence,
    max_prev = max_prev, prev_cat = prev_cat,
    outlier_rater = outlier_rater, outlier_kappa = outlier_kappa,
    outlier_gain = outlier_gain, best_rater = best_rater,
    best_rater_kappa = best_rater_kappa, rater_spread = rater_spread,
    n_raters_hidden = n_raters_hidden,
    worst_cat = worst_cat, worst_val = worst_val,
    best_cat = best_cat, best_val = best_val, cat_degenerate = cat_degenerate,
    n_unanimous = n_unanimous, n_split = n_split,
    agreement_df = agreement_df, rater_df = rater_df, category_df = category_df,
    item_df = item_df, marginal_df = marginal_df, methods_df = methods_df,
    metrics = metrics, json_output = json_output, boot_B = B
  )
}
Your data has more stories to tell.Run any analysis on your own data , R modules you own and can re-run, interactive reports, AI insights, and PDF export. 14 days of full access when you finish onboarding.
Try Free — No SignupSign Up Free

Your turn

Bring your own data and the question you actually need answered.

CympleData Scientist Send me your data and question, I’ll send you the analytics. ds@mcpanalytics.ai

Cite this analysis

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing