Standard Kappa Agreement
Executive Summary

Executive Summary

Kappa 0.621 (substantial) on 300 items

Items Rated
300
Categories
4
Percent Agreement
72.0%
Cohen's Kappa
0.621
Kappa 95% CI
0.551 to 0.690
Agreement Strength
substantial
Weighted Kappa
0.83
Kappa vs Chance p
< 0.001
Reviewer A Rating and Reviewer B Rating picked the same category on 72.0% of 300 items. Cohen's kappa is 0.621 (95% CI 0.551 to 0.690), which the Landis & Koch convention calls SUBSTANTIAL agreement, and which is clearly better than chance (z = 18.33, p < 0.001). Chance alone would have produced 26.2% agreement given each rater's own habits, leaving 73.8% of the scale available above chance, of which 62.1% was achieved. Because the categories are ordered, near misses can earn partial credit: linear-weighted kappa is 0.732 and quadratic-weighted kappa 0.830. Quote the weighting you used — quadratic weights are the most forgiving of the three and will always give the largest number. The single most common disagreement was Reviewer A Rating choosing "Good" while Reviewer B Rating chose "Fair", on 15 items (5.0% of the total). Agreement was weakest on "Fair", where the raters concurred on 64.90% of the times either of them reached for that label. One caution that applies whatever the number: kappa measures whether two people make the same call, not whether the call is right.
What this means

The short answer

The two reviewers agreed on 72.0% of 300 items, and Cohen's kappa is 0.621 (95% CI 0.551 to 0.690)—substantially better than chance (p < 0.001). They are reading consistently rather than just both defaulting to the same category.

The detail

Observed agreement: 72.0%. Chance agreement: 26.2%. Cohen's kappa: 0.621 (z = 18.33, p < 0.001). The Landis & Koch convention calls this "substantial." Because the categories are ordered on a quality scale, weighted kappa is also reported: linear-weighted kappa is 0.732 and quadratic-weighted kappa is 0.830. The single most common disagreement was Reviewer A choosing "Good" while Reviewer B chose "Fair" (15 items, 5.0% of total). Agreement was weakest on "Fair" (64.90% of items where either rater used that label). Kappa measures agreement, not correctness.

What this can't tell you

The confidence interval assumes large-sample behaviour on a matrix with 6 thin cells; treat the width as indicative. Kappa does not establish whether either reviewer is accurate—that requires comparison to a gold standard.

Overview

Analysis Overview

Cohen's kappa on 300 items judged by Reviewer A Rating and Reviewer B Rating across 4 categories.

N Items300
N Categories4
Percent Agreement72
Kappa0.621
What this means

The short answer

Two raters picking the same category on 72% of items looks strong, but raw agreement alone can't tell whether they're reading carefully or just both leaning on the same popular category by chance. Cohen's kappa corrects for that baseline and shows 0.621—substantially better than chance—meaning the raters achieved 62.1% of the agreement possible above what their own habits would produce anyway.

The detail

The 300 items were judged across 4 categories. Raw percent agreement is 72.0%; chance agreement is 26.2% (computed from each rater's marginal distribution); Cohen's kappa is 0.621 (95% CI 0.551 to 0.690). The significance test yields z = 18.33, p < 0.001, confirming the agreement is far better than chance. The Landis & Koch convention labels this "substantial." The confidence interval rests on a large-sample approximation applied to a matrix with 6 thin cells (fewer than 5 items each), so the width is indicative rather than precise.

What this can't tell you

Kappa measures whether the two reviewers make the same call, not whether either call is correct. A gold standard comparison would be needed to assess accuracy. The interval's precision is limited by sparse cells in the confusion matrix; a larger sample would sharpen the bounds.

Data Preparation

Data Quality

300 items used of 300 rows loaded.

Initial Rows300
Final Rows300
Rows Removed0
Incomplete Dropped0
Case Variants Merged0
What this means

The short answer

All 300 rows loaded carried complete judgements from both reviewers and were retained for analysis. No rows were dropped, no labels merged, and no categories pooled. The data entered analysis intact.

The detail

300 rows were loaded and 300 items were analysed. Rows removed: 0. Incomplete pairs dropped: 0. Case variants merged: 0. Labels were trimmed of surrounding spaces before comparison; blanks and placeholders were treated as missing rather than as a category. Each row represents one item with one judgement from Reviewer A Rating and one from Reviewer B Rating.

What this can't tell you

Six of the 16 cells in the 4-by-4 confusion matrix hold fewer than 5 items, making the confidence interval a large-sample approximation applied to sparse data. Consider whether a finer-grained or extended review cycle could populate the rare cell combinations, which would firm up the interval bounds.

Visualization

Where the Two Raters Land

Every combination of Reviewer A Rating and Reviewer B Rating's choices, counted.

What this means

The short answer

The diagonal (agreement) holds 72.0% of the 300 items. The brightest off-diagonal cell is "Good" vs. "Fair" (15 items, 5.0%), indicating that pair of category definitions needs clarification. Disagreements are spread across the scale rather than concentrated in one category.

The detail

The 4-by-4 matrix shows every combination of Reviewer A and Reviewer B choices. Diagonal cells (agreement): Poor–Poor 40 items (13.33%), Fair–Fair 50 items (16.67%), Good–Good 70 items (23.33%), Excellent–Excellent 56 items (18.67%), totalling 216 items (72.0%). The largest off-diagonal cell is Reviewer A = "Good" vs. Reviewer B = "Fair" with 15 items (5.0%). The categories are ordered on a quality scale, so near-diagonal disagreements are near misses; far-diagonal cells are serious ones. Asymmetry check: "Poor" shows the largest marginal gap at 1.0% (Reviewer A 17.3% vs. Reviewer B 18.3%), indicating minimal rater bias.

What this can't tell you

The heatmap reveals which category pairs collide most often but not why. Clarifying the "Good"/"Fair" boundary would likely improve agreement; consider reviewing the rubric entries for those two categories with both reviewers.

Data Table

Agreement Statistics

Percent agreement, chance agreement, kappa with its interval, and the weighting variants.

StatisticEstimateCI LowCI HighInterpretation
Raw percent agreement0.72The two raters chose the same category for 72.0% of the 300 items.
Agreement expected by chance0.2621What two raters with these same category habits would agree on by luck alone (26.2%).
Cohen's kappa0.62060.55110.6962.1% of the agreement left available above chance was achieved — substantial on the Landis & Koch convention.
Weighted kappa (linear)0.73160.67760.7856Near-miss disagreements between adjacent categories count as partial credit, in proportion to how far apart they are.
Weighted kappa (quadratic)0.830.78670.8734The same idea with distance squared: near misses are forgiven much more, far misses punished much harder. Always the most flattering of the three.
Prevalence-and-bias-adjusted kappa0.6267A diagnostic, not a replacement: what kappa would be if every category were equally common and both raters used them equally often. A large gap from Cohen's kappa means skewed marginals are doing the work.
What this means

The short answer

Kappa of 0.621 means the raters achieved 62.1% of the agreement possible above what their category habits alone would produce. The 95% confidence interval (0.551 to 0.690) confirms the pattern is real, though the interval width reflects sparse cells in the matrix.

The detail

Raw percent agreement: 72.0%. Agreement expected by chance: 26.2%. Cohen's kappa: 0.621 (CI 0.551 to 0.690). The standard error is 0.035 (Fleiss-Cohen-Everitt). The prevalence-and-bias-adjusted kappa is 0.627, within 0.01 of Cohen's kappa, so the category mix is not distorting the headline. Weighted kappa (linear): 0.732 (CI 0.678 to 0.786); weighted kappa (quadratic): 0.830 (CI 0.787 to 0.873). The confidence interval is a large-sample approximation; 6 of 16 cells hold fewer than 5 items, so read the bounds as indicative.

What this can't tell you

Whether 0.621 is "good enough" depends on the use case—the Landis & Koch labels are convention only. Sparse cells limit confidence interval precision; a larger sample would tighten the bounds. The analysis does not address whether either rater is correct.

Visualization

Which Categories They Fight Over

Per-category agreement between Reviewer A Rating and Reviewer B Rating.

What this means

The short answer

"Excellent" shows the strongest per-category agreement at 80.00%, and "Fair" the weakest at 64.90%. The 15.10 percentage-point spread across categories is modest, so disagreement is not concentrated in one rubric entry; the whole rubric needs refinement.

The detail

Per-category agreement (of all uses of each category by either rater, the share both used it on the same item): Poor 74.8%, Fair 64.9%, Good 70.4%, Excellent 80.0%. The range is 15.10 percentage points. "Excellent" is the most reliable; "Fair" lags furthest behind. No single category is carrying the disagreement.

What this can't tell you

Per-category agreement hides rater bias—a category one rater uses freely and the other avoids will score low here even when overall kappa is healthy. Consider whether "Fair" needs sharper boundaries or whether both raters need recalibration on the entire scale. A category-specific analysis with a larger sample would clarify whether the "Fair" weakness is systematic or noise.

Visualization

How Each Rater Uses the Scale

Category shares for Reviewer A Rating and Reviewer B Rating side by side.

What this means

The short answer

Both reviewers use the four categories in nearly identical proportions. "Good" is most frequent (33.2% and 33.0%), and the largest difference between raters is 1.0% on "Poor." This balanced distribution keeps the chance-agreement baseline low, allowing kappa to stay close to what the raw agreement of 72.0% suggests.

The detail

Reviewer A: Poor 17.3%, Fair 25.7%, Good 33.3%, Excellent 23.7%. Reviewer B: Poor 18.3%, Fair 25.7%, Good 33.0%, Excellent 23.0%. The largest difference is "Poor" at 1.0 percentage points (Reviewer A 17.3% vs. Reviewer B 18.3%). No category shows systematic bias—one rater reaching for it far more than the other. The chance-agreement calculation depends on these marginals: 26.2% is the sum of (Reviewer A share × Reviewer B share) across all four categories.

What this can't tell you

Balanced marginals alone do not guarantee high agreement on individual items. Raters can use categories in similar proportions yet disagree on which items belong in which category. The marginals confirm there is no systematic rater bias, but item-level calibration differences could still exist; consider a focused rubric review on the "Good"/"Fair" boundary, where the largest off-diagonal disagreement occurs.

Data Table

Methods & Disclosure

Every formula behind the numbers, and what they cannot decide.

ItemDetail
Design300 items each judged once by Reviewer A Rating and once by Reviewer B Rating, across 4 categories.
Percent agreementDiagonal of the confusion matrix over the total: 72.0%.
Chance agreementSum over categories of (Reviewer A Rating's share) x (Reviewer B Rating's share): 26.2%.
Cohen's kappa(observed - expected) / (1 - expected) = (0.720 - 0.262) / (1 - 0.262) = 0.621.
Standard error and CIFleiss-Cohen-Everitt asymptotic standard error, 0.035; the 95% interval is the estimate plus or minus 1.96 standard errors. It is a large-sample approximation, and 6 of the 16 matrix cells hold fewer than 5 items, so treat the interval as indicative.
Significance testz = kappa / SE under the null of chance-level agreement = 18.33, p < 0.001. The null variance is a different quantity from the one behind the confidence interval.
Category orderOrderedness was detected, not assumed: the labels are all points on a quality scale, so they were ordered along it.
WeightingWeighted kappa is reported because the categories are ordered — the labels are all points on a quality scale, so they were ordered along it.
Benchmark labelsLandis & Koch (1977) labels (slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.
What kappa is notKappa measures agreement, not accuracy: two raters can agree completely and both be wrong. It also describes these two raters on these items, and does not generalise to raters or items outside this set.
What this means

The short answer

All calculations are closed-form arithmetic on the 4-by-4 confusion matrix—no model fitting or simulation. The module detected that the categories are ordered on a quality scale, which enables weighted kappa. Kappa measures agreement between these two raters on these items and says nothing about whether either is correct.

The detail

Design: 300 items, each judged once by Reviewer A and once by Reviewer B, across 4 categories. Percent agreement: diagonal of matrix / total = 72.0%. Chance agreement: sum of (Reviewer A share × Reviewer B share) = 26.2%. Cohen's kappa: (observed − expected) / (1 − expected) = (0.720 − 0.262) / (1 − 0.262) = 0.621. Standard error (Fleiss-Cohen-Everitt): 0.035; 95% CI is estimate ± 1.96 SE. Significance test: z = kappa / SE = 18.33, p < 0.001 (null: chance-level agreement). Orderedness was detected from labels (all points on a quality scale); this fires weighted kappa. Weighting: linear (0.732) and quadratic (0.830) variants are reported. Benchmark labels (Landis & Koch 1977) are convention only.

What this can't tell you

Kappa describes agreement between these two raters on these items; it does not generalise to other raters or items. Kappa measures agreement, not accuracy—two raters can agree completely and both be wrong. A gold standard comparison would be needed to assess correctness. Six matrix cells are sparse (< 5 items), so the confidence interval is indicative; a larger sample would sharpen the bounds.

Rate this report Was this the answer you needed?
The exact source that produced this report — yours to keep, read, and re-run.
Download PDF
How this was computed method · R source · citation
The code that did it

Rater Agreement — Cohen's Kappa

Two raters, one categorical judgement per item: how much do they agree, and how much of that agreement is more than chance would have produced? The analysis builds the full confusion matrix, reports raw percent agreement alongside Cohen's kappa with its standard error and 95% confidence interval, adds linear- and quadratic-weighted kappa when the categories turn out to be ORDERED (detected from the level names, never assumed), breaks agreement down category by category, and explains the kappa paradox with this dataset's own numbers whenever skewed marginals are depressing kappa below what the raw agreement suggests.

Why This Method?

Percent agreement on its own is not interpretable: two raters who both answer "Yes" to almost everything will agree 90% of the time without looking at a single item. Kappa subtracts the agreement chance alone would deliver given each rater's own habits, and expresses what is left as a fraction of what was achievable.

What This Analysis Covers

  • The k x k confusion matrix as a heatmap, plus the largest disagreement
  • Raw percent agreement, chance agreement, and Cohen's kappa with SE + CI
  • A z-test of kappa against zero (no better than chance)
  • Linear- and quadratic-weighted kappa when the categories are ordered
  • Per-category agreement — which categories the raters actually fight over
  • Each rater's marginal distribution, and the kappa-paradox explanation

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {rater_1, rater_2}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

Rule 1 — the labels are numbers (1, 2, 3 / 0-10 scales).

num <- suppressWarnings(as.numeric(clean))
  if (!anyNA(num) && length(unique(num)) == length(num)) {
    return(list(ordered = TRUE, order = levs[order(num)],
                rule = "the labels are numbers, so they were ordered by value"))
  }

Rule 2 — the labels start with a number ("1 - Poor", "2 = Fair").

lead <- suppressWarnings(as.numeric(sub("^\\s*([-+]?[0-9]+(\\.[0-9]+)?)\\s*[-=.):|].*$",
                                          "\\1", clean)))
  has_lead <- grepl("^\\s*[-+]?[0-9]+(\\.[0-9]+)?\\s*[-=.):|]", clean)
  if (all(has_lead) && !anyNA(lead) && length(unique(lead)) == length(lead)) {
    return(list(ordered = TRUE, order = levs[order(lead)],
                rule = "every label starts with a number, so they were ordered by that number"))
  }

Rule 3 — the labels are all drawn from one recognised ordinal vocabulary.

scales <- list(
    "an agreement scale" = c("strongly disagree", "disagree", "somewhat disagree",
                             "slightly disagree", "neutral", "neither agree nor disagree",
                             "slightly agree", "somewhat agree", "agree", "strongly agree"),
    "a quality scale" = c("very poor", "poor", "below average", "fair", "average",
                          "satisfactory", "good", "very good", "excellent", "outstanding"),
    "a frequency scale" = c("never", "rarely", "seldom", "sometimes", "occasionally",
                            "often", "frequently", "usually", "always"),
    "a severity scale" = c("none", "minimal", "mild", "moderate", "severe",
                           "very severe", "extreme", "critical"),
    "a magnitude scale" = c("very low", "low", "medium", "moderate", "high", "very high"),
    "a satisfaction scale" = c("very dissatisfied", "dissatisfied", "neutral",
                               "satisfied", "very satisfied"),
    "a likelihood scale" = c("very unlikely", "unlikely", "possible", "likely",
                             "very likely", "certain"),
    "a priority scale" = c("trivial", "minor", "moderate", "major", "critical", "blocker")
  )
  low <- tolower(clean)
  if (length(levs) >= 3 && !any(duplicated(low))) {
    for (nm in names(scales)) {
      sc <- scales[[nm]]
      if (all(low %in% sc)) {
        return(list(ordered = TRUE, order = levs[order(match(low, sc))],
                    rule = paste0("the labels are all points on ", nm,
                                  ", so they were ordered along it")))
      }
    }
  }

  list(ordered = FALSE, order = levs,
       rule = "no numeric values, numeric prefixes, or recognised ordinal wording were found in the labels")
}

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed  $ row accounting (items)
  #   r1_h / r2_h              $ humanized names of the two mapped columns
  #   n_items / k_cats         $ usable items, categories analysed
  #   levs                     $ character — categories in analysis order
  #   n_dropped / n_case_merged / n_lumped   $ cleaning counts
  #   case_example             $ character(2) — the merged spelling pair, or NULL
  #   ordered_flag / order_rule              $ orderedness verdict + why
  #   p_o / p_e / kappa / kappa_se / ci_low / ci_high / z_stat / p_value
  #   kappa_linear / kappa_quadratic (+ _se) $ NA when categories are nominal
  #   band                     $ Landis & Koch label for the unweighted kappa
  #   pabak                    $ prevalence/bias-adjusted kappa (diagnostic)
  #   paradox                  $ logical — skewed marginals are depressing kappa
  #   max_prev / prev_cat / max_bias / bias_cat
  #   confusion_df / kappa_df / category_df / marginal_df / methods_df
  #   top_off_*                $ the single largest disagreement cell
  #   n_sparse                 $ confusion cells holding fewer than 5 items
  #   metrics / json_output
  # === /SHARED EXPORTS ===

  r1_h <- humanize_semantic("rater_1", col_map)
  r2_h <- humanize_semantic("rater_2", col_map)

Step 1: Check the mapped columns

initial_rows <- nrow(df)
  for (req in c("rater_1", "rater_2")) {
    if (is.null(df[[req]])) {
      stop(sprintf("column_mapping must map both judgement columns(%s and %s) — one column per rater, one row per item.",
                   r1_h, r2_h))
    }
  }

Step 2: Normalise the two judgement columns to clean category labels.

Categories are text by definition; numbers are accepted and kept as labels (a 1-5 scale is a set of five categories, not a measurement).

as_label <- function(v) {
    s <- trimws(as.character(v))
    s[is.na(s)] <- ""
    s[tolower(s) %in% c("", "na", "n/a", "null", "none given", "-", "--", "?")] <- ""
    s
  }
  a <- as_label(df$rater_1)
  b <- as_label(df$rater_2)

  keep <- a != "" & b != ""
  n_dropped <- sum(!keep)
  a <- a[keep]; b <- b[keep]

  n_items <- length(a)
  if (n_items < 20) {
    stop(sprintf("Only %d item(s) have a judgement from both %s and %s — kappa needs at least 20 to be worth reporting.",
                 n_items, r1_h, r2_h))
  }

Step 3: Merge labels that differ only in capitalisation, and say so.

all_lab <- c(a, b)
  by_case <- split(all_lab, tolower(all_lab))
  canon <- list()
  n_case_merged <- 0L
  case_example <- NULL
  for (key in names(by_case)) {
    spellings <- by_case[[key]]
    uq <- unique(spellings)
    tab <- sort(table(spellings), decreasing = TRUE)
    winner <- names(tab)[1]
    canon[[key]] <- winner
    if (length(uq) > 1) {
      n_case_merged <- n_case_merged + sum(spellings != winner)
      if (is.null(case_example)) {
        case_example <- c(setdiff(uq, winner)[1], winner)
      }
    }
  }
  a <- unlist(canon[tolower(a)], use.names = FALSE)
  b <- unlist(canon[tolower(b)], use.names = FALSE)

Step 4: Reject identifier-like columns; lump a long tail into "Other".

all_lab <- c(a, b)
  freq <- sort(table(all_lab), decreasing = TRUE)
  n_lumped <- 0L
  if (length(freq) > 25) {
    stop(sprintf("%s and %s together hold %d distinct values — that looks like free text or an identifier, not a set of categories. Map the two columns holding each rater&#x27;s category choice.",
                 r1_h, r2_h, length(freq)))
  }
  if (length(freq) > 12) {
    keep_lab <- names(freq)[1:11]
    n_lumped <- length(freq) - 11L
    a[!(a %in% keep_lab)] <- "Other"
    b[!(b %in% keep_lab)] <- "Other"
  }

Step 5: Refuse the degenerate cases with a message naming the column.

ua <- unique(a); ub <- unique(b)
  if (length(ua) < 2) {
    stop(sprintf("%s gave the same answer(\"%s\") to every item, so there is nothing for kappa to measure — agreement beyond chance is undefined when one rater never varies.",
                 r1_h, ua[1]))
  }
  if (length(ub) < 2) {
    stop(sprintf("%s gave the same answer(\"%s\") to every item, so there is nothing for kappa to measure — agreement beyond chance is undefined when one rater never varies.",
                 r2_h, ub[1]))
  }

Step 6: Decide the category order (detected, not assumed)

levs_all <- sort(unique(c(a, b)))
  det <- detect_category_order(levs_all)
  if (det$ordered) {
    levs <- det$order
  } else {

Nominal: order by how much of the data each category carries, so the heatmap reads busiest-first rather than alphabetically.

tot <- sapply(levs_all, function(l) sum(a == l) + sum(b == l))
    levs <- levs_all[order(-tot, levs_all)]
  }
  k <- length(levs)

  fa <- factor(a, levels = levs)
  fb <- factor(b, levels = levs)
  tab <- table(fa, fb)
  N <- matrix(as.numeric(tab), nrow = k, ncol = k,
              dimnames = list(levs, levs))
  n <- sum(N)
  p <- N / n
  rowm <- rowSums(p)
  colm <- colSums(p)

  final_rows <- n_items
  rows_removed <- initial_rows - final_rows

Step 7: Cohen's kappa, its standard error, CI, and a z-test vs chance

I <- diag(1, k)
  base <- kappa_general(p, I, n)
  p_o <- base$po; p_e <- base$pe
  kappa <- base$kappa; kappa_se <- base$se
  ci_low  <- if (is.na(kappa) || is.na(kappa_se)) NA_real_ else kappa - 1.96 * kappa_se
  ci_high <- if (is.na(kappa) || is.na(kappa_se)) NA_real_ else kappa + 1.96 * kappa_se

Variance under the null (kappa = 0) is a different quantity from the variance used for the CI — this is the one the significance test needs.

var0 <- (p_e + p_e^2 - sum(rowm * colm * (rowm + colm))) / (n * (1 - p_e)^2)
  se0 <- if (is.finite(var0) && var0 > 0) sqrt(var0) else NA_real_
  z_stat <- if (is.na(kappa) || is.na(se0)) NA_real_ else kappa / se0
  p_value <- if (is.na(z_stat)) NA_real_ else 2 * pnorm(-abs(z_stat))

Step 8: Weighted kappa — only when the categories are actually ordered

kappa_linear <- NA_real_; kappa_linear_se <- NA_real_
  kappa_quadratic <- NA_real_; kappa_quadratic_se <- NA_real_
  weight_note <- ""
  if (k < 3) {
    weight_note <- paste0(
      "Weighted kappa is not reported: with only ", k,
      " categories every disagreement is already the maximum possible one, so any weighting scheme collapses back to the unweighted value.")
  } else if (!det$ordered) {
    weight_note <- paste0(
      "Weighted kappa is not reported, because the categories do not appear to be ordered — ",
      det$rule,
      ". Weighting assumes some disagreements are worse than others, which is only meaningful on a scale that has a direction. If your categories ARE ordered, rename them so the order is visible(for example \"1 - \", \"2 - \", \"3 - \" prefixes) and re-run.")
  } else {
    idx <- seq_len(k)
    D <- abs(outer(idx, idx, "-")) / (k - 1)
    Wl <- 1 - D
    Wq <- 1 - D^2
    kl <- kappa_general(p, Wl, n)
    kq <- kappa_general(p, Wq, n)
    kappa_linear <- kl$kappa; kappa_linear_se <- kl$se
    kappa_quadratic <- kq$kappa; kappa_quadratic_se <- kq$se
    weight_note <- paste0(
      "Weighted kappa is reported because the categories are ordered — ",
      det$rule, ".")
  }

  band <- landis_koch(kappa)

Step 9: Prevalence and bias — the two things that move kappa without

either rater changing how well they agree.

avg_marg <- (rowm + colm) / 2
  prev_i <- which.max(avg_marg)
  max_prev <- avg_marg[prev_i]
  prev_cat <- levs[prev_i]
  bias_vec <- abs(rowm - colm)
  bias_i <- which.max(bias_vec)
  max_bias <- bias_vec[bias_i]
  bias_cat <- levs[bias_i]
  pabak <- (k * p_o - 1) / (k - 1)
  paradox <- isTRUE(p_o >= 0.70 && !is.na(kappa) && kappa < 0.60 && max_prev >= 0.60)

Step 10: Confusion matrix (long form for the heatmap)

confusion_df <- data.frame(
    rater_1_category = rep(levs, times = k),
    rater_2_category = rep(levs, each = k),
    n_items = as.numeric(N[cbind(rep(seq_len(k), times = k),
                                 rep(seq_len(k), each = k))]),
    stringsAsFactors = FALSE
  )
  confusion_df$share_pct <- round(100 * confusion_df$n_items / n, 2)
  n_sparse <- sum(N < 5)

Largest off-diagonal cell, found NA-safely and only among real cells.

off <- N; diag(off) <- NA_real_
  off_ok <- which(!is.na(off) & off > 0)
  if (length(off_ok) > 0) {
    best <- off_ok[which.max(off[off_ok])]
    bi <- ((best - 1) %% k) + 1
    bj <- ((best - 1) %/% k) + 1
    top_off_n <- off[best]
    top_off_r1 <- levs[bi]
    top_off_r2 <- levs[bj]
  } else {
    top_off_n <- 0; top_off_r1 <- NA_character_; top_off_r2 <- NA_character_
  }

Step 11: Per-category agreement (proportion of specific agreement:

twice the agreed items over the times either rater used the category)

denom <- rowSums(N) + colSums(N)
  spec <- ifelse(denom > 0, 200 * diag(N) / denom, NA_real_)
  category_df <- data.frame(
    category = levs,
    agreement_pct = round(spec, 1),
    stringsAsFactors = FALSE
  )
  ok_cat <- which(!is.na(category_df$agreement_pct))
  worst_cat <- if (length(ok_cat) > 0)
    category_df$category[ok_cat[which.min(category_df$agreement_pct[ok_cat])]] else NA_character_
  worst_val <- if (length(ok_cat) > 0)
    min(category_df$agreement_pct[ok_cat], na.rm = TRUE) else NA_real_
  best_cat <- if (length(ok_cat) > 0)
    category_df$category[ok_cat[which.max(category_df$agreement_pct[ok_cat])]] else NA_character_
  best_val <- if (length(ok_cat) > 0)
    max(category_df$agreement_pct[ok_cat], na.rm = TRUE) else NA_real_

Step 12: Each rater's marginal distribution (the paradox's evidence)

marginal_df <- data.frame(
    category = rep(levs, times = 2),
    rater = c(rep(r1_h, k), rep(r2_h, k)),
    share_pct = round(100 * c(rowm, colm), 1),
    stringsAsFactors = FALSE
  )

Step 13: The kappa results table

kappa_rows <- list(
    list("Raw percent agreement", p_o, NA_real_, NA_real_,
         sprintf("The two raters chose the same category for %s of the %s items.",
                 pct1(p_o), format(n, big.mark = ","))),
    list("Agreement expected by chance", p_e, NA_real_, NA_real_,
         sprintf("What two raters with these same category habits would agree on by luck alone(%s).",
                 pct1(p_e))),
    list("Cohen&#x27;s kappa", kappa, ci_low, ci_high,
         sprintf("%s of the agreement left available above chance was achieved — %s on the Landis & Koch convention.",
                 pct1(kappa), band))
  )
  if (!is.na(kappa_linear)) {
    kappa_rows[[length(kappa_rows) + 1]] <- list(
      "Weighted kappa(linear)", kappa_linear,
      kappa_linear - 1.96 * kappa_linear_se, kappa_linear + 1.96 * kappa_linear_se,
      "Near-miss disagreements between adjacent categories count as partial credit, in proportion to how far apart they are.")
    kappa_rows[[length(kappa_rows) + 1]] <- list(
      "Weighted kappa(quadratic)", kappa_quadratic,
      kappa_quadratic - 1.96 * kappa_quadratic_se, kappa_quadratic + 1.96 * kappa_quadratic_se,
      "The same idea with distance squared: near misses are forgiven much more, far misses punished much harder. Always the most flattering of the three.")
  }
  kappa_rows[[length(kappa_rows) + 1]] <- list(
    "Prevalence-and-bias-adjusted kappa", pabak, NA_real_, NA_real_,
    "A diagnostic, not a replacement: what kappa would be if every category were equally common and both raters used them equally often. A large gap from Cohen&#x27;s kappa means skewed marginals are doing the work.")

  kappa_df <- data.frame(
    statistic = sapply(kappa_rows, function(r) r[[1]]),
    estimate = round(sapply(kappa_rows, function(r) as.numeric(r[[2]])), 4),
    ci_low = round(sapply(kappa_rows, function(r) as.numeric(r[[3]])), 4),
    ci_high = round(sapply(kappa_rows, function(r) as.numeric(r[[4]])), 4),
    interpretation = sapply(kappa_rows, function(r) r[[5]]),
    stringsAsFactors = FALSE
  )

Step 14: Methods disclosure

methods_df <- data.frame(
    item = c(
      "Design",
      "Percent agreement",
      "Chance agreement",
      "Cohen&#x27;s kappa",
      "Standard error and CI",
      "Significance test",
      "Category order",
      "Weighting",
      "Benchmark labels",
      "What kappa is not"
    ),
    detail = c(
      sprintf("%s items each judged once by %s and once by %s, across %d categories.",
              format(n, big.mark = ","), r1_h, r2_h, k),
      sprintf("Diagonal of the confusion matrix over the total: %s.", pct1(p_o)),
      sprintf("Sum over categories of(%s&#x27;s share) x (%s's share): %s.",
              r1_h, r2_h, pct1(p_e)),
      sprintf("(observed - expected) / (1 - expected) = (%s - %s) / (1 - %s) = %s.",
              r3(p_o), r3(p_e), r3(p_e), r3(kappa)),
      sprintf("Fleiss-Cohen-Everitt asymptotic standard error, %s; the 95%% interval is the estimate plus or minus 1.96 standard errors. It is a large-sample approximation%s.",
              r3(kappa_se),
              if (n_sparse > 0) sprintf(", and %d of the %d matrix cells hold fewer than 5 items, so treat the interval as indicative", n_sparse, k * k) else ""),
      sprintf("z = kappa / SE under the null of chance-level agreement = %s, %s. The null variance is a different quantity from the one behind the confidence interval.",
              r2(z_stat), fmt_pp(p_value)),
      sprintf("Orderedness was detected, not assumed: %s.", det$rule),
      weight_note,
      "Landis & Koch(1977) labels(slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.",
      "Kappa measures agreement, not accuracy: two raters can agree completely and both be wrong. It also describes these two raters on these items, and does not generalise to raters or items outside this set."
    ),
    stringsAsFactors = FALSE
  )

Step 15: Headline metrics + the computed one-paragraph answer

metrics <- list(
    `Items Rated`        = n,
    `Categories`         = k,
    `Percent Agreement`  = pct1(p_o),
    `Cohen&#x27;s Kappa`      = round(kappa, 3),
    `Kappa 95% CI`       = paste0(r3(ci_low), " to ", r3(ci_high)),
    `Agreement Strength` = band,
    `Weighted Kappa`     = if (is.na(kappa_quadratic)) "not applicable" else round(kappa_quadratic, 3),
    `Kappa vs Chance p`  = fmt_p(p_value)
  )

  paradox_sentence <- if (paradox) {
    paste0(
      " Read those two numbers together before quoting either: raw agreement is high(", pct1(p_o),
      ") while kappa is only ", r3(kappa), ", which is the well-known kappa paradox rather than a contradiction. ",
      pct1(max_prev), " of all judgements fell into the single category \"", prev_cat,
      "\", so two raters with these habits would already have agreed on ", pct1(p_e),
      " of items by chance; only ", pct1(1 - p_e), " of the scale was left for skill to win, and ",
      pct1(kappa), " of that remainder was won.")
  } else {
    paste0(
      " Chance alone would have produced ", pct1(p_e), " agreement given each rater&#x27;s own habits, leaving ",
      pct1(1 - p_e), " of the scale available above chance, of which ", pct1(kappa), " was achieved.")
  }

  weighted_sentence <- if (!is.na(kappa_linear)) {
    paste0(" The categories are ordered(", det$rule,
           "), so near misses can be given partial credit: linear-weighted kappa is ", r3(kappa_linear),
           " and quadratic-weighted kappa ", r3(kappa_quadratic),
           " — both higher than the unweighted value because most disagreements are between neighbouring categories.")
  } else {
    paste0(" ", weight_note)
  }

  json_output <- list(
    answer = paste0(
      r1_h, " and ", r2_h, " agreed on ", pct1(p_o), " of ", format(n, big.mark = ","),
      " items across ", k, " categories. Cohen&#x27;s kappa is ", r3(kappa),
      " (95% CI ", r3(ci_low), " to ", r3(ci_high), "), ", band,
      " agreement on the Landis & Koch convention, and ",
      if (!is.na(p_value) && p_value < 0.05)
        paste0("clearly better than chance(", fmt_pp(p_value), ")")
      else
        paste0("not distinguishable from chance(", fmt_pp(p_value), ")"),
      ".", paradox_sentence, weighted_sentence,
      " Agreement was weakest on \"", worst_cat, "\" (", r2(worst_val),
      "% of the times either rater used it) and strongest on \"", best_cat, "\" (",
      r2(best_val), "%). Kappa measures agreement, not correctness."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "confusion_heatmap", "kappa_results",
        "category_agreement", "rater_marginals", "methods"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed, n_dropped = n_dropped,
    n_case_merged = n_case_merged, case_example = case_example,
    n_lumped = n_lumped,
    r1_h = r1_h, r2_h = r2_h,
    n_items = n, k_cats = k, levs = levs,
    ordered_flag = det$ordered, order_rule = det$rule,
    p_o = p_o, p_e = p_e, kappa = kappa, kappa_se = kappa_se,
    ci_low = ci_low, ci_high = ci_high,
    z_stat = z_stat, p_value = p_value,
    kappa_linear = kappa_linear, kappa_linear_se = kappa_linear_se,
    kappa_quadratic = kappa_quadratic, kappa_quadratic_se = kappa_quadratic_se,
    weight_note = weight_note, band = band, pabak = pabak,
    paradox = paradox, paradox_sentence = paradox_sentence,
    max_prev = max_prev, prev_cat = prev_cat,
    max_bias = max_bias, bias_cat = bias_cat,
    n_sparse = n_sparse,
    top_off_n = top_off_n, top_off_r1 = top_off_r1, top_off_r2 = top_off_r2,
    worst_cat = worst_cat, worst_val = worst_val,
    best_cat = best_cat, best_val = best_val,
    confusion_df = confusion_df, kappa_df = kappa_df,
    category_df = category_df, marginal_df = marginal_df,
    methods_df = methods_df,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell.Run any analysis on your own data: R modules you own, interactive reports, AI insights, and PDF export. 500 free credits when you finish onboarding.
Try Free — No SignupSign Up Free

Your turn

Bring your own data and the question you actually need answered.

CympleData Scientist Send me your data and question, I’ll send you the analytics. ds@mcpanalytics.ai

Cite this analysis

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing