Executive Summary
Fleiss' kappa 0.670 (substantial) across 5 raters on 9,986 items
Five raters achieved Fleiss' kappa of 0.67 (95% CI 0.664–0.677) on 9,986 items, substantially better than the 33.3% chance agreement their category frequencies would produce. Krippendorff's alpha confirms 0.67. Annotator 1 is most aligned (rater kappa 0.718); annotator 2 is least aligned (rater kappa 0.651), but removing annotator 2 raises panel kappa only to 0.683, indicating the disagreement is systemic rather than person-driven. Of 9,986 items, 5,479 (54.9%) drew unanimous verdicts; 155 (1.6%) produced no majority. Agreement measures consistency, not correctness—a panel can agree unanimously and be wrong.
Analysis Overview
Fleiss' kappa and Krippendorff's alpha on 9,986 items judged by 5 raters across 3 categories.
Five raters judged 9,986 sentence pairs across three categories (entailment, neutral, contradiction). Raw pairwise agreement reached 78.0%, but this inflates when raters lean on the same category by chance alone. Fleiss' kappa (0.67) removes that floor by computing what the panel's category habits would produce randomly (33.3%) and reporting the share of remaining headroom won. Krippendorff's alpha (0.67) asks the same question through disagreement instead and tolerates missing ratings. The two coefficients nearly coincide here because data is complete. With exactly two raters, Cohen's kappa applies; with continuous measurements, an intraclass correlation replaces kappa entirely because kappa discards the scale.
Data Quality
9,986 items and 49,930 judgements from 9,986 rows loaded.
The short answer
The input file was in wide form (one row per item, one column per rater) and was correctly detected from the data: each of the 5 rater columns held only 3 distinct labels instead of one value per row. No ratings were missing—all 5 raters judged all 9,986 items, yielding 49,930 judgements. Both Fleiss' kappa and Krippendorff's alpha were computed on identical complete data, so they should (and do) converge.
The detail
9,986 rows loaded, resolving to 9,986 items, 5 raters, and 49,930 judgements across 3 categories. The file was recognised as wide form because the 5 mapped columns showed 3, 3, 3, 3, 3 distinct values respectively—the hallmark of one column per rater. The panel resolved to annotator 1, annotator 2, annotator 3, annotator 4, annotator 5. No ratings are missing; every rater judged every item. The 95% confidence interval around kappa runs 0.664 to 0.677, a width of 0.014, anchored to 9,986 items on the analysis basis. Labels were trimmed of surrounding spaces, and blanks and standard placeholders were treated as no judgement rather than as a category.
What this can't tell you
Complete data eliminates the need to choose between coefficients on the basis of missing patterns. If ratings had been sparse, Krippendorff's alpha would have used every item with two or more judgements while Fleiss' kappa would have dropped incomplete items—a distinction that matters when missingness is not random.
Agreement Statistics
Pairwise agreement, chance agreement, Fleiss' kappa and Krippendorff's alpha with bootstrap intervals.
| Statistic | Estimate | CI Low | CI High | Interpretation |
|---|---|---|---|---|
| Raw pairwise agreement | 0.78 | — | — | Across every pair of raters who judged the same item, 78.0% of those pairs chose the same category. |
| Agreement expected by chance | 0.3334 | — | — | What a panel using the categories this often would have agreed on by luck alone (33.3%). |
| Fleiss' kappa | 0.67 | 0.6635 | 0.6774 | 67.0% of the agreement left available above chance was achieved, on the 9,986 item(s) every one of the 5 raters judged — substantial on the Landis & Koch convention. |
| Krippendorff's alpha (nominal) | 0.67 | 0.6625 | 0.6778 | The disagreement-based coefficient, computed on all 9,986 items with two or more judgements — it does not require every rater to have judged every item. |
Pairwise agreement of 78.0% against chance agreement of 33.3% leaves 66.7% headroom, of which kappa captures 67.0%. The 95% bootstrap interval (0.6635 to 0.6774) derives from item resampling, not formula—the classical Fleiss standard error applies only to the null hypothesis, not to intervals around non-zero estimates. Fleiss' kappa and nominal Krippendorff's alpha match within 0.000, which is correct when data are complete and identical; they are two arithmetic routes to one question. Ordinal alpha was not computed because category labels carry no numeric order, numeric prefix, or ordinal wording. If categories are ordered, rename them (e.g., "1 - entailment") and re-run. Whether 0.67 meets your standard depends on the decision it serves.
Which Rater Is Out of Step
Each rater's agreement with the rest of the panel, and the panel's kappa without them.
Annotator 1 leads at rater kappa 0.7178 (81.19% raw agreement with the panel); annotator 2 trails at rater kappa 0.6507 (76.72%). The spread between highest and lowest is 0.067, so no single rater dominates the disagreement. Removing annotator 2 would raise panel kappa only to 0.6829, a gain of 0.013—negligible. Removing annotator 1 would lower it to 0.6382, showing that alignment is distributed across the panel, not concentrated in one outlier. A low rater-level score reflects threshold drift (applying the rubric differently) rather than careless error, and can occur whether that rater or the rest of the panel is correct.
Which Categories the Panel Fights Over
Category-specific kappa across 3 categories.
The short answer
Contradiction was the most reliably coded category (kappa 0.759), neutral the least (kappa 0.568), a gap of 0.191. No single category is carrying the disagreement; the spread is modest enough that improvement must come from the rubric as a whole, not from fixing one category definition.
The detail
Category-specific kappa: entailment 0.6826 (33.97% of judgements, 16,959 uses); neutral 0.5684 (33.21%, 16,581 uses); contradiction 0.7594 (32.83%, 16,390 uses). All three categories appear in roughly equal proportions, so none is a rare label thinly evidenced. The 0.191-point spread between contradiction and neutral is the largest single-category gap. A category that one rater reaches for freely and others barely touch scores low here even when overall kappa looks healthy—this detail a single summary coefficient hides.
What this can't tell you
Category-specific kappa does not isolate whether disagreement stems from the category's definition or from how individual raters apply it. To locate the rubric's weakest point, sample items coded as neutral where raters split, and compare their wording or context to items where neutral was unanimous. The equal distribution of judgements across categories means the chance-agreement floor is not inflated by any single dominant class.
How Often the Panel Reached a Verdict
Consensus strength across 9,986 items on the analysis basis.
The short answer
More than half the items (5,479 of 9,986, or 54.9%) drew unanimous agreement. A further 4,352 (43.6%) produced a majority without unanimity. Only 155 items (1.6%) produced no majority at all—the concrete work list where the rubric is silent or the item is genuinely ambiguous.
The detail
Consensus bands: unanimous (every rater agreed) 5,479 items (54.87%); strong majority (more than two-thirds) 2,845 items (28.49%); simple majority (more than half) 1,507 items (15.09%); no majority (raters split) 155 items (1.55%). Bands are set from each item's largest block of agreeing raters, so an evenly split item needs no arbitrary winner. With 5 raters the possible consensus values are coarse (unanimous, 4-of-5, 3-of-5, tie); empty bands in a visualization mean arithmetic simply cannot land there, not that no item did.
What this can't tell you
Concentration of disagreement in a small minority of items does not mean the rubric is clear; it may mean most items are unambiguous while a hard core are genuinely borderline. Re-reading a sample of the 155 no-majority items usually explains the headline kappa faster than the headline number itself does. Consider exporting those 155 items separately to identify patterns in their wording, context, or source.
How Each Rater Uses the Scale
Category shares per rater, side by side.
All five raters distribute their judgements across the three categories with no dramatic skew. Entailment ranges from 33.35% (annotator 1) to 34.34% (annotator 5); neutral from 32.71% (annotator 4) to 33.73% (annotator 2); contradiction from 32.30% (annotator 2) to 33.35% (annotator 1). The even spread is the mechanical reason the chance-agreement floor stays at 33.3%, allowing the 78.0% raw agreement to translate into a kappa of 0.67 rather than being inflated by category bias. A rater whose bars deviate visibly from the group applies a different threshold, not random noise, and depresses agreement even when individual judgements are defensible.
Methods & Disclosure
Every formula behind the numbers, and what they cannot decide.
| Item | Detail |
|---|---|
| Design | 9,986 item(s) judged by 5 rater(s) across 3 categories, 49,930 judgement(s) in total. |
| Input shape | Detected from the data, not from the mapping: the 5 mapped columns were read as one column per rater, because each holds only a handful of repeated labels (3, 3, 3, 3, 3 distinct) rather than one value per row. |
| Raw pairwise agreement | Mean over items of the share of rater PAIRS on that item choosing the same category: 78.0%. |
| Chance agreement | Sum of the squared overall category shares (entailment 34.0%, neutral 33.2%, contradiction 32.8%): 33.3%. |
| Fleiss' kappa | (observed - expected) / (1 - expected) = (0.780 - 0.333) / (1 - 0.333) = 0.670, computed on the 9,986 item(s) every one of the 5 raters judged. |
| Significance test | Fleiss' null-hypothesis standard error is 0.002, giving z = 299.29 and p < 0.001 against the hypothesis that agreement is no better than chance. That variance describes kappa when kappa is truly zero, so it belongs to this test and NOT to the confidence interval. |
| Confidence intervals | A seeded nonparametric bootstrap resampling ITEMS with replacement (500 resamples, fixed seed), taking the 2.5th and 97.5th percentiles. The bootstrap standard error of kappa is 0.004. This is used rather than a closed form because the null variance above is the wrong quantity for an interval around a non-zero estimate, and alpha has no simple closed-form variance. |
| Missing ratings | No ratings are missing: every rater judged every item, so Fleiss' kappa and Krippendorff's alpha are computed on exactly the same data. |
| Krippendorff's alpha | Built from the coincidence matrix: alpha = 1 - observed disagreement / expected disagreement, with (n-1) weighting so it is defined for small samples. Nominal alpha treats every disagreement as equal. Ordinal alpha is not reported for these categories. |
| Category order | Orderedness was detected, not assumed: no numeric values, numeric prefixes, or recognised ordinal wording were found in the labels. |
| Per-category kappa | Fleiss' category-specific coefficient, generalized to a variable number of raters per item: one minus the observed splits on that category over the splits expected from its overall share. |
| Per-rater figures | A rater's agreement figure is the share of the OTHER raters' judgements on the same items that matched this rater's own, chance-corrected against the same 33.3% floor. The leave-one-out column recomputes Fleiss' kappa for the panel with that rater removed. |
| Benchmark labels | Landis & Koch (1977) labels (slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for. |
| What agreement is not | Agreement is not correctness: a panel can agree unanimously and be unanimously wrong. These figures describe this panel on these items and do not generalise to other raters or other items. |
The short answer
All figures except the confidence intervals are closed-form arithmetic on the 9,986-by-3 matrix of rater counts per item per category. The intervals come from resampling items 500 times with a fixed seed, producing reproducible but approximate bounds that widen when item counts are small. Two judgement calls were made: input shape (detected as wide form from the data) and category order (detected as nominal—no numeric values, prefixes, or ordinal wording found in the labels).
The detail
Design: 9,986 items, 5 raters, 3 categories, 49,930 judgements. Raw pairwise agreement: mean share of rater pairs per item choosing the same category = 78.0%. Chance agreement: sum of squared overall category shares (0.340² + 0.332² + 0.328²) = 0.333. Fleiss' kappa: (0.780 − 0.333) / (1 − 0.333) = 0.670, on 9,986 items where all 5 raters judged. Significance: null standard error 0.002, z = 299.29, p < 0.001. Confidence intervals: nonparametric bootstrap resampling items with replacement (500 resamples, seeded), taking 2.5th and 97.5th percentiles; bootstrap standard error 0.004. Krippendorff's alpha: 1 − (observed disagreement / expected disagreement), using (n−1) weighting. Category order: not detected; no ordinal alpha reported.
What this can't tell you
Agreement is not correctness. No coefficient on this page establishes whether the panel is right. The figures describe this panel on these items only; a different rater set or item pool with a different category distribution will produce a different kappa from the same rubric. The Landis & Koch bands are convention without theoretical foundation. Consider whether agreement is uniform across item subgroups (e.g., by source, difficulty, or domain) by requesting a stratified re-run or a finer-grained export.
Multi-Rater Agreement — Fleiss Kappa
Three or more raters, one categorical judgement per item: how much do they agree, how much of that is more than chance would have produced, and which rater is pulling the group apart? The analysis computes Fleiss' kappa with a standard error and confidence interval, Krippendorff's alpha (nominal, and ordinal when the categories turn out to be ordered), a per-category kappa showing which categories the raters fight over, a per-rater agreement with the consensus plus a leave-one-rater-out kappa that names the outlier, and the item-level distribution of how strong the consensus actually was.
Why This Method?
Raw agreement across a panel is not interpretable on its own: a panel that answers "Yes" to almost everything will agree constantly without reading a single item. Fleiss' kappa subtracts the agreement chance alone would deliver given how often each category is used overall. Krippendorff's alpha answers the same question through a different route and, unlike Fleiss', keeps working when not every rater judged every item.
What This Analysis Covers
- Raw pairwise agreement, chance agreement, and Fleiss' kappa with SE + CI
- Krippendorff's alpha, nominal and (when the labels are ordered) ordinal
- Explicit accounting for missing ratings — what each statistic used
- Per-category kappa: which categories the panel cannot pin down
- Per-rater agreement with the consensus and leave-one-rater-out kappa
- The item-level consensus distribution, and the kappa paradox when it fires
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {rater_1 .. rater_N}. Both real-world layouts are accepted and the shape is detected from the data, never assumed. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
Rule 1 — the labels are numbers (1, 2, 3 / 0-10 scales).
num <- suppressWarnings(as.numeric(clean))
if (!anyNA(num) && length(unique(num)) == length(num)) {
return(list(ordered = TRUE, order = levs[order(num)],
rule = "the labels are numbers, so they were ordered by value"))
}Rule 2 — the labels start with a number ("1 - Poor", "2 = Fair").
lead <- suppressWarnings(as.numeric(sub("^\\s*([-+]?[0-9]+(\\.[0-9]+)?)\\s*[-=.):|].*$",
"\\1", clean)))
has_lead <- grepl("^\\s*[-+]?[0-9]+(\\.[0-9]+)?\\s*[-=.):|]", clean)
if (all(has_lead) && !anyNA(lead) && length(unique(lead)) == length(lead)) {
return(list(ordered = TRUE, order = levs[order(lead)],
rule = "every label starts with a number, so they were ordered by that number"))
}Rule 3 — the labels are all drawn from one recognised ordinal vocabulary.
scales <- list(
"an agreement scale" = c("strongly disagree", "disagree", "somewhat disagree",
"slightly disagree", "neutral", "neither agree nor disagree",
"slightly agree", "somewhat agree", "agree", "strongly agree"),
"a quality scale" = c("very poor", "poor", "below average", "fair", "average",
"satisfactory", "good", "very good", "excellent", "outstanding"),
"a frequency scale" = c("never", "rarely", "seldom", "sometimes", "occasionally",
"often", "frequently", "usually", "always"),
"a severity scale" = c("none", "minimal", "mild", "moderate", "severe",
"very severe", "extreme", "critical"),
"a magnitude scale" = c("very low", "low", "medium", "moderate", "high", "very high"),
"a satisfaction scale" = c("very dissatisfied", "dissatisfied", "neutral",
"satisfied", "very satisfied"),
"a likelihood scale" = c("very unlikely", "unlikely", "possible", "likely",
"very likely", "certain"),
"a decision scale" = c("reject", "major revision", "revise", "minor revision",
"accept with changes", "accept"),
"a priority scale" = c("trivial", "minor", "moderate", "major", "critical", "blocker")
)
low <- tolower(clean)
if (length(levs) >= 3 && !any(duplicated(low))) {
for (nm in names(scales)) {
sc <- scales[[nm]]
if (all(low %in% sc)) {
return(list(ordered = TRUE, order = levs[order(match(low, sc))],
rule = paste0("the labels are all points on ", nm,
", so they were ordered along it")))
}
}
}
list(ordered = FALSE, order = levs,
rule = "no numeric values, numeric prefixes, or recognised ordinal wording were found in the labels")
}Step 1: Collect the mapped judgement columns, in mapped order
rc <- grep("^rater_[0-9]+$", names(df), value = TRUE)
rc <- rc[order(as.numeric(sub("^rater_", "", rc)))]
ch <- humanize_semantic(rc, col_map)
if (length(rc) < 3) {
stop(sprintf("Only %d judgement column(s) were mapped(%s). This analysis needs three or more raters. For exactly two raters the right tool is Cohen's kappa, not Fleiss' — or, if your file has one row per rating, map the three columns holding the item, the rater and the judgement.",
length(rc),
if (length(ch) > 0) paste(ch, collapse = ", ") else "none"))
}
cols <- lapply(rc, function(cn) as_label(df[[cn]]))
names(cols) <- rc
d_distinct <- vapply(cols, function(v) length(unique(v[v != ""])), integer(1))Step 2: Detect the input SHAPE from the data, not from the mapping.
Wide means one column per rater and one row per item. Long means one row per rating, with an item column, a rater column and a judgement column. A long file is unmistakable in the counts: the item column carries many distinct values, each repeating once per rater, while the other two carry only a handful.
shape <- "wide"
shape_rule <- ""
li <- ri <- ji <- NA_integer_
if (length(rc) == 3) {
ii <- which.max(d_distinct)
others <- setdiff(seq_len(3), ii)
avg_rep <- if (d_distinct[ii] > 0) initial_rows / d_distinct[ii] else 0
if (d_distinct[ii] >= 10 && avg_rep >= 2.5 && min(d_distinct[others]) >= 2 &&
d_distinct[ii] > 3 * max(d_distinct[others])) {
item_v <- cols[[ii]]Which of the remaining two is the RATER? The one for which each item appears at most once — a rater judges an item once, whereas several raters give an item the same judgement all the time.
dupr <- vapply(others, function(o)
sum(duplicated(paste0(item_v, "\r", cols[[o]]))), numeric(1))
ri <- others[which.min(dupr)]
ji <- setdiff(others, ri)
li <- ii
shape <- "long"
shape_rule <- sprintf(
"%s holds %d distinct values each repeating about %s times, while %s and %s hold only %d and %d distinct values — that is one row per rating, not one column per rater",
ch[li], d_distinct[li], r2(avg_rep), ch[ri], ch[ji],
d_distinct[ri], d_distinct[ji])
}
}Step 3: Flatten either shape into one table of (item, rater, label)
n_dup_ratings <- 0L
if (shape == "long") {
item_v <- cols[[li]]; rater_v <- cols[[ri]]; lab_v <- cols[[ji]]
ok <- item_v != "" & rater_v != ""
n_blank_rows <- sum(!ok)
item_v <- item_v[ok]; rater_v <- rater_v[ok]; lab_v <- lab_v[ok]The same rater judging the same item twice is a data error, not a second opinion; the first judgement is kept and the count is reported.
key <- paste0(item_v, "\r", rater_v)
dupd <- duplicated(key)
n_dup_ratings <- sum(dupd)
item_v <- item_v[!dupd]; rater_v <- rater_v[!dupd]; lab_v <- lab_v[!dupd]
if (length(unique(rater_v)) > 20) {
stop(sprintf("%s holds %d distinct raters. Above 20 raters this looks like an identifier column rather than a panel; map the column naming who gave each judgement.",
ch[ri], length(unique(rater_v))))
}
row_unit <- "rating"
shape_rule <- paste0("the three mapped columns were read as a long file because ", shape_rule)
} else {Wide: a column holding a distinct value on nearly every row is an identifier or free text, not a category, and is refused by name. Both an absolute cap and a proportional test are needed — a key column in a short file has few distinct values in absolute terms but repeats nothing.
n_filled <- vapply(cols, function(v) sum(v != ""), integer(1))
bad <- which(d_distinct > 25 | (d_distinct >= 10 & d_distinct > 0.5 * n_filled))
if (length(bad) > 0) {
b <- bad[1]
why <- if (d_distinct[b] > 25)
"which is too many to be a set of categories"
else
sprintf("across %d judgement(s) — nearly one per row, so almost nothing repeats",
n_filled[b])
stop(sprintf("%s holds %d distinct values, %s. That makes it an identifier or free text rather than a rater's category choice. Map one column per rater, each holding that rater's category. If your file instead has one row per rating, map exactly three columns: the item, the rater, and the judgement.",
ch[b], d_distinct[b], why))
}
n_blank_rows <- 0L
item_v <- rep(as.character(seq_len(initial_rows)), times = length(rc))
rater_v <- rep(ch, each = initial_rows)
lab_v <- unlist(cols, use.names = FALSE)
keep <- lab_v != ""
item_v <- item_v[keep]; rater_v <- rater_v[keep]; lab_v <- lab_v[keep]
row_unit <- "item"
shape_rule <- sprintf(
"the %d mapped columns were read as one column per rater, because each holds only a handful of repeated labels(%s distinct) rather than one value per row",
length(rc), paste(d_distinct, collapse = ", "))
}Step 4: Merge labels that differ only in capitalisation, and say so.
by_case <- split(lab_v, tolower(lab_v))
canon <- list(); n_case_merged <- 0L; case_example <- NULL
for (keyc in names(by_case)) {
spellings <- by_case[[keyc]]
uq <- unique(spellings)
tab <- sort(table(spellings), decreasing = TRUE)
winner <- names(tab)[1]
canon[[keyc]] <- winner
if (length(uq) > 1) {
n_case_merged <- n_case_merged + sum(spellings != winner)
if (is.null(case_example)) case_example <- c(setdiff(uq, winner)[1], winner)
}
}
lab_v <- unlist(canon[tolower(lab_v)], use.names = FALSE)Step 5: Refuse free text; lump a long tail of rare labels into "Other".
freq <- sort(table(lab_v), decreasing = TRUE)
n_lumped <- 0L
if (length(freq) > 25) {
stop(sprintf("The judgement columns(%s) together hold %d distinct values — that looks like free text rather than a set of categories. Map the columns holding each rater's category choice.",
paste(ch, collapse = ", "), length(freq)))
}
if (length(freq) > 12) {
keep_lab <- names(freq)[1:11]
n_lumped <- length(freq) - 11L
lab_v[!(lab_v %in% keep_lab)] <- "Other"
}
rater_names <- sort(unique(rater_v))
n_raters <- length(rater_names)
if (n_raters < 3) {
stop(sprintf("Only %d rater(s) were found in %s. Fleiss' kappa needs three or more; with exactly two raters use Cohen's kappa instead.",
n_raters, if (shape == "long") ch[ri] else paste(ch, collapse = ", ")))
}
levs_all <- sort(unique(lab_v))
if (length(levs_all) < 2) {
stop(sprintf("Every judgement in %s is \"%s\", so there is only one category and agreement beyond chance is undefined — nothing varies for kappa to explain.",
paste(ch, collapse = ", "), levs_all[1]))
}Step 6: Decide the category order (detected, not assumed)
det <- detect_category_order(levs_all)
if (det$ordered) {
levs <- det$order
} else {
tot <- vapply(levs_all, function(l) sum(lab_v == l), numeric(1))
levs <- levs_all[order(-tot, levs_all)]
}
k <- length(levs)Step 7: Build the item-by-category count matrix
items <- unique(item_v)
Nmat <- matrix(0, nrow = length(items), ncol = k,
dimnames = list(NULL, levs))
im <- match(item_v, items)
cm <- match(lab_v, levs)
tabm <- table(factor(im, levels = seq_along(items)),
factor(cm, levels = seq_len(k)))
Nmat[] <- as.numeric(tabm)
m_i <- rowSums(Nmat)
n_ratings <- sum(m_i)
all_idx <- which(m_i >= 2)
n_items_dropped <- length(items) - length(all_idx)
complete_idx <- which(m_i == n_raters)
n_items_partial <- sum(m_i >= 2 & m_i < n_raters)
n_missing <- n_raters * length(items) - n_ratings
if (length(all_idx) < 20) {
stop(sprintf("Only %d item(s) carry judgements from at least two of %s. Multi-rater agreement needs at least 20 such items before the numbers mean anything.",
length(all_idx), paste(ch, collapse = ", ")))
}Step 8: Fleiss' kappa. It is a complete-cases statistic: the classical
coefficient assumes every item was judged by the same number of raters. When ratings are missing the module reports BOTH the complete-case value (the one other software produces) and a generalized value over every item with two or more ratings — and says which items each one used.
fleiss_calc <- function(ii) {
m <- m_i[ii]
sq <- rowSums(Nmat[ii, , drop = FALSE]^2)
P_i <- (sq - m) / (m * (m - 1))
P_bar <- mean(P_i)
tot <- sum(m)
p <- colSums(Nmat[ii, , drop = FALSE]) / tot
P_e <- sum(p^2)
kap <- if ((1 - P_e) > 1e-12) (P_bar - P_e) / (1 - P_e) else NA_real_
list(P_bar = P_bar, P_e = P_e, kappa = kap, p = p, tot = tot,
n = length(ii), m_const = if (length(unique(m)) == 1) m[1] else NA_real_)
}
basis <- if (length(complete_idx) >= 20) "complete" else "all"
basis_idx <- if (basis == "complete") complete_idx else all_idx
fb <- fleiss_calc(basis_idx)
fa <- fleiss_calc(all_idx)
p_bar <- fb$P_bar; p_e <- fb$P_e; kappa <- fb$kappa
p_j <- fb$p
kappa_all <- fa$kappa
n_basis <- fb$nFleiss' null-hypothesis standard error (Fleiss 1971), which exists only when every item on the basis carries the same number of raters. It is the variance of kappa when kappa is truly zero, so it belongs to the significance test and NOT to the confidence interval.
se_null <- NA_real_; z_stat <- NA_real_; p_value <- NA_real_
if (!is.na(fb$m_const) && fb$m_const >= 2) {
R <- fb$m_const
S2 <- sum(p_j^2); S3 <- sum(p_j^3)
inner <- S2 - (2 * R - 3) * S2^2 + 2 * (R - 2) * S3
if (is.finite(inner) && inner > 0 && (1 - S2) > 1e-12) {
se_null <- sqrt(2 / (n_basis * R * (R - 1)) * inner) / (1 - S2)
z_stat <- kappa / se_null
p_value <- 2 * pnorm(-abs(z_stat))
}
}Step 9: Krippendorff's alpha, via the coincidence matrix. Alpha is the
statistic that tolerates missing ratings, so it uses EVERY item carrying at least two judgements — including the ones Fleiss had to set aside.
Ocon <- t(apply(Nmat, 1, function(row) {
m <- sum(row)
if (m < 2) return(rep(0, k * k))
as.vector((outer(row, row) - diag(row, nrow = k)) / (m - 1))
}))
if (k == 1) Ocon <- matrix(Ocon, ncol = 1)
LO <- outer(seq_len(k), seq_len(k), pmin)
HI <- outer(seq_len(k), seq_len(k), pmax)
ordinal_ok <- det$ordered && k >= 3
alpha_calc <- function(ii) {
o <- matrix(colSums(Ocon[ii, , drop = FALSE]), nrow = k, ncol = k)
nc <- rowSums(o)
ntot <- sum(nc)
a_nom <- NA_real_; a_ord <- NA_real_
De_n <- ntot^2 - sum(nc^2)
if (is.finite(De_n) && De_n > 1e-12 && ntot > 1) {
Do_n <- ntot - sum(diag(o))
a_nom <- 1 - (ntot - 1) * Do_n / De_n
}
if (ordinal_ok && ntot > 1) {
cum0 <- c(0, cumsum(nc))
S <- matrix(cum0[HI + 1] - cum0[LO], nrow = k, ncol = k)
Dm <- (S - outer(nc, nc, function(x, y) (x + y) / 2))^2
De_o <- sum(outer(nc, nc) * Dm)
if (is.finite(De_o) && De_o > 1e-12) {
a_ord <- 1 - (ntot - 1) * sum(o * Dm) / De_o
}
}
list(nominal = a_nom, ordinal = a_ord)
}
aa <- alpha_calc(all_idx)
alpha_nom <- aa$nominal; alpha_ord <- aa$ordinal
alpha_nom_basis <- alpha_calc(basis_idx)$nominalStep 10: Confidence intervals by resampling ITEMS. The closed-form
variance above is a null-hypothesis quantity and is wrong for an interval around a non-zero estimate; alpha has no simple closed form at all. A seeded nonparametric bootstrap over items answers both, and treats items as the sampling unit, which is what they are.
B <- 500L
set.seed(42)
nb <- length(basis_idx); na_ <- length(all_idx)
bk <- numeric(B); ban <- numeric(B); bao <- numeric(B)
for (b in seq_len(B)) {
ib <- basis_idx[sample.int(nb, nb, replace = TRUE)]
bk[b] <- fleiss_calc(ib)$kappa
ia <- all_idx[sample.int(na_, na_, replace = TRUE)]
ab <- alpha_calc(ia)
ban[b] <- ab$nominal; bao[b] <- ab$ordinal
}
qci <- function(v) {
v <- v[is.finite(v)]
if (length(v) < 50) return(c(NA_real_, NA_real_, NA_real_))
c(as.numeric(quantile(v, 0.025, names = FALSE)),
as.numeric(quantile(v, 0.975, names = FALSE)), sd(v))
}
qk <- qci(bk); qan <- qci(ban); qao <- qci(bao)
ci_low <- qk[1]; ci_high <- qk[2]; se_boot <- qk[3]
alpha_nom_lo <- qan[1]; alpha_nom_hi <- qan[2]
alpha_ord_lo <- qao[1]; alpha_ord_hi <- qao[2]
band <- landis_koch(kappa)Step 11: Per-category kappa — which categories the panel fights over.
Fleiss' category-specific coefficient, generalized to a variable number of raters per item.
mb <- m_i[basis_idx]
Nb <- Nmat[basis_idx, , drop = FALSE]
tot_b <- sum(mb)
cat_kappa <- vapply(seq_len(k), function(j) {
pj <- sum(Nb[, j]) / tot_b
den <- tot_b * pj * (1 - pj)
if (!is.finite(den) || den <= 1e-12) return(NA_real_)
obs <- sum(Nb[, j] * (mb - Nb[, j]) / (mb - 1))
1 - obs / den
}, numeric(1))
category_df <- data.frame(
category = levs,
category_kappa = round(cat_kappa, 4),
share_pct = round(100 * colSums(Nb) / tot_b, 2),
n_uses = as.numeric(colSums(Nb)),
stringsAsFactors = FALSE
)
ok_cat <- which(!is.na(cat_kappa))
worst_cat <- if (length(ok_cat) > 0) levs[ok_cat[which.min(cat_kappa[ok_cat])]] else NA_character_
worst_val <- if (length(ok_cat) > 0) min(cat_kappa[ok_cat]) else NA_real_
best_cat <- if (length(ok_cat) > 0) levs[ok_cat[which.max(cat_kappa[ok_cat])]] else NA_character_
best_val <- if (length(ok_cat) > 0) max(cat_kappa[ok_cat]) else NA_real_With exactly two categories every disagreement is the same disagreement, so both category-specific kappas equal each other and the overall coefficient. Naming a "worst" category there would be an artefact of tie-breaking.
cat_degenerate <- k == 2 ||
(length(ok_cat) > 1 && (best_val - worst_val) < 1e-9)Step 12: Per-rater agreement with the rest of the panel, and the
leave-one-rater-out kappa. The second is the actionable one: it says what the panel's agreement would be if this rater were not in it.
basis_items <- items[basis_idx]
in_basis <- item_v %in% basis_items
bi_item <- item_v[in_basis]; bi_rater <- rater_v[in_basis]; bi_lab <- lab_v[in_basis]
bi_row <- match(bi_item, basis_items)
bi_col <- match(bi_lab, levs)
m_of_row <- mb[bi_row]
n_same <- Nb[cbind(bi_row, bi_col)]
rater_obs <- numeric(n_raters); rater_kap <- numeric(n_raters)
rater_wo <- numeric(n_raters); rater_n <- numeric(n_raters)
for (r in seq_len(n_raters)) {
sel <- bi_rater == rater_names[r]
den <- sum(m_of_row[sel] - 1)
rater_n[r] <- sum(sel)
rater_obs[r] <- if (den > 0) sum(n_same[sel] - 1) / den else NA_real_
rater_kap[r] <- if (is.na(rater_obs[r]) || (1 - p_e) <= 1e-12) NA_real_
else (rater_obs[r] - p_e) / (1 - p_e)Leave-one-rater-out: strip this rater's ratings and recompute Fleiss on the items that still carry two or more judgements.
Nw <- Nb
idxw <- which(sel)
if (length(idxw) > 0) {
dec <- table(factor(bi_row[idxw], levels = seq_len(nrow(Nb))),
factor(bi_col[idxw], levels = seq_len(k)))
Nw <- Nw - matrix(as.numeric(dec), nrow = nrow(Nb), ncol = k)
}
mw <- rowSums(Nw)
keepw <- which(mw >= 2)
if (length(keepw) >= 5) {
sqw <- rowSums(Nw[keepw, , drop = FALSE]^2)
Pw <- (sqw - mw[keepw]) / (mw[keepw] * (mw[keepw] - 1))
totw <- sum(mw[keepw])
pw <- colSums(Nw[keepw, , drop = FALSE]) / totw
Pew <- sum(pw^2)
rater_wo[r] <- if ((1 - Pew) > 1e-12) (mean(Pw) - Pew) / (1 - Pew) else NA_real_
} else {
rater_wo[r] <- NA_real_
}
}
rater_df <- data.frame(
rater = rater_names,
agreement_pct = round(100 * rater_obs, 2),
rater_kappa = round(rater_kap, 4),
kappa_without_rater = round(rater_wo, 4),
items_rated = rater_n,
stringsAsFactors = FALSE
)
ok_r <- which(!is.na(rater_kap))
outlier_rater <- if (length(ok_r) > 0) rater_names[ok_r[which.min(rater_kap[ok_r])]] else NA_character_
outlier_kappa <- if (length(ok_r) > 0) min(rater_kap[ok_r]) else NA_real_
best_rater <- if (length(ok_r) > 0) rater_names[ok_r[which.max(rater_kap[ok_r])]] else NA_character_
best_rater_kappa <- if (length(ok_r) > 0) max(rater_kap[ok_r]) else NA_real_
outlier_gain <- if (!is.na(outlier_rater) && !is.na(rater_wo[match(outlier_rater, rater_names)]))
rater_wo[match(outlier_rater, rater_names)] - kappa else NA_real_
rater_spread <- if (length(ok_r) > 1) max(rater_kap[ok_r]) - min(rater_kap[ok_r]) else NA_real_Step 13: Item-level consensus. Banded by each item's MODAL share —
the largest number of raters landing on one category, over the raters who judged it. Only the maximum COUNT is used, never its position, so an evenly split item needs no arbitrary winner.
modal_share <- apply(Nb, 1, max) / mb
band_labels <- c("No majority — the raters split",
"Simple majority — more than half",
"Strong majority — more than two thirds",
"Unanimous — every rater agreed")
band_idx <- ifelse(modal_share >= 1, 4L,
ifelse(modal_share > 2 / 3, 3L,
ifelse(modal_share > 0.5, 2L, 1L)))
item_df <- data.frame(
consensus_band = band_labels,
n_items = as.numeric(tabulate(band_idx, nbins = 4)),
stringsAsFactors = FALSE
)
item_df$share_pct <- round(100 * item_df$n_items / n_basis, 2)
n_unanimous <- item_df$n_items[4]
n_split <- item_df$n_items[1]Step 14: Per-rater marginals — how each rater uses the scale. These
are the habits behind both the chance floor and any outlier.
marg <- matrix(0, nrow = n_raters, ncol = k, dimnames = list(rater_names, levs))
tabr <- table(factor(bi_rater, levels = rater_names),
factor(bi_lab, levels = levs))
marg[] <- as.numeric(tabr)
marg_share <- marg / pmax(rowSums(marg), 1)
overall_share <- colSums(marg) / sum(marg)
divergence <- rowSums(abs(sweep(marg_share, 2, overall_share)))
show_r <- rater_names
n_raters_hidden <- 0L
if (n_raters > 8) {
show_r <- rater_names[order(-divergence)][1:8]
n_raters_hidden <- n_raters - 8L
}
marginal_df <- data.frame(
category = rep(levs, times = length(show_r)),
rater = rep(show_r, each = k),
share_pct = round(100 * as.vector(t(marg_share[show_r, , drop = FALSE])), 2),
stringsAsFactors = FALSE
)
prev_i <- which.max(p_j)
max_prev <- p_j[prev_i]
prev_cat <- levs[prev_i]
paradox <- isTRUE(p_bar >= 0.70 && !is.na(kappa) && kappa < 0.60 && max_prev >= 0.60)Step 15: The results table
basis_phrase <- if (basis == "complete")
sprintf("the %s item(s) every one of the %d raters judged", fmt_n(n_basis), n_raters)
else
sprintf("all %s item(s) carrying at least two judgements", fmt_n(n_basis))
rows <- list(
list("Raw pairwise agreement", p_bar, NA_real_, NA_real_,
sprintf("Across every pair of raters who judged the same item, %s of those pairs chose the same category.",
pct1(p_bar))),
list("Agreement expected by chance", p_e, NA_real_, NA_real_,
sprintf("What a panel using the categories this often would have agreed on by luck alone(%s).",
pct1(p_e))),
list("Fleiss' kappa", kappa, ci_low, ci_high,
sprintf("%s of the agreement left available above chance was achieved, on %s — %s on the Landis & Koch convention.",
pct1(kappa), basis_phrase, band))
)
if (n_missing > 0 && basis == "complete") {
rows[[length(rows) + 1]] <- list(
"Fleiss' kappa (all items, generalized)", kappa_all, NA_real_, NA_real_,
sprintf("The same coefficient generalized to a variable number of raters per item, so it uses all %s items with two or more judgements instead of only the %s complete ones.",
fmt_n(length(all_idx)), fmt_n(n_basis)))
}
rows[[length(rows) + 1]] <- list(
"Krippendorff's alpha (nominal)", alpha_nom, alpha_nom_lo, alpha_nom_hi,
sprintf("The disagreement-based coefficient, computed on all %s items with two or more judgements — it does not require every rater to have judged every item.",
fmt_n(length(all_idx))))
if (!is.na(alpha_ord)) {
rows[[length(rows) + 1]] <- list(
"Krippendorff's alpha (ordinal)", alpha_ord, alpha_ord_lo, alpha_ord_hi,
"The same coefficient with an ordinal distance between categories, so a near-miss between neighbouring levels counts as a smaller disagreement than a jump across the scale.")
}
if (n_missing > 0 && basis == "complete" && !is.na(alpha_nom_basis)) {
rows[[length(rows) + 1]] <- list(
"Krippendorff's alpha (complete items only)", alpha_nom_basis, NA_real_, NA_real_,
"Alpha restricted to the same complete items Fleiss' kappa used, so the two headline coefficients can be compared on identical data.")
}
agreement_df <- data.frame(
statistic = vapply(rows, function(r) r[[1]], character(1)),
estimate = round(vapply(rows, function(r) as.numeric(r[[2]]), numeric(1)), 4),
ci_low = round(vapply(rows, function(r) as.numeric(r[[3]]), numeric(1)), 4),
ci_high = round(vapply(rows, function(r) as.numeric(r[[4]]), numeric(1)), 4),
interpretation = vapply(rows, function(r) r[[5]], character(1)),
stringsAsFactors = FALSE
)Step 16: Methods disclosure
se_line <- if (is.na(se_null)) {
sprintf("The classical null-hypothesis standard error is not available here, because the items on the analysis basis do not all carry the same number of raters. The significance test is therefore not reported; read the bootstrap interval instead, which does not need that assumption.")
} else {
sprintf("Fleiss' null-hypothesis standard error is %s, giving z = %s and %s against the hypothesis that agreement is no better than chance. That variance describes kappa when kappa is truly zero, so it belongs to this test and NOT to the confidence interval.",
r3(se_null), r2(z_stat), fmt_pp(p_value))
}
methods_df <- data.frame(
item = c(
"Design",
"Input shape",
"Raw pairwise agreement",
"Chance agreement",
"Fleiss' kappa",
"Significance test",
"Confidence intervals",
"Missing ratings",
"Krippendorff's alpha",
"Category order",
"Per-category kappa",
"Per-rater figures",
"Benchmark labels",
"What agreement is not"
),
detail = c(
sprintf("%s item(s) judged by %d rater(s) across %d categories, %s judgement(s) in total.",
fmt_n(length(items)), n_raters, k, fmt_n(n_ratings)),
sprintf("Detected from the data, not from the mapping: %s.", shape_rule),
sprintf("Mean over items of the share of rater PAIRS on that item choosing the same category: %s.",
pct1(p_bar)),
sprintf("Sum of the squared overall category shares(%s): %s.",
paste(sprintf("%s %s", levs, vapply(p_j, pct1, character(1))), collapse = ", "),
pct1(p_e)),
sprintf("(observed - expected) / (1 - expected) = (%s - %s) / (1 - %s) = %s, computed on %s.",
r3(p_bar), r3(p_e), r3(p_e), r3(kappa), basis_phrase),
se_line,
sprintf("A seeded nonparametric bootstrap resampling ITEMS with replacement(%d resamples, fixed seed), taking the 2.5th and 97.5th percentiles. The bootstrap standard error of kappa is %s. This is used rather than a closed form because the null variance above is the wrong quantity for an interval around a non-zero estimate, and alpha has no simple closed-form variance.",
B, r3(se_boot)),
if (n_missing > 0)
sprintf("%s of the %s possible rater-by-item judgements are absent. Fleiss' kappa is a complete-cases statistic and used %s; Krippendorff's alpha used all %s items with at least two judgements; %s item(s) carrying a single judgement contribute to neither, because agreement needs at least two opinions.",
fmt_n(n_missing), fmt_n(n_raters * length(items)), basis_phrase,
fmt_n(length(all_idx)), fmt_n(n_items_dropped))
else
"No ratings are missing: every rater judged every item, so Fleiss' kappa and Krippendorff's alpha are computed on exactly the same data.",
sprintf("Built from the coincidence matrix: alpha = 1 - observed disagreement / expected disagreement, with(n-1) weighting so it is defined for small samples. Nominal alpha treats every disagreement as equal. %s",
if (!is.na(alpha_ord))
"Ordinal alpha weights a disagreement by the distance between the categories along the detected order."
else
"Ordinal alpha is not reported for these categories."),
sprintf("Orderedness was detected, not assumed: %s.", det$rule),
"Fleiss' category-specific coefficient, generalized to a variable number of raters per item: one minus the observed splits on that category over the splits expected from its overall share.",
sprintf("A rater's agreement figure is the share of the OTHER raters' judgements on the same items that matched this rater's own, chance-corrected against the same %s floor. The leave-one-out column recomputes Fleiss' kappa for the panel with that rater removed.",
pct1(p_e)),
"Landis & Koch(1977) labels(slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.",
"Agreement is not correctness: a panel can agree unanimously and be unanimously wrong. These figures describe this panel on these items and do not generalise to other raters or other items."
),
stringsAsFactors = FALSE
)Step 17: Headline metrics + the computed one-paragraph answer
metrics <- list(
`Items Rated` = as.numeric(length(items)),
`Raters` = as.numeric(n_raters),
`Categories` = as.numeric(k),
`Raw Agreement` = pct1(p_bar),
`Fleiss Kappa` = round(kappa, 3),
`Fleiss 95% CI` = paste0(r3(ci_low), " to ", r3(ci_high)),
`Krippendorff Alpha` = round(alpha_nom, 3),
`Agreement Strength` = band,
`Least Aligned Rater` = if (is.na(outlier_rater)) "not identifiable" else outlier_rater,
`Kappa vs Chance p` = fmt_p(p_value)
)
paradox_sentence <- if (paradox) {
paste0(
" Read the two agreement numbers together before quoting either: raw pairwise agreement is high(",
pct1(p_bar), ") while kappa is only ", r3(kappa),
", which is the well-known kappa paradox rather than a contradiction. ",
pct1(max_prev), " of all judgements fell into the single category \"", prev_cat,
"\", so a panel with these habits would already have agreed on ", pct1(p_e),
" of rater pairs by chance; only ", pct1(1 - p_e),
" of the scale was left for skill to win, and ", pct1(kappa), " of that remainder was won.")
} else {
paste0(
" Chance alone would have produced ", pct1(p_e),
" pairwise agreement given how often each category is used, leaving ",
pct1(1 - p_e), " of the scale available above chance, of which ",
pct1(kappa), " was achieved.")
}
missing_sentence <- if (n_missing > 0) {
paste0(" ", fmt_n(n_missing), " of the ", fmt_n(n_raters * length(items)),
" possible judgements are missing, which the two coefficients handle differently: Fleiss' kappa is a complete-cases statistic and used ",
basis_phrase, ", while Krippendorff's alpha used all ", fmt_n(length(all_idx)),
" items carrying at least two judgements",
if (n_items_dropped > 0)
paste0(" and ", fmt_n(n_items_dropped),
" item(s) with a single judgement were used by neither")
else "",
". Where ratings are missing, alpha is the coefficient to quote.")
} else ""
ordinal_sentence <- if (!is.na(alpha_ord)) {
paste0(" The categories are ordered(", det$rule,
"), so ordinal alpha(", r3(alpha_ord),
") is also reported; it counts a disagreement between neighbouring levels as smaller than a jump across the scale, which is why it sits above the nominal value of ",
r3(alpha_nom), ".")
} else {
paste0(" Ordinal alpha is not reported because ", det$rule,
"; treating unordered labels as if some disagreements were milder than others would invent structure the data does not carry.")
}
outlier_sentence <- if (!is.na(outlier_rater) && !is.na(outlier_gain)) {
paste0(" ", outlier_rater, " is the least aligned member of the panel(rater kappa ",
r3(outlier_kappa), " against ", r3(best_rater_kappa), " for ", best_rater,
"); with that rater excluded the panel's kappa would be ",
r3(kappa + outlier_gain),
if (outlier_gain > 0) paste0(", ", r3(outlier_gain), " higher than it is now")
else paste0(", ", r3(abs(outlier_gain)), " lower than it is now"),
".")
} else ""
json_output <- list(
answer = paste0(
fmt_n(n_raters), " raters judged ", fmt_n(length(items)), " items across ", k,
" categories, agreeing on ", pct1(p_bar),
" of rater pairs. Fleiss' kappa is ", r3(kappa),
" (95% CI ", r3(ci_low), " to ", r3(ci_high), "), ", band,
" agreement on the Landis & Koch convention, and ",
if (!is.na(p_value) && p_value < 0.05)
paste0("clearly better than chance(", fmt_pp(p_value), ")")
else if (!is.na(p_value))
paste0("not distinguishable from chance(", fmt_pp(p_value), ")")
else
"could not be tested against chance because the raters did not all judge the same items",
"; Krippendorff's alpha is ", r3(alpha_nom), ".",
paradox_sentence, missing_sentence, ordinal_sentence, outlier_sentence,
if (cat_degenerate)
paste0(" With only ", k,
" categories every disagreement is the same disagreement, so the category-specific kappas are necessarily equal to each other and to the overall coefficient — there is no per-category story to tell here.")
else
paste0(" The panel is least consistent on the category \"", worst_cat,
"\" (category kappa ", r3(worst_val), ") and most consistent on \"",
best_cat, "\" (", r3(best_val), ")."),
" Agreement is not correctness: a panel can agree unanimously and be unanimously wrong."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "agreement_results", "rater_consensus",
"category_kappa", "item_consensus", "rater_marginals", "methods"),
function(cid) list(id = cid, metrics = metrics)
)
)
final_rows <- if (row_unit == "item") length(all_idx) else n_ratings
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = max(0, initial_rows - final_rows),
row_unit = row_unit, shape = shape, shape_rule = shape_rule,
n_blank_rows = n_blank_rows, n_dup_ratings = n_dup_ratings,
n_case_merged = n_case_merged, case_example = case_example, n_lumped = n_lumped,
col_names = ch, rater_names = rater_names, n_raters = n_raters,
levs = levs, k_cats = k, n_items = length(items), n_ratings = n_ratings,
n_missing = n_missing, n_items_partial = n_items_partial,
n_items_dropped = n_items_dropped, n_all = length(all_idx),
basis = basis, n_basis = n_basis, n_basis_ratings = tot_b,
basis_phrase = basis_phrase,
p_bar = p_bar, p_e = p_e, p_j = p_j, kappa = kappa, kappa_all = kappa_all,
se_null = se_null, se_boot = se_boot, z_stat = z_stat, p_value = p_value,
ci_low = ci_low, ci_high = ci_high,
alpha_nom = alpha_nom, alpha_nom_lo = alpha_nom_lo, alpha_nom_hi = alpha_nom_hi,
alpha_ord = alpha_ord, alpha_nom_basis = alpha_nom_basis,
ordered_flag = det$ordered, order_rule = det$rule, band = band,
paradox = paradox, paradox_sentence = paradox_sentence,
missing_sentence = missing_sentence, ordinal_sentence = ordinal_sentence,
outlier_sentence = outlier_sentence,
max_prev = max_prev, prev_cat = prev_cat,
outlier_rater = outlier_rater, outlier_kappa = outlier_kappa,
outlier_gain = outlier_gain, best_rater = best_rater,
best_rater_kappa = best_rater_kappa, rater_spread = rater_spread,
n_raters_hidden = n_raters_hidden,
worst_cat = worst_cat, worst_val = worst_val,
best_cat = best_cat, best_val = best_val, cat_degenerate = cat_degenerate,
n_unanimous = n_unanimous, n_split = n_split,
agreement_df = agreement_df, rater_df = rater_df, category_df = category_df,
item_df = item_df, marginal_df = marginal_df, methods_df = methods_df,
metrics = metrics, json_output = json_output, boot_B = B
)
}Your turn
Bring your own data and the question you actually need answered.
CympleData Scientist Send me your data and question, I’ll send you the analytics. ds@mcpanalytics.ai