ds@mcpanalytics.aiA data scientist you can email your data and question to, and get a reliable response. Not AI slop.A data scientist you can email. Not AI slop.
Free, no account required

How Reliable Are Your Raters? Find Out In Minutes

Upload your ratings, map who rated what, and get the full intraclass correlation report — ICC(2,1), ICC(3,1), average-measure versions, Koo & Li bands, and exactly which ICC to report. Free.

Encrypted & deleted in 7 days
PDF & citation included

Free analyses run on up to 10,000 rows. Larger files are randomly sampled to that size, so sign up to analyze your full dataset.

📊
-
Rows
-
Columns
-
Numeric

Running intraclass correlation (rater reliability) analysis...

Computing intraclass correlations...

Your report is ready

Sent to . Inside: the full ICC table with Koo & Li bands, rater bias chart, subject-level agreement plot, variance breakdown, R code, and AI insights.

Analyze another file
Want more than this?
Cymple, Data Scientist
Want the full analysis on this data? Send it to me with your question and I’ll send you the analytics.
Send Cymple my data →
Sample Output

Every report includes interactive charts, tables, and AI insights

Upload your data to get your own report

View all case studies See all free tools
The method, explained

Which ICC do you actually need?

The method, in full

The practical guide behind this tool: when it applies, how to read the output, and the traps that make it say the wrong thing.

Read the guide →

The same analysis, worked end to end

Real data, the R Markdown that produced every figure, and the numbers checked three ways.

Download it and re-run it yourself: this is what an answer looks like when someone asks you to prove it.

See the worked example →

How it works

The analysis builds the subject-by-rater grid (averaging duplicate ratings, excluding subjects not rated by everyone), then computes the whole ICC family from ANOVA mean squares: one-way ICC(1,1); two-way ICC(2,1) for absolute agreement and ICC(3,1) for consistency; and the average-measure ICC(2,k) and ICC(3,k). An F-test checks that raters distinguish subjects at all, Koo & Li (2016) bands translate each value into poor/moderate/good/excellent, and a variance decomposition splits disagreement into systematic rater bias versus random noise.

Use it for any reliability study — multiple raters scoring the same subjects, one instrument measured repeatedly, or several devices measuring the same samples — whenever the rating is numeric.

Not for categorical ratings (use Cohen's or Fleiss' kappa), for a single rater with no repeats, or when different subjects are rated on different scales.

Built for: Researchers, clinicians, QA leads, and ML teams measuring rater or instrument reliability

Typical data source: A long-format spreadsheet of ratings: what was rated, who rated it, and the score

HealthcareResearchEducationManufacturingMachine LearningSports Science

What data do you need?

Long format — one row per rating. For example, four clinicians scoring the same 30 patient scans:

sample_id (categorical) rater (categorical) quality_score (numeric)
S01 Dr. Adams 46.7
S01 Dr. Baker 48.9
S02 Dr. Adams 52.3

Minimum 10 rows · Best with 5-500 subjects and 2-10 raters (10-5,000 rows)

What's in the report?

Standard-library analysis: how consistent are your raters, instruments, or repeated measurements? Upload long-format ratings (one row per rating) and get the full intraclass correlation family — ICC(1,1), ICC(2,1), ICC(3,1) and the average-measure versions — computed from ANOVA variance components, with Koo & Li interpretation bands, a systematic rater-bias check, a subject-level agreement plot, and a variance breakdown showing exactly where the disagreement comes from. Built for reliability studies: inter-rater agreement, test-retest, instrument comparison.

📋

ICC Results — Which One to Report

Every ICC form side by side — single vs average measure, agreement vs consistency — each with its Koo & Li band and a rule for when to report it.

📊

Systematic Rater Bias

Each rater's average score on the same subjects; any gap is pure systematic leniency or severity.

🔵

Subject-Level Agreement

The first two raters plotted subject by subject against the perfect-agreement diagonal — offsets and outliers are visible at a glance.

📊

Where the Disagreement Comes From

How much of the variation is real subject differences versus rater bias versus noise — the anatomy of your reliability number.

🤖

AI Insights

Plain-English interpretation of what the numbers mean, what's significant, and what to do next.

The Question This Answers

Are my raters interchangeable?

Map what was rated, who rated it, and the score. You get the full ICC family with Koo & Li bands, a rater-bias chart that shows who scores systematically high or low, and a plain-language rule for which ICC to put in your paper.

Questions?

See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.

Your data has more stories to tell

Run any analysis on your own data: R analyses, interactive reports, AI insights, and PDF export.

Try Free, No Credit Card
Powered by MCP Analytics

Your turn

Bring your own data and the question you actually need answered.

CympleData Scientist Send me your data and question, I’ll send you the analytics. ds@mcpanalytics.ai