Free, no account required

Same Cases, Judged Twice. Does One Say Yes More Often? Answer In Minutes

Upload one row per case with two yes/no verdicts and get the full McNemar report: all four variants with the one you are quoting named, the disagreements that carry the evidence, both positive rates, and agreement with kappa alongside. Free.

24,000+ analyses run
Encrypted & deleted in 7 days
PDF & citation included

Free analyses run on up to 100,000 rows. Larger files are randomly sampled to that size, so sign up to analyze your full dataset.

📊
-
Rows
-
Columns
-
Numeric

Running mcnemar's test analysis...

Testing your paired verdicts...

Your report is ready

Sent to . Inside: the paired table, all four McNemar variants with the headline named, the disagreement split, marginal rates, agreement and kappa, R code, and AI insights.

Analyze another file
Sample Output

Every report includes interactive charts, tables, and AI insights

Upload your data to get your own report

View all case studies See all free tools
Why an analysis, not a chat answer

Ask twice, get two answers? Watch how we fix that.

How it works

The analysis builds the paired two-by-two table from one row per case, then reports all four McNemar variants from the same table: the uncorrected chi-square as the headline (Fagerland, Lydersen and Laake 2013 recommend against the continuity correction), the continuity-corrected chi-square that R returns by default, the exact binomial test, and the mid-p test. Only the two discordant cells enter the test, so the split between them is shown directly, with the marginal positive rate for each judging as the effect size to report beside the p-value. Percent agreement and Cohen's kappa are computed alongside so the reader can see that a high agreement score and a decisive McNemar result are not in conflict: they read different cells of the same table.

Use it whenever the same cases carry two binary verdicts: two reviewers scoring the same applications, two diagnostic tests run on the same patients, one model evaluated before and after a change on the same test set, or the same people answering yes or no before and after an intervention.

Not for two separate groups of subjects, which is the ordinary chi-square test of independence. Not for more than two response categories on the same cases, which is Stuart-Maxwell. Not for more than two judgings, which is Cochran's Q. Not for numeric ratings, which is intraclass correlation.

Built for: Researchers, clinicians, risk and credit teams, QA leads, and ML engineers comparing two paired yes/no judgments

Typical data source: A spreadsheet with one row per case and two yes/no columns: what was judged, and the two verdicts

HealthcareResearchFinancial ServicesMachine LearningEducationManufacturing

What data do you need?

One row per case, with the two verdicts side by side. For example, two loan officers deciding on the same 200 applications:

case_id (categorical) reviewer_a (categorical) reviewer_b (categorical)
case_001 approve approve
case_002 deny approve
case_003 approve approve

Minimum 10 rows · Best with 50 or more cases, and ideally 10 or more disagreements (any number of rows up to 100,000)

What's in the report?

Standard-library analysis: when the same cases are judged twice, does one judging say yes more often? Upload one row per case with two binary verdict columns and get the paired two-by-two table, all four McNemar variants (uncorrected headline per Fagerland 2013, continuity-corrected, exact binomial, mid-p) with the variant named on the page, the disagreement split that carries the evidence, the two marginal positive rates as the effect size, and percent agreement plus kappa alongside as an explicit contrast. Built for paired-proportions questions: two reviewers on the same files, two diagnostic tests on the same patients, one classifier before and after a change.

📋

McNemar Results: Name Your Variant

All four variants side by side, with the one your software would have picked marked, so a reader can reproduce your number.

📊

The Disagreements Carry the Evidence

Only the cases the two judgings disagreed on. An even split is what chance looks like; a lopsided one has a direction.

📊

Marginal Rates: The Effect Size

Each judging's yes rate on identical cases. The gap between them is the effect size to report next to the p-value.

📋

Agreement vs McNemar: Different Questions

The full paired table with the tested cells marked, so it is visible that agreement and this test are reading different parts of it.

🤖

AI Insights

Plain-English interpretation of what the numbers mean, what's significant, and what to do next.

The Question This Answers

Do my two reviewers approve at the same rate?

Map the case id and each reviewer's verdict. You get the paired table, all four McNemar variants with the headline named, the split of the cases they disagreed on, and both reviewers' approval rates on identical files. A lopsided split means one reviewer holds a looser bar, which is a calibration problem rather than a competence one.

Questions?

See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.

Your data has more stories to tell

Run any analysis on your own data: validated R analyses, interactive reports, AI insights, and PDF export.

Try Free, No Credit Card
Powered by MCP Analytics