Upload one row per case with two yes/no verdicts and get the full McNemar report: all four variants with the one you are quoting named, the disagreements that carry the evidence, both positive rates, and agreement with kappa alongside. Free.
Free analyses run on up to 100,000 rows. Larger files are randomly sampled to that size, so sign up to analyze your full dataset.
Testing your paired verdicts...
Sent to . Inside: the paired table, all four McNemar variants with the headline named, the disagreement split, marginal rates, agreement and kappa, R code, and AI insights.
Analyze another fileThe analysis builds the paired two-by-two table from one row per case, then reports all four McNemar variants from the same table: the uncorrected chi-square as the headline (Fagerland, Lydersen and Laake 2013 recommend against the continuity correction), the continuity-corrected chi-square that R returns by default, the exact binomial test, and the mid-p test. Only the two discordant cells enter the test, so the split between them is shown directly, with the marginal positive rate for each judging as the effect size to report beside the p-value. Percent agreement and Cohen's kappa are computed alongside so the reader can see that a high agreement score and a decisive McNemar result are not in conflict: they read different cells of the same table.
Use it whenever the same cases carry two binary verdicts: two reviewers scoring the same applications, two diagnostic tests run on the same patients, one model evaluated before and after a change on the same test set, or the same people answering yes or no before and after an intervention.
Not for two separate groups of subjects, which is the ordinary chi-square test of independence. Not for more than two response categories on the same cases, which is Stuart-Maxwell. Not for more than two judgings, which is Cochran's Q. Not for numeric ratings, which is intraclass correlation.
Built for: Researchers, clinicians, risk and credit teams, QA leads, and ML engineers comparing two paired yes/no judgments
Typical data source: A spreadsheet with one row per case and two yes/no columns: what was judged, and the two verdicts
One row per case, with the two verdicts side by side. For example, two loan officers deciding on the same 200 applications:
Minimum 10 rows · Best with 50 or more cases, and ideally 10 or more disagreements (any number of rows up to 100,000)
Standard-library analysis: when the same cases are judged twice, does one judging say yes more often? Upload one row per case with two binary verdict columns and get the paired two-by-two table, all four McNemar variants (uncorrected headline per Fagerland 2013, continuity-corrected, exact binomial, mid-p) with the variant named on the page, the disagreement split that carries the evidence, the two marginal positive rates as the effect size, and percent agreement plus kappa alongside as an explicit contrast. Built for paired-proportions questions: two reviewers on the same files, two diagnostic tests on the same patients, one classifier before and after a change.
All four variants side by side, with the one your software would have picked marked, so a reader can reproduce your number.
Only the cases the two judgings disagreed on. An even split is what chance looks like; a lopsided one has a direction.
Each judging's yes rate on identical cases. The gap between them is the effect size to report next to the p-value.
The full paired table with the tested cells marked, so it is visible that agreement and this test are reading different parts of it.
Plain-English interpretation of what the numbers mean, what's significant, and what to do next.
Do my two reviewers approve at the same rate?
Map the case id and each reviewer's verdict. You get the paired table, all four McNemar variants with the headline named, the split of the cases they disagreed on, and both reviewers' approval rates on identical files. A lopsided split means one reviewer holds a looser bar, which is a calibration problem rather than a competence one.
See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.
Run any analysis on your own data: validated R analyses, interactive reports, AI insights, and PDF export.
Try Free, No Credit CardTell us what went wrong, in your own words. We capture the page you're on automatically, so no need to describe where you are.