Upload a CSV, map the outcome and the score your cutoff rule is applied to, and get the sharp regression discontinuity estimate with a confidence interval, the plot it comes from, the effective sample near the cutoff, and the three checks that decide whether it means anything. Free.
Free analyses run on up to 10,000 rows. Larger files are randomly sampled to that size — sign up to analyze your full dataset.
Fitting local linear regressions either side of the cutoff...
Sent to — The discontinuity estimate with a robust confidence interval, the RD plot with binned means and both fitted lines, manipulation and covariate-balance tests, bandwidth sensitivity, R code, and AI insights.
Analyze another fileThe cutoff is taken from the 'cutoff' module parameter, else read off the treatment column, else assumed and flagged loudly. The discontinuity is estimated by local linear regression fitted separately either side of the cutoff, weighting each row by a triangular kernel that falls from 1 at the cutoff to 0 at the edge of the bandwidth; the gap between the two fitted lines at the cutoff is the estimate, with a heteroskedasticity-robust (HC1) standard error and a t-based 95% interval. The bandwidth comes from an Imbens-Kalyanaraman plug-in rule implemented directly — pilot density and residual variance at the cutoff, a global cubic for the third derivative, side-specific quadratics for the curvature difference, and the IK regularisation terms, with the triangular-kernel constant 3.4375 — then widened or narrowed so at least 15 rows sit each side. Three diagnostics follow: a McCrary-style density test (fine histogram with the cutoff on a bin edge, triangular-kernel local linear fits of the bin frequencies either side extrapolated to the cutoff, log-difference with its asymptotic standard error); covariate balance, running each mapped pre-determined column through the same estimator at the same bandwidth; and bandwidth sensitivity, recomputing the estimate from half to double the chosen bandwidth. If a treatment column is supplied and treatment is not a deterministic function of the cutoff, the design is reported as fuzzy and the sharp estimate is withheld.
Use it whenever a rule on a numeric score, measure, or index decides who gets something — an eligibility threshold, a discount tier, an audit trigger, a class-size cap — and you want the causal effect of crossing that line without running an experiment.
Not when treatment was assigned by anything other than a threshold on a measurable running variable (use the group-comparison, propensity-matching, or difference-in-differences tools), not when the change happened at a point in TIME across everyone at once (use interrupted time series), and not when compliance with the cutoff is partial — that is a fuzzy design needing a two-stage estimator, which this analysis detects and declines rather than approximating.
Built for: Economists, policy analysts, growth and pricing teams, and researchers evaluating threshold rules they did not randomise
Typical data source: Any spreadsheet with one row per person, account, or order: the score or measure a rule is applied to, and what happened afterwards
One row per unit: the score the rule is applied to, whether the unit got the treatment, anything fixed beforehand, and the outcome:
Minimum 30 rows · Best with 500-20,000 rows, with at least a few hundred within reach of the cutoff
Standard-library analysis: did crossing the threshold cause the change? Map the outcome and the running variable a cutoff rule is applied to — an exam score, a revenue band, an eligibility index, a queue position — and get the sharp regression discontinuity estimate: local linear regression with a triangular kernel on each side of the cutoff at a data-driven bandwidth, the discontinuity with a robust 95% confidence interval, the scatter of outcome against the running variable with binned means and both fitted lines, the effective sample size actually near the cutoff, and the three diagnostics that decide whether the design is credible — a McCrary-style manipulation test, covariate balance at the cutoff, and bandwidth sensitivity.
Binned means of the outcome across the running variable with both local linear fits and the cutoff marked — the gap where the lines meet the cutoff is the estimate.
The cutoff, the bandwidth, the fitted level either side, the jump with its robust interval, and the effective sample the whole result rests on.
A density test for units sorting across the threshold — the failure that kills the design outright.
Pre-determined columns run through the same estimator; any of them jumping at the cutoff means the design is compromised.
The estimate and interval recomputed from half to double the chosen bandwidth — a finding should survive the whole range.
The design, the estimator, the bandwidth rule, what each diagnostic found, and the limits — including that the estimate applies at the cutoff only.
Plain-English interpretation — what the numbers mean, what's significant, and what to do next.
Did the eligibility rule actually change anything?
Map the outcome and the score the rule is applied to. You get the jump at the cutoff with a confidence interval, the picture that jump comes from, and the three checks that decide whether it means anything — whether people gamed the score, whether other characteristics jump at the line too, and whether the number survives a range of bandwidths.
See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.
Run any analysis on your own data — validated R analyses, interactive reports, AI insights, and PDF export.
Try Free — No Credit CardTell us what went wrong, in your own words. We capture the page you're on automatically, so no need to describe where you are.