Upload a CSV, map an outcome, a confounded treatment, and an instrument. You get the first-stage F and the instrument-strength verdict, the ordinary and instrumented estimates side by side with the difference between them, an over-identification test, and the exclusion restriction stated plainly as an assumption. Free.
Free analyses run on up to 10,000 rows. Larger files are randomly sampled to that size — sign up to analyze your full dataset.
Fitting both stages and testing the instrument...
Sent to — first-stage F and instrument-strength verdict, ordinary versus instrumented estimates with correct 2SLS standard errors, over-identification and endogeneity tests, R code, and AI insights.
Analyze another fileThe first stage regresses the treatment on the instruments plus any controls with base least squares, and the F statistic comparing that fit against the same regression WITHOUT the instruments measures instrument strength on the excluded instruments alone. The second stage regresses the outcome on the fitted treatment from the first stage plus the same controls. The reported standard error is the two-stage least squares standard error: the second-stage regression's own figure is wrong because it measures residuals against the fitted treatment, so the residual variance is rebuilt from the structural equation — the outcome minus the fitted coefficients applied to the ACTUAL treatment — and used to scale the second-stage cross-product inverse. An ordinary least squares regression is run alongside for contrast, a Sargan over-identification test is computed when there are more excluded instrument parameters than confounded regressors, and a Wu-Hausman test asks whether the two estimates differ by more than sampling noise.
Use it when the treatment was not randomly assigned, you believe something unmeasured moves both the treatment and the outcome, and you have a variable that plausibly shifted the treatment for reasons unrelated to the outcome — a lottery, a distance, an eligibility cutoff, an administrative rule, a supply shock.
Not when you have no credible instrument, in which case matching or a difference-in-differences design may fit better. Not when the instrument fails the first-stage strength check. Not for more than one confounded treatment column at a time, and not when the design already gives you a comparison group directly, such as a randomized experiment or a sharp cutoff.
Built for: Economists, policy analysts, data scientists, and researchers estimating causal effects from observational data
Typical data source: One row per unit with an outcome, a treatment the unit chose or was sorted into, and at least one column that nudged the treatment for an outside reason
One row per unit, with an outcome, a confounded treatment, at least one instrument, and any controls. For example, returns to schooling instrumented by distance to the nearest college:
Minimum 50 rows · Best with 300-10,000 rows, one confounded treatment, one to three instruments
Standard-library analysis: what is the causal effect when the treatment is confounded? Map an outcome, a confounded treatment, and at least one instrument — a variable that shifts the treatment but has no other route to the outcome — and get two-stage least squares built stage by stage: the first-stage F on the excluded instruments and the instrument-strength verdict it implies, the ordinary least squares estimate and the instrumented estimate side by side with the difference between them made explicit, an over-identification (Sargan) test when more than one instrument is supplied, an endogeneity (Wu-Hausman) test, and a plain statement of the exclusion restriction the whole estimate rests on.
The first-stage F on the excluded instruments — the single number that decides whether the instrumented estimate may be read at all.
The first-stage relationship as a picture: a visible slope is relevance, a flat cloud is a weak instrument.
The confounded estimate next to the instrumented one — the distance between the bars is what the confounding was doing.
Both estimates with intervals and p-values, using the correct two-stage least squares standard error rather than the second stage's own.
What was tested versus what was assumed — including the exclusion restriction, which no statistic can check.
Plain-English interpretation — what the numbers mean, what's significant, and what to do next.
The plain regression says one thing — but people chose the treatment themselves
Map the outcome, the treatment, and an instrument that nudged the treatment without touching the outcome any other way. You get the ordinary estimate and the instrumented estimate side by side; the gap between them is the size of what selection was adding to the ordinary number.
See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.
Run any analysis on your own data — validated R analyses, interactive reports, AI insights, and PDF export.
Try Free — No Credit CardTell us what went wrong, in your own words. We capture the page you're on automatically, so no need to describe where you are.