You just validated your new blood pressure monitor against the hospital's gold standard device. You measured 40 patients with both methods, calculated correlation (r = 0.94), and declared victory. Your device agrees with the standard.

Wrong. Dead wrong.

High correlation doesn't mean the methods agree—it means they move together. Your new device could systematically read 20 mmHg higher than the gold standard and still show perfect correlation. That's not agreement. That's disaster waiting for clinical deployment.

Before you draw conclusions about whether two measurement methods can be used interchangeably, you need to check the experimental design. Correlation is interesting. Agreement requires a different analysis entirely: Bland-Altman.

Quick Win: Spot Bias in 60 Seconds

Plot differences against averages. If the cloud of points centers far from zero, you have systematic bias. If the spread widens as measurements increase, you have proportional bias. Both patterns invalidate simple interchangeability claims—and you can see them before you even calculate limits of agreement.

Why Method Comparison Fails Without Proper Agreement Analysis

Here's the problem with how most people compare measurement methods: they use the wrong statistical test. They calculate Pearson correlation or run a t-test and call it done. Both approaches miss the point entirely.

A t-test tells you if the mean values differ between methods. But identical means don't guarantee agreement—one method could consistently overestimate in some ranges and underestimate in others, averaging out to no difference. Useless for deciding if you can swap methods.

Correlation tells you if two methods move in the same direction. Two methods could have correlation of 0.99 but differ by 50 units on every single measurement. Perfect correlation, zero agreement.

What you actually need to know: How much do individual measurements differ between methods, and is that difference clinically or practically acceptable?

That's what Bland-Altman analysis answers. It was introduced by Martin Bland and Douglas Altman in 1986 specifically because researchers kept using correlation to assess agreement—and making dangerous conclusions as a result.

The Bland-Altman approach is deceptively simple: calculate the difference between paired measurements (Method A - Method B), plot those differences against the average of the two methods, and establish limits of agreement that capture where 95% of differences fall. If those limits are narrow enough for your purposes, the methods agree. If not, they don't.

The Three Mistakes That Invalidate Your Limits of Agreement

Before we get to the mechanics, let's address the common pitfalls that make Bland-Altman analyses worthless. These mistakes show up in published papers, regulatory submissions, and method validation reports. They're common, they're subtle, and they completely undermine your conclusions.

Mistake 1: Ignoring Proportional Bias

Standard Bland-Altman assumes the difference between methods stays constant across the measurement range. Plot your differences against averages—if you see a systematic trend (differences getting larger or smaller as the average increases), you have proportional bias.

When proportional bias exists, you can't use fixed limits of agreement. The limits need to vary with the magnitude of measurement. Solution: regress the differences on the averages and construct regression-based limits of agreement. This isn't optional—using fixed limits when proportional bias exists gives you wrong answers.

Quick diagnostic: If the correlation between differences and averages exceeds |0.1|, check for proportional bias formally. If significant, switch to regression-based LoA.

Mistake 2: Trusting the 1.96 Multiplier With Non-Normal Differences

The limits of agreement formula (mean difference ± 1.96 × SD) assumes the differences follow a normal distribution. If they don't—if you have outliers, skewness, or heavy tails—the 1.96 multiplier doesn't capture 95% of the data.

Check normality of differences with a histogram or Q-Q plot before you calculate limits. If you see clear departure from normality, you have three options:

  • Transform the data: Log transformation often helps with right-skewed differences (common with ratios or concentrations)
  • Use non-parametric limits: Calculate the 2.5th and 97.5th percentiles directly instead of mean ± 1.96 × SD
  • Report it honestly: State that differences aren't normally distributed and interpret limits accordingly

What you can't do: ignore the problem and report standard LoA anyway. That's methodological malpractice.

Mistake 3: Confusing Statistical Agreement With Clinical Agreement

This is the most insidious mistake because it's not a statistical error—it's a logical one. Bland-Altman tells you where 95% of differences fall. It does not tell you if that range is acceptable.

If your limits of agreement for glucose measurements are ±15 mg/dL, is that good enough? Statistics can't answer that. The clinical team has to decide if a potential 15 mg/dL difference between methods matters for treatment decisions. If you're titrating insulin doses, maybe ±5 mg/dL is the maximum acceptable difference. If you're screening for diabetes, maybe ±20 mg/dL is fine.

Define your acceptable difference threshold before you run the analysis. Then compare your calculated LoA to that threshold. If your LoA fall within the acceptable range, the methods agree for your purpose. If not, they don't—regardless of how "tight" the limits look statistically.

Common Pitfall: The Correlation Trap

If you catch yourself calculating Pearson correlation between two methods and using it to claim agreement, stop immediately. High correlation means the methods move together. It says nothing about bias or interchangeability. A method that systematically reads 50% higher than the gold standard can have correlation of 1.0. Use correlation to assess relationship strength, not agreement.

How to Set Up a Proper Method Comparison Experiment

Before you analyze anything, you need data from a well-designed comparison study. Here's how to set it up so your Bland-Altman analysis actually means something.

Sample Size: How Many Paired Measurements?

Minimum 30-50 paired observations for stable limits of agreement. Below 30, your confidence intervals on the LoA become so wide they're nearly useless. For detecting proportional bias reliably, aim for 50+ measurements spread across the full range of clinically relevant values.

If you're comparing methods across different conditions (e.g., different patient populations, different measurement ranges, different operators), you need adequate sample size within each subgroup. Pooling heterogeneous data and running one Bland-Altman analysis hides important variation.

Measurement Protocol: Eliminate Extraneous Variation

The only difference between your paired measurements should be the method itself. Everything else—timing, subject state, operator, environmental conditions—should be as similar as possible.

Best practice: measure with both methods in quick succession (to minimize biological variation), randomize the order (to avoid systematic order effects), and use the same operator when possible (to eliminate inter-operator variation from your agreement assessment).

If you're comparing a new automated method against a manual gold standard, and you have your junior technician run the automated method while your senior tech runs the gold standard, you're confounding method differences with operator differences. Don't do that.

Measurement Range: Cover the Full Clinical Spectrum

Your limits of agreement are only valid for the measurement range you studied. If you validate a glucose meter using samples from 70-140 mg/dL (normal range) but then use it for diabetic patients with values up to 400 mg/dL, you're extrapolating beyond your validation data.

Proportional bias often emerges at the extremes of the measurement range. Make sure your comparison study includes enough low-range and high-range measurements to detect it. A sample of 50 measurements clustered around the midpoint is less useful than 50 measurements spanning the full range.

Calculating and Interpreting Limits of Agreement: A Worked Example

Let's work through a real example: validating a new portable spirometer against a clinical-grade device. We measured forced vital capacity (FVC) in 45 patients with both devices, randomizing the order of measurements.

Step 1: Calculate the Differences and Averages

For each patient, calculate:

  • Difference = New method - Gold standard
  • Average = (New method + Gold standard) / 2

The choice of which method to subtract from which is arbitrary—it just determines the sign of your bias. Consistency matters. If you define difference as New - Gold, a positive mean difference means your new method reads higher.

Step 2: Check for Systematic Bias

Calculate the mean difference. In our spirometer example, mean difference = -0.08 L. The new device reads, on average, 80 mL lower than the gold standard. Is that acceptable? Depends on your clinical tolerance—but at least now you know it exists.

A mean difference of zero would indicate no systematic bias (methods agree on average). Non-zero mean indicates fixed bias—one method consistently reads higher or lower than the other.

Step 3: Check for Proportional Bias

Plot differences against averages. Do the differences show a trend? In our case, the correlation between differences and averages is r = -0.03 (basically zero), and visual inspection shows no trend. No proportional bias detected.

If we had seen proportional bias, we'd need to fit a regression (differences ~ averages) and calculate regression-based limits. Since we don't, we proceed with standard limits.

Step 4: Check Normality of Differences

Plot a histogram of the differences or run a Shapiro-Wilk test. In our spirometer data, differences are approximately normal (no severe outliers, no skewness). We're good to use the parametric formula.

Step 5: Calculate Limits of Agreement

Standard deviation of differences = 0.22 L. Our limits of agreement are:

  • Upper LoA = mean difference + 1.96 × SD = -0.08 + 1.96(0.22) = +0.35 L
  • Lower LoA = mean difference - 1.96 × SD = -0.08 - 1.96(0.22) = -0.51 L

Interpretation: 95% of measurements will differ by between -0.51 L and +0.35 L between the new spirometer and the gold standard. Most of the time the new device reads lower (negative bias), but occasionally it could read up to 0.35 L higher.

Step 6: Compare to Clinical Acceptability

Is ±0.5 L acceptable for clinical FVC measurements? If you're screening for obstructive lung disease, probably yes—that's within typical test-retest variability. If you're making decisions about lung transplant eligibility based on precise FVC cutoffs, probably not.

This is where clinical expertise enters. The statistics show you the range of agreement. You decide if that range is good enough.

Try It Yourself: Free Bland-Altman Analysis Tool

Upload your CSV with paired measurements and get instant Bland-Altman plots, limits of agreement, bias assessment, and proportional bias diagnostics. No statistics background required—the tool checks assumptions and flags common mistakes automatically.

Run Free Bland-Altman Analysis

What the Bland-Altman Plot Actually Shows You

The plot itself is where the value lives. You're plotting difference (y-axis) against average (x-axis) for each paired measurement, with horizontal lines showing the mean difference and limits of agreement.

Here's what to look for—these are the quick-win visual diagnostics that catch problems before you even report numbers:

Pattern 1: Horizontal Band of Points

Ideal scenario. Points scattered randomly around the mean difference with constant spread across the measurement range. This means:

  • No proportional bias (spread doesn't change with magnitude)
  • Consistent agreement across the range
  • Standard limits of agreement are appropriate

If you see this pattern, and your LoA are within acceptable clinical bounds, you're done. The methods agree.

Pattern 2: Cloud Centered Far From Zero

Points cluster around a mean difference well above or below zero. This indicates systematic (fixed) bias. One method consistently reads higher or lower than the other across all measurements.

Fixed bias doesn't necessarily mean the methods don't agree—it just means you need to apply a correction factor. If Method A always reads 5 units higher than Method B, you can subtract 5 from Method A readings and restore agreement. The question is whether that's practical in your context.

Pattern 3: Funnel or Trend

Differences increase or decrease systematically as the average increases. This is proportional bias—agreement deteriorates (or improves) depending on measurement magnitude.

You'll see this when one method has different scaling or when measurement error is proportional to the true value. Standard LoA don't work here. You need regression-based limits that widen or narrow across the range.

Pattern 4: Outliers or Clusters

One or more points sitting far outside the main cloud. Investigate these before you calculate final limits:

  • Measurement error (transcription mistake, device malfunction)?
  • Real biological variation (patient moved, coughed, had acute event)?
  • Subgroup with different agreement characteristics?

Don't automatically exclude outliers, but understand what they represent. If they're genuine measurement failures, remove them. If they're real biological variation, keep them—they're part of the agreement picture.

Advanced Scenarios: When Standard Bland-Altman Isn't Enough

The basic Bland-Altman approach works for simple paired measurements with no complications. Real-world method comparison studies are rarely that clean. Here's how to handle common complications.

Repeated Measurements Per Subject

Standard Bland-Altman assumes independent observations. If you have multiple measurements per subject (e.g., you measured each patient three times with each method), you're violating that assumption. Observations from the same subject are correlated.

Two solutions:

  1. Average the replicates: Calculate the mean of the three measurements for each method, then run Bland-Altman on the averaged values. Simple, but throws away information about within-subject variability.
  2. Use mixed-effects Bland-Altman: Account for within-subject correlation explicitly using random effects. More complex, but gives you proper confidence intervals on the LoA.

What you can't do: treat all measurements as independent. That inflates your sample size artificially and makes your limits too narrow.

Comparison Across Multiple Methods

You're comparing three or four different measurement devices simultaneously. Don't run all pairwise Bland-Altman comparisons and call it comprehensive—you'll make multiple comparisons without adjustment, inflating your false positive rate.

Better approach: designate one method as the reference (ideally the gold standard) and compare all other methods against that single reference. This keeps your comparisons independent and interpretable.

Methods With Different Units or Scales

Sometimes you're comparing methods that measure the same construct but report in different units (e.g., glucose in mg/dL vs mmol/L, or blood pressure from an invasive catheter vs non-invasive cuff).

Convert to common units before running Bland-Altman. If one method measures on a ratio scale and the other on an interval scale, you may need to analyze percent differences instead of absolute differences.

When Bland-Altman Analysis Fails (And What to Use Instead)

Bland-Altman is powerful for continuous measurements with paired observations. It's not universal. Here's when it breaks down and what to reach for instead.

When You Don't Have Paired Measurements

Bland-Altman requires paired measurements—you measured the same subject/sample with both methods. If your two methods were used on different subjects or samples, you can't calculate individual differences. You're stuck comparing population distributions instead.

Alternative: Two-sample t-test for mean differences, or non-parametric equivalents like Mann-Whitney if distributions aren't normal. But understand you've lost the ability to assess individual-level agreement.

When Measurements Are Categorical

Bland-Altman works for continuous outcomes. If you're comparing two methods that produce categorical results (e.g., two radiologists classifying tumors as benign/malignant), you need agreement statistics for categorical data: Cohen's kappa or weighted kappa.

When the Gold Standard Is Imperfect

Traditional Bland-Altman treats both methods symmetrically—neither is assumed to be the "true" value. If you actually have a gold standard (one method is known to be more accurate), you might want error-based assessment instead of agreement-based.

Consider plotting (New method - Gold standard) directly, rather than differences against averages. This shows you prediction error explicitly and makes calibration easier to assess.

When You Need to Account for Covariates

Agreement might vary by patient characteristics (age, disease severity, operator experience). Standard Bland-Altman pools everyone together, hiding subgroup differences.

Solution: Stratified analysis (separate Bland-Altman plots for each subgroup) or regression-based approaches that model agreement as a function of covariates. More complex, but necessary when agreement isn't constant across your population.

Quick Diagnostic Checklist

Before you finalize your Bland-Altman analysis, verify:

  • ☑ Paired measurements from same subjects/samples
  • ☑ Sample size ≥30 (preferably ≥50)
  • ☑ Measurements span full clinical range
  • ☑ Differences approximately normal (or transformed)
  • ☑ No proportional bias (or regression-based LoA used)
  • ☑ No repeated measures (or properly accounted for)
  • ☑ LoA compared to pre-specified acceptable difference

Real-World Example: Where Bland-Altman Catches What Correlation Misses

A medical device company developed a fingertip pulse oximeter for home use. They validated it against a hospital-grade device in 60 patients. Correlation between the two devices: r = 0.92. They submitted to regulators claiming substantial agreement.

Regulators requested Bland-Altman analysis. Here's what it revealed:

Mean difference: -1.2% oxygen saturation (home device reads systematically lower). Limits of agreement: -5.3% to +2.9%. That means 95% of home device readings could be up to 5.3% lower than the hospital device.

Clinical context: In the 90-100% saturation range (normal), a 5% difference might be tolerable. But in the 85-90% range (hypoxemic), a 5% underestimate means the home device could read 85% when true saturation is 90%—potentially delaying critical medical intervention.

The correlation was excellent. The agreement was inadequate for the intended use. Bland-Altman caught it; correlation didn't. The device went back for recalibration.

This is why you run Bland-Altman. Not because the formula is complicated (it isn't), but because it answers the right question: Can I use these methods interchangeably for individual patients?

How to Report Bland-Altman Results (So Reviewers Don't Reject Your Paper)

You've run the analysis, checked assumptions, and confirmed your methods agree (or don't). Now you need to report it in a way that passes peer review and regulatory scrutiny. Here's what to include.

Essential Elements to Report

  1. Sample size and study design: How many paired observations? How were measurements collected (randomized order, same operator, etc.)?
  2. Mean difference and 95% confidence interval: This is your bias estimate with uncertainty
  3. Limits of agreement (both limits) with 95% CIs: Yes, the LoA themselves have confidence intervals—report them
  4. Assessment of proportional bias: Report the correlation between differences and averages, and whether you used regression-based LoA
  5. Normality assessment: State whether differences were normally distributed and what you did if they weren't
  6. Clinical interpretation: Compare your LoA to your pre-specified acceptable difference and state whether agreement is adequate for your purpose

The Bland-Altman Plot Itself

Include the plot in your results. It should show:

  • All data points (difference vs average)
  • Horizontal line at mean difference
  • Horizontal lines at upper and lower LoA
  • Optionally, dashed lines showing 95% CIs on the LoA
  • Clear axis labels with units

Make the plot readable. Don't cram 500 overlapping points into a tiny figure. Use semi-transparent points if you have many observations.

What Not to Report

Don't report Pearson correlation between the two methods as evidence of agreement. Reviewers who know Bland-Altman will flag this immediately as misunderstanding the purpose of the analysis.

Don't report only the limits of agreement without the mean difference. The mean tells you about fixed bias; the limits tell you about random disagreement. You need both.

Don't claim "excellent agreement" because your LoA look narrow without defining what "acceptable" means in your context. Let the reader judge whether ±10 units is acceptable for their application.

Frequently Asked Questions

When should I use Bland-Altman analysis instead of correlation?

Use Bland-Altman analysis when you need to assess whether two measurement methods can be used interchangeably. Correlation tells you if methods move together (useless for agreement), while Bland-Altman shows systematic bias and the range of likely differences. If you're validating a new device against a gold standard or comparing lab techniques, you need Bland-Altman, not correlation.

What are the limits of agreement and how do I interpret them?

The limits of agreement (LoA) are mean difference ± 1.96 × SD of differences, capturing the range where 95% of differences fall. If your LoA are ±10 mmHg for blood pressure, 95% of measurements will differ by that much or less between methods. Whether that's acceptable depends on your clinical or practical tolerance—statistics can't answer that for you.

What are the most common mistakes in Bland-Altman analysis?

The three critical mistakes: (1) Not checking for proportional bias—when differences grow with measurement magnitude, you need regression-based LoA. (2) Ignoring non-normality of differences—outliers or skewness violate the 1.96 multiplier assumption. (3) Using correlation to assess agreement—high correlation doesn't mean methods agree, it just means they move together.

How many measurements do I need for a valid Bland-Altman plot?

Minimum 30-50 paired measurements for stable limits of agreement. With fewer than 30, your confidence intervals on the LoA become too wide to be useful. For detecting proportional bias reliably, aim for 50+ across the full measurement range. If you're comparing methods across different patient groups or conditions, you need adequate sample size in each subgroup.

Can I use Bland-Altman analysis if my measurements have repeated readings per subject?

Standard Bland-Altman assumes independent observations. With repeated measures (multiple readings per subject), you need to account for within-subject correlation—either average the replicates first or use mixed-effects extensions of Bland-Altman. Ignoring this clustering inflates your precision estimates and makes your limits of agreement too narrow.

The Bottom Line: Agreement Requires More Than Moving Together

Here's what matters for method comparison: correlation tells you if two methods track together. Bland-Altman tells you if you can swap them.

If you're validating a new measurement device, comparing lab assays, or deciding if two clinical tests can be used interchangeably, you need Bland-Altman analysis. The plot shows you bias, the limits show you variability, and together they answer the question that actually matters: Will individual measurements from these methods differ by more than I can tolerate?

The method is simple. The interpretation requires domain knowledge. Statistics shows you the limits of agreement—you decide if they're acceptable.

Before you conclude that two methods agree, check the experimental design. Did you use paired measurements? Did you cover the full range? Did you check for proportional bias? Did you verify normality of differences? Did you compare your LoA to clinical acceptability?

If the answer to any of those is no, your agreement claim isn't supported. Fix the design, rerun the analysis, and report it honestly.

Want to run your own Bland-Altman analysis? Upload your CSV and get instant plots, bias assessment, limits of agreement, and assumption checks. The tool flags proportional bias automatically and warns you when assumptions are violated. You'll see exactly what patterns are in your data—and whether your methods actually agree.

Key Takeaway: The Fast Diagnosis Pattern

Plot differences against averages before you calculate anything. Cloud centered far from zero? Fixed bias. Cloud that widens or narrows? Proportional bias. Outliers sitting far from the pack? Investigate them. These visual patterns catch agreement failures instantly—and often reveal problems that summary statistics hide. The Bland-Altman plot isn't just a reporting requirement. It's your primary diagnostic tool.

Related Articles