Standard Bland Altman
Executive Summary

Executive Summary

Do Device Reading and Lab Reading agree well enough to be used interchangeably?

Pairs Analysed
150
Bias (Mean Difference)
2.025
Lower Limit of Agreement
-5.84
Upper Limit of Agreement
9.89
Within Limits
96.00%
Proportional Bias
not detected
Correlation r
0.98
Across 150 paired measurements, Device Reading reads on average 2.02 higher than Lab Reading (bias = 2.02, 95% CI 1.38 to 2.67 — a real systematic offset). The 95% limits of agreement run from -5.84 to 9.89: for a new measurement, the difference between the two methods is expected to fall in this range about 95% of the time. Here 96.00% of the observed differences fall inside (about 95% is expected by construction). No proportional bias was detected (slope -0.004, p = 0.810), so the limits apply across the measured range. The two methods correlate strongly (r = 0.980), yet they do not agree: correlation only shows the readings rise and fall together, while the bias of 2.02 is a real systematic offset — high correlation is not agreement. Whether these limits are acceptable depends on your clinical/technical tolerance: compare the limits of agreement against the largest disagreement you could tolerate in practice. The analysis describes how far apart the methods are — it cannot decide whether that is close enough for your use.
Suggested Interpretation

The short answer

The device reads 2.02 units higher on average—a real systematic offset (95% CI 1.38 to 2.67). The two methods are expected to differ by anywhere from -5.84 to 9.89 for a typical measurement, and 96.00% of observed pairs fell inside those limits. The disagreement does not grow with measurement size (p = 0.810). Despite strong correlation (r = 0.980), the systematic offset of 2.02 means the methods do not agree—high correlation alone is not agreement. Whether these limits are acceptable depends entirely on your tolerance.

The detail

Bias = 2.025 (95% CI 1.38 to 2.67, excludes zero). Lower limit of agreement = -5.84 (95% CI -6.95 to -4.73); upper limit = 9.89 (95% CI 8.78 to 11.00). Proportional bias check: slope -0.004, p = 0.810 (not significant). Coverage: 96.00% of 150 differences within limits (≈95% expected by construction). Pearson r = 0.980.

What this can't tell you

The analysis describes how far apart the methods are but cannot decide whether that distance is clinically or technically acceptable. You must set your own tolerance threshold and compare it against the limits of agreement.

Overview

Analysis Overview

Bland-Altman agreement between Device Reading and Lab Reading across 150 paired measurements.

N Pairs150
Bias2.025
Loa Low-5.84
Loa High9.89
Suggested Interpretation

The short answer

The analysis uses differences rather than correlation because correlation only shows whether two methods move together—they can track almost perfectly while one reads systematically higher. Here, both the bias (2.02) and the limits of agreement (-5.84 to 9.89) capture how far apart the methods actually are, which is what determines whether they can be used interchangeably.

The detail

Across 150 paired measurements, the analysis plots each pair's difference (Device Reading minus Lab Reading) against the pair's mean. The bias of 2.02 is the systematic offset; the 95% limits of agreement of -5.84 to 9.89 describe the expected range of disagreement for a typical new measurement. A regression then checks whether this disagreement is stable across the measured range.

What this can't tell you

The analysis cannot decide whether these limits are acceptable for your use. That is a tolerance question: you must compare the limits against the largest disagreement you could tolerate in practice.

Data Preparation

Data Quality

Pair completeness and measurement checks.

Initial Rows150
Final Rows150
Rows Removed0
Suggested Interpretation

The short answer

All 150 rows formed complete pairs with no missing values dropped. Both columns were checked for numeric validity and variation before analysis. Critically, both measurements must be on the same units—the analysis cannot detect unit mismatches from the data alone.

The detail

150 rows were loaded and 150 rows remained after preprocessing (0 rows removed). Both the Device Reading and Lab Reading columns were coerced to numbers and verified for variation before the differences were calculated.

What this can't tell you

If the two columns are on different units (e.g., one in mg/dL and one in mmol/L), the differences are meaningless and the analysis cannot detect that automatically. You must verify unit consistency independently.

Visualization

Bland-Altman Plot

Difference between Device Reading and Lab Reading plotted against their mean, with bias and limit lines.

Suggested Interpretation

The short answer

The plot shows each pair's difference against their mean, with the bias line at 2.02 and agreement limits at -5.84 and 9.89. The cloud is flat and even around the bias line—disagreement is stable across the measured range. 96.00% of points sit inside the limits, as expected by construction.

The detail

Each of the 150 points plots the mean of Device Reading and Lab Reading on the horizontal axis against their difference (Device Reading minus Lab Reading) on the vertical. The middle reference line marks the bias (2.0249); the outer lines mark the 95% limits (-5.8401 and 9.89). The absence of a funnel shape (where spread grows with magnitude) supports using raw rather than percentage differences. The even scatter around the bias line indicates the disagreement does not systematically worsen at higher or lower measurements.

What this can't tell you

Visual inspection of scatter cannot formally test whether the cloud is truly flat or whether apparent variation is consistent with random sampling. The proportional-bias regression (slope p = 0.810) provides that formal test.

Data Table

Agreement Statistics

Bias, SD of differences, and limits of agreement — each with a 95% CI.

MeasureEstimateCI LowCI HighInterpretation
Bias (mean difference)2.0251.3782.672On average Device Reading reads 2.02 higher than Lab Reading — the CI excludes zero, a real systematic offset.
SD of differences4.013The spread of the per-row disagreements between the two methods.
Lower limit of agreement-5.84-6.949-4.731For a new measurement, Device Reading minus Lab Reading is expected to stay above this about 97.5% of the time.
Upper limit of agreement9.898.78111For a new measurement, Device Reading minus Lab Reading is expected to stay below this about 97.5% of the time.
Suggested Interpretation

The short answer

The device reads 2.02 higher on average, and this offset is statistically real (95% CI 1.38 to 2.67 excludes zero). For a new measurement, the difference is expected to land between -5.84 and 9.89 about 95% of the time. The confidence intervals around each limit (-6.95 to -4.73 and 8.78 to 11.00) show the honest range of plausible limits given the sample size.

The detail

Bias (mean difference) = 2.025 (95% CI 1.378 to 2.672). SD of differences = 4.013. Lower limit of agreement = -5.84 (95% CI -6.949 to -4.731). Upper limit of agreement = 9.89 (95% CI 8.781 to 10.999). Each limit's CI reflects the classical Bland-Altman standard error and the uncertainty in estimating the true limits from 150 pairs.

What this can't tell you

Whether these limits are acceptable for interchangeable use. Acceptability is a tolerance question: compare the limits against the largest disagreement you could tolerate in practice. The analysis quantifies disagreement but cannot decide whether it is close enough for your application.

Data Table

Proportional Bias Check

Does the disagreement between the methods change with the size of the measurement?

TermEstimateStd ErrorP ValueInterpretation
Intercept2.4271.7010.156The expected difference between the methods at a (hypothetical) measurement of zero.
Slope (difference vs magnitude)-0.0040.01660.810Not significant (p = 0.810): no evidence that the disagreement grows or shrinks with the size of the measurement.
Suggested Interpretation

The short answer

The disagreement between the device and lab method does not change with measurement size. The slope of difference against magnitude is -0.004 (p = 0.810), not statistically significant. This means a single pair of limits applies across the entire measured range.

The detail

Regression of the difference on the mean measurement yields a slope of -0.004 (standard error 0.0166, p = 0.810). This non-significant result indicates no evidence that disagreement grows or shrinks as the measurement gets larger. The intercept (2.4267, p = 0.156) is not significantly different from the bias itself, supporting the use of a constant bias and fixed limits across the range.

What this can't tell you

A non-significant p-value does not prove the slope is exactly zero—only that the observed slope is consistent with no real relationship. With 150 pairs, the analysis has reasonable power to detect a proportional trend if one exists at a meaningful scale, but cannot rule out very small effects.

Data Table

Methods & Disclosure

How the bias, limits, and checks are computed, and what they can and cannot decide.

ItemDetail
Difference conventionEvery row's difference is Device Reading minus Lab Reading; the magnitude axis is the mean of the two readings.
Bias and its 95% CIBias = mean difference = 2.02; 95% CI 1.38 to 2.67 (t distribution, 149 degrees of freedom).
Limits of agreementBias plus or minus 1.96 times the SD of the differences (4.01): -5.84 to 9.89.
CIs for the limitsClassical Bland-Altman standard error for a limit, SD x sqrt(1/n + 1.96^2/(2(n-1))) = 0.56, giving lower limit CI -6.95 to -4.73 and upper limit CI 8.78 to 11.00.
Proportional-bias checkOrdinary regression of the difference on the mean of the two methods; slope -0.004 (p = 0.810).
Coverage check96.00% of the 150 observed differences fall inside the limits (about 95% is expected by construction).
Correlation vs agreementPearson r between Device Reading and Lab Reading is 0.980. Correlation measures whether the readings rise and fall together; agreement asks how far apart they are. High correlation is not agreement.
AcceptabilityWhether these limits are acceptable depends on your clinical/technical tolerance: compare the limits of agreement against the largest disagreement you could tolerate in practice. The analysis describes how far apart the methods are — it cannot decide whether that is close enough for your use.
Suggested Interpretation

The short answer

The bias is the mean of the per-row differences (Device Reading minus Lab Reading), computed with a t-based 95% confidence interval. The limits of agreement are bias ± 1.96 × SD of differences (4.01), each with its own CI. A regression checks whether disagreement changes with magnitude. Pearson r = 0.980 is reported to show that high correlation is not agreement.

The detail

Difference convention: Device Reading minus Lab Reading for every row. Bias = 2.02, 95% CI 1.38 to 2.67 (t distribution, 149 df). Limits of agreement = 2.02 ± 1.96 × 4.013 = -5.84 to 9.89. Standard error for each limit = SD × √(1/n + 1.96²/(2(n−1))) = 0.56, giving lower CI -6.95 to -4.73 and upper CI 8.78 to 11.00. Proportional-bias regression: slope -0.004, p = 0.810. Coverage: 96.00% of 150 differences inside limits. Pearson r = 0.980 (co-movement, not agreement).

What this can't tell you

Acceptability depends on your clinical or technical tolerance. The analysis measures disagreement but cannot decide whether the limits are close enough for your use.

Methodology

Methodology

Statistical methodology and diagnostics for Method Agreement — Bland-Altman

Statistical Method

Method Agreement — Bland-Altman

Standard-library analysis: do two measurement methods agree well enough to be used interchangeably? Map two numeric columns measuring the same thing on the same units — a new device against a reference, two instruments, two assays, two raters scoring a continuous quantity — and get the Bland-Altman plot with bias and limit lines, the bias (mean difference) with a 95% confidence interval, the 95% limits of agreement each with its own confidence interval, a proportional-bias check, the observed share of points inside the limits, and the correlation coefficient explicitly contrasted with agreement.

Data
N = 150 observations
Assumptions
  • Each row is one subject or sample measured once by each method, on the same units
  • Both measurements are numeric or cleanly convertible
  • The differences are approximately normally distributed for the 95% limits to have their nominal coverage
  • The disagreement is roughly constant across the measured range — the proportional-bias check tests this
Limitations
  • The analysis quantifies agreement but cannot judge whether it is acceptable — that depends on the user's clinical/technical tolerance
  • Rows missing either method's value are dropped as incomplete pairs
  • At least 10 complete pairs are required, and limit estimates from small samples carry wide confidence intervals
  • When proportional bias is detected, the simple limits of agreement mislead — regression-based limits or percentage differences are then the better tool
Software & Citation
MCP Analytics · mcpanalytics.ai
Code Appendix

Analysis Code

Complete R source code for this analysis

Method Agreement — Bland-Altman

Do two measurement methods agree well enough to be used interchangeably? Each row is one subject or sample measured by both methods on the same units. The analysis works on the per-row difference between the methods: the bias (mean difference) with a 95% confidence interval, the 95% limits of agreement (bias ± 1.96·SD of the differences) each with its own confidence interval, a proportional-bias check (regression of the difference on the magnitude), and the share of points inside the limits.

Why This Method?

Correlation cannot answer an agreement question: two methods can correlate almost perfectly while one reads systematically higher than the other. Bland-Altman analysis instead describes how far apart the two methods are expected to be for a typical measurement — a bias plus a range — which is the quantity a practitioner can actually judge against a tolerance.

What This Analysis Covers

  • The Bland-Altman scatter (difference vs mean, with bias and limit lines)
  • Bias, SD of differences, and limits of agreement, each with 95% CIs
  • A proportional-bias regression (does the disagreement grow with magnitude?)
  • The observed share of points inside the limits, and the correlation r

explicitly contrasted with agreement

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {method_1, method_2}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed  $ row accounting
  #   n / n_dropped        $ complete pairs used + incomplete rows dropped
  #   m1_h / m2_h          $ humanized user names for the two methods
  #   bias / sd_d          $ mean difference (method_1 - method_2) + its SD
  #   bias_lo / bias_hi    $ 95% CI for the bias
  #   loa_low / loa_high   $ limits of agreement = bias -/+ 1.96*sd_d
  #   loa_low_lo/loa_low_hi/loa_high_lo/loa_high_hi  $ 95% CIs for each limit
  #   pct_within           $ observed % of points inside the limits
  #   r_methods            $ Pearson r between the two methods
  #   slope/slope_se/slope_p/intercept  $ regression of difference on mean
  #   prop_bias            $ TRUE when the slope is significant at p<0.05
  #   bias_sig             $ TRUE when the bias CI excludes zero
  #   high_r               $ TRUE when r_methods >= 0.8
  #   contrast_fires       $ high_r AND bias_sig — the correlation-vs-agreement teaching point
  #   dir_word             $ "higher"/"lower" consistent with the sign of bias
  #   ba_points_df         $ mean_value, difference (<=1000 sampled)
  #   agreement_df         $ measure, estimate, ci_low, ci_high, interpretation
  #   prop_bias_df         $ term, estimate, std_error, p_value, interpretation
  #   methods_df           $ item, detail
  #   tolerance_note       $ the "acceptability is the user's call" sentence
  #   metrics / json_output
  # === /SHARED EXPORTS ===

Step 1: Resolve mapped columns (humanized for all prose)

initial_rows <- nrow(df)
  m1_h <- humanize_semantic("method_1", col_map)
  m2_h <- humanize_semantic("method_2", col_map)
  if (!("method_1" %in% names(df)) || !("method_2" %in% names(df))) {
    stop(sprintf("Bland-Altman agreement needs both &#x27;%s' (the first method) and '%s' (the second method) mapped.",
                 m1_h, m2_h))
  }

Step 2: Coerce both measurements to numeric (95% rule)

coerce_num <- function(v, label_h) {
    if (is.numeric(v)) return(v)
    ch <- as.character(v)
    non_blank <- !is.na(ch) & trimws(ch) != ""
    conv <- suppressWarnings(as.numeric(ch))
    if (sum(non_blank) == 0 ||
        sum(!is.na(conv[non_blank])) < 0.95 * sum(non_blank)) {
      stop(sprintf("The column &#x27;%s' does not look numeric — fewer than 95%% of its values parse as numbers. Pick a numeric column.",
                   label_h))
    }
    conv
  }
  v1 <- coerce_num(df$method_1, m1_h)
  v2 <- coerce_num(df$method_2, m2_h)

Step 3: Keep row-wise complete pairs; require enough of them

keep <- !is.na(v1) & !is.na(v2)
  n_dropped <- sum(!keep)
  v1 <- v1[keep]; v2 <- v2[keep]
  n <- length(v1)
  final_rows <- n
  rows_removed <- initial_rows - final_rows
  if (n < 10) {
    stop(sprintf("Only %d complete %s / %s pairs remained after dropping incomplete rows — at least 10 are needed for a Bland-Altman agreement analysis.",
                 n, m1_h, m2_h))
  }

Step 4: Guard degenerate inputs (constant columns, identical differences)

if (isTRUE(stats::var(v1) == 0)) {
    stop(sprintf("The column &#x27;%s' is constant — every value is identical — so method agreement cannot be assessed. Map a measurement that varies.", m1_h))
  }
  if (isTRUE(stats::var(v2) == 0)) {
    stop(sprintf("The column &#x27;%s' is constant — every value is identical — so method agreement cannot be assessed. Map a measurement that varies.", m2_h))
  }

Step 5: Differences, bias, and its 95% CI (classical formulas)

d  <- v1 - v2
  mn <- (v1 + v2) / 2
  bias <- mean(d)
  sd_d <- stats::sd(d)
  if (is.na(sd_d) || sd_d == 0) {
    stop(sprintf("&#x27;%s' and '%s' differ by exactly the same amount on every row, so the limits of agreement collapse to a point — there is no variation in the differences to analyse.",
                 m1_h, m2_h))
  }
  tq <- stats::qt(0.975, n - 1)
  se_bias <- sd_d / sqrt(n)
  bias_lo <- bias - tq * se_bias
  bias_hi <- bias + tq * se_bias

Step 6: Limits of agreement, each with its own 95% CI

loa_low  <- bias - 1.96 * sd_d
  loa_high <- bias + 1.96 * sd_d
  se_loa <- sd_d * sqrt(1 / n + 1.96^2 / (2 * (n - 1)))
  loa_low_lo  <- loa_low  - tq * se_loa
  loa_low_hi  <- loa_low  + tq * se_loa
  loa_high_lo <- loa_high - tq * se_loa
  loa_high_hi <- loa_high + tq * se_loa

Step 7: Observed coverage of the limits (should be near 95%)

n_within <- sum(d >= loa_low & d <= loa_high)
  pct_within <- 100 * n_within / n

Step 8: Correlation between the methods (to contrast with agreement)

r_methods <- suppressWarnings(tryCatch(stats::cor(v1, v2), error = function(e) NA_real_))

Step 9: Proportional-bias check — regression of difference on mean

slope <- NA_real_; slope_se <- NA_real_; slope_p <- NA_real_
  intercept <- NA_real_; intercept_se <- NA_real_; intercept_p <- NA_real_
  if (isTRUE(stats::var(mn) > 0)) {
    fit <- tryCatch(stats::lm(d ~ mn), error = function(e) NULL)
    if (!is.null(fit)) {
      cf <- tryCatch(summary(fit)$coefficients, error = function(e) NULL)
      if (!is.null(cf) && nrow(cf) == 2) {
        intercept    <- cf[1, 1]; intercept_se <- cf[1, 2]; intercept_p <- cf[1, 4]
        slope        <- cf[2, 1]; slope_se     <- cf[2, 2]; slope_p     <- cf[2, 4]
      }
    }
  }
  prop_bias <- !is.na(slope_p) && slope_p < 0.05

Step 10: Verdict flags — all direction/consistency language computed

bias_sig <- !is.na(bias_lo) && !is.na(bias_hi) && (bias_lo > 0 || bias_hi < 0)
  high_r <- !is.na(r_methods) && r_methods >= 0.8
  contrast_fires <- high_r && bias_sig
  dir_word <- if (bias > 0) "higher" else if (bias < 0) "lower" else "the same on average"
  tolerance_note <- paste0(
    "Whether these limits are acceptable depends on your clinical/technical tolerance: ",
    "compare the limits of agreement against the largest disagreement you could tolerate ",
    "in practice. The analysis describes how far apart the methods are — it cannot decide ",
    "whether that is close enough for your use."
  )

Step 11: Bland-Altman scatter dataset — <=1000 sampled, fixed seed

set.seed(42)
  sidx <- if (n > 1000) sample(n, 1000) else seq_len(n)
  ba_points_df <- data.frame(
    mean_value = round(mn[sidx], 4),
    difference = round(d[sidx], 4),
    stringsAsFactors = FALSE
  )
  ba_points_df <- ba_points_df[order(ba_points_df$mean_value), , drop = FALSE]
  rownames(ba_points_df) <- NULL

Step 12: Agreement table — bias, SD, limits, all with CIs

agreement_df <- data.frame(
    measure = c("Bias(mean difference)", "SD of differences",
                "Lower limit of agreement", "Upper limit of agreement"),
    estimate = round(c(bias, sd_d, loa_low, loa_high), 3),
    ci_low  = round(c(bias_lo, NA_real_, loa_low_lo,  loa_high_lo), 3),
    ci_high = round(c(bias_hi, NA_real_, loa_low_hi,  loa_high_hi), 3),
    interpretation = c(
      sprintf("On average %s reads %s %s than %s%s.",
              m1_h,
              if (bias == 0) "" else r2(abs(bias)),
              if (bias >= 0) "higher" else "lower", m2_h,
              if (bias_sig) " — the CI excludes zero, a real systematic offset"
              else " — the CI includes zero, so no systematic offset is established"),
      "The spread of the per-row disagreements between the two methods.",
      sprintf("For a new measurement, %s minus %s is expected to stay above this about 97.5%% of the time.", m1_h, m2_h),
      sprintf("For a new measurement, %s minus %s is expected to stay below this about 97.5%% of the time.", m1_h, m2_h)
    ),
    stringsAsFactors = FALSE
  )

Step 13: Proportional-bias table

prop_bias_df <- data.frame(
    term = c("Intercept", "Slope(difference vs magnitude)"),
    estimate  = round(c(intercept, slope), 4),
    std_error = round(c(intercept_se, slope_se), 4),
    p_value   = c(fmt_p(intercept_p), fmt_p(slope_p)),
    interpretation = c(
      "The expected difference between the methods at a(hypothetical) measurement of zero.",
      if (prop_bias)
        sprintf("Significant(%s): the disagreement between %s and %s changes with the size of the measurement — agreement varies with magnitude and the simple limits mislead.",
                fmt_pp(slope_p), m1_h, m2_h)
      else if (!is.na(slope_p))
        sprintf("Not significant(%s): no evidence that the disagreement grows or shrinks with the size of the measurement.",
                fmt_pp(slope_p))
      else
        "The regression could not be computed(no variation in the measurement magnitude)."
    ),
    stringsAsFactors = FALSE
  )

Step 14: Methods table

methods_df <- data.frame(
    item = c("Difference convention", "Bias and its 95% CI", "Limits of agreement",
             "CIs for the limits", "Proportional-bias check", "Coverage check",
             "Correlation vs agreement", "Acceptability"),
    detail = c(
      sprintf("Every row&#x27;s difference is %s minus %s; the magnitude axis is the mean of the two readings.", m1_h, m2_h),
      sprintf("Bias = mean difference = %s; 95%% CI %s to %s(t distribution, %d degrees of freedom).",
              r2(bias), r2(bias_lo), r2(bias_hi), n - 1),
      sprintf("Bias plus or minus 1.96 times the SD of the differences(%s): %s to %s.",
              r2(sd_d), r2(loa_low), r2(loa_high)),
      sprintf("Classical Bland-Altman standard error for a limit, SD x sqrt(1/n + 1.96^2/(2(n-1))) = %s, giving lower limit CI %s to %s and upper limit CI %s to %s.",
              r2(se_loa), r2(loa_low_lo), r2(loa_low_hi), r2(loa_high_lo), r2(loa_high_hi)),
      sprintf("Ordinary regression of the difference on the mean of the two methods; slope %s(%s).",
              r3(slope), fmt_pp(slope_p)),
      sprintf("%s%% of the %s observed differences fall inside the limits(about 95%% is expected by construction).",
              r2(pct_within), format(n, big.mark = ",")),
      sprintf("Pearson r between %s and %s is %s. Correlation measures whether the readings rise and fall together; agreement asks how far apart they are. High correlation is not agreement.",
              m1_h, m2_h, r3(r_methods)),
      tolerance_note
    ),
    stringsAsFactors = FALSE
  )

  metrics <- list(
    `Pairs Analysed`            = n,
    `Bias(Mean Difference)`    = round(bias, 3),
    `Lower Limit of Agreement`  = round(loa_low, 3),
    `Upper Limit of Agreement`  = round(loa_high, 3),
    `Within Limits`             = paste0(r2(pct_within), "%"),
    `Proportional Bias`         = if (prop_bias) "detected" else "not detected",
    `Correlation r`             = if (is.na(r_methods)) NA_real_ else round(r_methods, 3)
  )

  contrast_clause <- if (contrast_fires) {
    paste0(" The two methods correlate strongly(r = ", r3(r_methods),
           "), yet they do not agree: the bias of ", r2(bias),
           " is a real systematic offset — high correlation is not agreement.")
  } else ""
  prop_clause <- if (prop_bias) {
    paste0(" Warning: the difference changes with the size of the measurement(slope ",
           r3(slope), ", ", fmt_pp(slope_p),
           "), so agreement varies with magnitude and the simple limits mislead.")
  } else ""

  json_output <- list(
    answer = paste0(
      "Bland-Altman agreement of ", m1_h, " versus ", m2_h, " across ",
      format(n, big.mark = ","), " paired measurements: the bias(", m1_h,
      " minus ", m2_h, ") is ", r2(bias), " (95% CI ", r2(bias_lo), " to ",
      r2(bias_hi), "), with 95% limits of agreement from ", r2(loa_low),
      " to ", r2(loa_high), "; ", r2(pct_within),
      "% of observed differences fall inside the limits. Proportional bias ",
      if (prop_bias) "was detected" else "was not detected", " (slope ",
      fmt_pp(slope_p), ").", contrast_clause, prop_clause,
      " Whether these limits are acceptable depends on your clinical/technical tolerance."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "bland_altman_plot",
        "agreement_table", "proportional_bias", "methods"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed, n = n, n_dropped = n_dropped,
    m1_h = m1_h, m2_h = m2_h,
    bias = bias, sd_d = sd_d, bias_lo = bias_lo, bias_hi = bias_hi,
    loa_low = loa_low, loa_high = loa_high,
    loa_low_lo = loa_low_lo, loa_low_hi = loa_low_hi,
    loa_high_lo = loa_high_lo, loa_high_hi = loa_high_hi,
    pct_within = pct_within, r_methods = r_methods,
    slope = slope, slope_se = slope_se, slope_p = slope_p,
    intercept = intercept,
    prop_bias = prop_bias, bias_sig = bias_sig, high_r = high_r,
    contrast_fires = contrast_fires, dir_word = dir_word,
    tolerance_note = tolerance_note,
    ba_points_df = ba_points_df, agreement_df = agreement_df,
    prop_bias_df = prop_bias_df, methods_df = methods_df,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell. Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No Signup Sign Up Free

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing