Executive Summary
Does Ad Spend relate to Revenue through Site Visits?
Ad Spend is associated with Revenue both through Site Visits and directly. The indirect pathway (through Site Visits) accounts for 54.8% of the total association. Ad Spend is associated with Site Visits at 0.612 per unit (path a, p < 0.001); Site Visits is associated with Revenue at 0.471 per unit among rows with the same Ad Spend (path b, p < 0.001). The indirect effect is 0.288 (95% bootstrap CI: 0.254 to 0.324, excluding zero). The direct effect is 0.238 and the total is 0.526. Causal interpretation requires three untestable assumptions: no unmeasured confounding, correct causal order, and no reverse causation. If these fail, the numbers remain accurate associations but are not causal effects.
Analysis Overview
Mediation analysis of Ad Spend and Revenue through Site Visits, across 2,000 rows.
This mediation analysis decomposes the association between Ad Spend and Revenue into two pathways: one running through Site Visits (indirect) and one not (direct). Three least-squares regressions supply the numbers: Ad Spend → Site Visits (path a), Site Visits → Revenue holding Ad Spend fixed (path b), and Ad Spend → Revenue alone (total path c). The indirect effect—the product of paths a and b—receives a bootstrap confidence interval rather than a normal-theory one, because products of estimated coefficients are skewed. The analysis answers which pathway dominates but cannot verify the causal assumptions that would license a causal reading: no unmeasured confounding of any pair, correct causal order (Ad Spend → Site Visits → Revenue), and no reverse causation. All 2,000 rows were complete and usable; no covariates were mapped.
Data Quality
Row completeness, column typing, and what was excluded.
All 2,000 rows loaded and retained; no rows were dropped for missing values. Every mapped column was usable as supplied. No covariates were mapped, so all reported associations are unadjusted—any variable not included in the model cannot be controlled for. Rows were retained rather than imputed because imputing a mediator would invent the quantity the analysis measures. The three columns (Ad Spend, Site Visits, Revenue) were recorded together on the same rows at one point in time, so causal ordering comes from domain knowledge, not from the data structure.
Effect Paths
Each arrow of the mediation model, with its estimate and uncertainty.
| Path | Estimate | Std Error | P Value | CI Low | CI High | Interpretation |
|---|---|---|---|---|---|---|
| a: Ad Spend → Site Visits | 0.6117 | 0.0219 | < 0.001 | — | — | A one-unit-higher Ad Spend is associated with a 0.612 higher Site Visits. |
| b: Site Visits → Revenue (holding Ad Spend fixed) | 0.4707 | 0.0253 | < 0.001 | — | — | Among rows with the same Ad Spend, a one-unit-higher Site Visits is associated with a 0.471 higher Revenue. |
| c-prime (direct): Ad Spend → Revenue (holding Site Visits fixed) | 0.2376 | 0.0292 | < 0.001 | — | — | The part of the Ad Spend-to-Revenue association that does not run through Site Visits. |
| c (total): Ad Spend → Revenue | 0.5255 | 0.0268 | < 0.001 | — | — | The whole Ad Spend-to-Revenue association, before splitting it up. |
| a×b (indirect, through Site Visits) | 0.2879 | — | < 0.001 (Sobel) | 0.254 | 0.3237 | The part that runs through Site Visits. Bootstrap 95% CI 0.254 to 0.324, which excludes zero. |
The path diagram shows four associations, each estimated by least-squares regression. Path a (Ad Spend → Site Visits) is 0.6117 (SE 0.0219, p < 0.001): a one-unit increase in Ad Spend is associated with a 0.612 increase in Site Visits. Path b (Site Visits → Revenue, holding Ad Spend fixed) is 0.4707 (SE 0.0253, p < 0.001): among rows with the same Ad Spend, one unit more Site Visits is associated with 0.471 more Revenue. The direct path c-prime (Ad Spend → Revenue, holding Site Visits fixed) is 0.2376 (SE 0.0292, p < 0.001). The total path c (Ad Spend → Revenue, unadjusted) is 0.5255 (SE 0.0268, p < 0.001). The indirect effect (a × b) is 0.2879, with a percentile bootstrap 95% CI of 0.254 to 0.324, which excludes zero. The Sobel test (p < 0.001) is reported for completeness but is not the decisive figure; the bootstrap interval is preferred because the product of two coefficients is skewed, not normally distributed.
Effect Decomposition
Total, direct, and indirect association side by side.
The total association of 0.5255 splits exactly into a direct component of 0.2376 and an indirect component of 0.2879, summing to 0.5255 by construction under least squares. The indirect pathway (through Site Visits) represents 54.8% of the total association. The indirect bar has a percentile bootstrap interval (0.254 to 0.324); the direct and total bars have normal-theory intervals (direct: 0.1804 to 0.2947; total: 0.473 to 0.578). These intervals come from different machinery and are not directly comparable in width. The indirect component is the largest share of the association, meaning Site Visits accounts for more than half the relationship between Ad Spend and Revenue. A large indirect bar reflects the pattern of associations in the data, not proof of a causal mechanism.
Bootstrap Distribution
Where the indirect effect of 'Ad Spend' on 'Revenue' through 'Site Visits' lands across 2,000 resamples of your rows.
The short answer
The indirect pathway—ad spend raising revenue through site visits—accounts for a substantial and stable portion of the total association. The bootstrap distribution places this indirect effect between 0.254 and 0.324 across resamples, with no estimates near zero, confirming the pathway is reliably present in your data.
The detail
The percentile bootstrap 95% confidence interval for the indirect effect runs from 0.254 to 0.324, derived from 2,000 resamples of your 2,000 rows. This interval sits entirely above zero and is skewed slightly right—a product of two coefficients typically is—which is why the percentile bootstrap (not the symmetric Sobel test) is the appropriate verdict. All 2,000 resamples yielded usable estimates.
What this can't tell you
This analysis assumes the causal order (ad spend → site visits → revenue) and rules out unmeasured confounding. If either assumption fails, the numbers remain accurate associations but not causal effects. The data were measured together, so causal direction comes from your process knowledge, not the data itself.
What This Rests On
The assumptions a causal reading requires, and which of them this dataset can check.
| Assumption | What It Means | Checkable Here |
|---|---|---|
| No unmeasured confounding | Nothing outside the model causes both Ad Spend and Site Visits, both Ad Spend and Revenue, or both Site Visits and Revenue. | No — requires knowledge outside the dataset |
| Causal ordering | Ad Spend comes before Site Visits, which comes before Revenue. | No — requires knowledge outside the dataset |
| No reverse causation | Revenue does not feed back into Site Visits. | No — requires knowledge outside the dataset |
| Design | Each row carries one observation of 3 columns measured together; the analysis has no record of anyone assigning Ad Spend. | No — the dataset records no assignment mechanism |
| Model form | Each path is a straight line and the errors are well behaved; Ad Spend and Revenue are treated as numeric. | Partly — the fitted models assume it |
| What this analysis can conclude | That Ad Spend, Site Visits and Revenue are associated in the pattern a mediation model would produce — not that Ad Spend causes Revenue by way of Site Visits. | Yes — this is what was computed |
Three untestable assumptions underpin a causal reading. (1) No unmeasured confounding: no variable outside the model causes both Ad Spend and Site Visits, both Ad Spend and Revenue, or both Site Visits and Revenue. (2) Correct causal order: Ad Spend precedes Site Visits, which precedes Revenue—not some other arrangement of these three columns. (3) No reverse causation: Revenue does not feed back into Site Visits. These columns were measured together on the same rows; the data provides no record of assignment, so causal order rests on domain knowledge. The dataset cannot check any of these. If any fails, the associations reported remain accurate descriptions of the data but are not causal effects. A plausible unmapped common cause of Ad Spend and Site Visits (for example, seasonal demand) would bias every estimate upward or downward, and the report cannot quantify by how much.
Methods & Disclosure
Every model, formula, and setting behind the numbers.
| Item | Detail |
|---|---|
| Path a | Ordinary least squares of Site Visits on Ad Spend: a = 0.612 (SE 0.022, p < 0.001). |
| Path b and c-prime | Ordinary least squares of Revenue on Ad Spend and Site Visits: b = 0.471 (SE 0.025, p < 0.001) and c-prime = 0.238 (SE 0.029, p < 0.001). |
| Path c (total) | Ordinary least squares of Revenue on Ad Spend: c = 0.526 (SE 0.027, p < 0.001). |
| Indirect effect | a times b = 0.288, with percentile bootstrap 95% CI 0.254 to 0.324 from 2,000 resamples (2,000 usable). |
| Bootstrap | Rows resampled with replacement, both regressions refitted on each resample, and the 2.5th and 97.5th percentiles of the resulting a times b values reported. Seed 20260728, so the interval is reproducible. |
| Sobel test | z = 15.496 (p < 0.001), computed as a times b divided by the square root of (b squared times the variance of a, plus a squared times the variance of b). Reported for completeness; its normality assumption is poor for a product of coefficients, so the bootstrap interval is the one to read. |
| Proportion mediated | Indirect divided by total = 0.548, so 54.8% of the Ad Spend-to-Revenue association runs through Site Visits. This ratio is only meaningful because the total association is itself distinguishable from zero and the direct and indirect parts point the same way. |
| Covariates | No covariates were mapped, so every association reported is unadjusted. |
| Decomposition check | With least squares on one sample and one covariate set, c must equal c-prime plus a times b exactly; the observed gap is 0.00000000, confirming the decomposition is arithmetically consistent. |
| What it does not establish | A mediation model fitted to columns measured together cannot establish causation. Reading this as "Ad Spend changes Site Visits, which in turn changes Revenue" requires three assumptions that no amount of arithmetic on this table can check: (1) no unmeasured variable causes both Ad Spend and Site Visits, or both Ad Spend and Revenue, or both Site Visits and Revenue; (2) the causal order really is Ad Spend then Site Visits then Revenue, and not some other ordering of the same three columns; (3) no reverse causation — Revenue does not feed back into Site Visits. These columns were measured together on the same rows, so the ordering comes from your knowledge of the process, not from the data. If any of the three fails, the numbers here are still an accurate description of the associations, but they are not causal effects. |
The short answer
Ad spend reaches revenue through both pathways: 54.8% of the total association flows through site visits (indirect effect = 0.288), while 45.2% flows directly (c-prime = 0.238). Both are statistically distinguishable from zero, so neither pathway dominates alone.
The detail
Path a (Ad Spend → Site Visits): 0.612 (SE 0.022, p < 0.001). Path b (Site Visits → Revenue): 0.471 (SE 0.025, p < 0.001). Their product (indirect effect) = 0.288, with percentile bootstrap 95% CI 0.254 to 0.324. Direct path c-prime (Ad Spend → Revenue, adjusted for Site Visits) = 0.238 (SE 0.029, p < 0.001). Total association c = 0.526 (SE 0.027, p < 0.001). Proportion mediated = 0.548 (54.8%). Bootstrap seed 20260728; Sobel z = 15.496 (p < 0.001), reported for reference but not as the verdict because the product is non-normal.
What this can't tell you
No covariates were included, so all estimates are unadjusted. Causal claims require ruling out unmeasured confounding and verifying causal order—both assumptions external to this data.
Methodology
Statistical methodology and diagnostics for Mediation & Moderation Analysis
Statistical Method
Standard-library analysis: does a predictor relate to an outcome through a third variable, or does its relationship with the outcome depend on one? Map a predictor, an outcome, and either a mediator or a moderator (plus optional covariates). Mediation splits the total association into the direct path and the indirect path that runs through the mediator — the a, b, c and c-prime paths with standard errors, the indirect effect a×b with a percentile bootstrap confidence interval, the proportion mediated where it is interpretable, the Sobel test reported alongside with its weakness stated, and the whole path diagram as a table. Moderation fits the interaction model and reports the conditional slope of the predictor at low, mean and high values of the moderator (or within each moderator level). Every narrative states, in your own column names, the assumption set a causal reading would rest on.
- Each row is one independent observation with the predictor, mediator (or moderator), and outcome measured on it
- The outcome and the mediator are numeric or cleanly convertible; the predictor is numeric or a two-value indicator
- Each path is approximately linear and the regression errors are well behaved
- For the decomposition to be read causally: no unmeasured variable causes both the predictor and the mediator, both the predictor and the outcome, or both the mediator and the outcome
- For the decomposition to be read causally: the causal order really is predictor, then mediator, then outcome, with no reverse causation
- On observational cross-sectional data this does NOT establish that the predictor causes the outcome through the mediator — the assumptions above are not testable from the data and the analysis states so in every narrative
- A single omitted common cause of any pair of the three variables biases the paths by an amount the analysis cannot estimate
- Only one mediator is analysed at a time; multiple or sequential mediators need a different model
- The mediator and outcome must be numeric — a binary mediator or outcome needs a different link function
Analysis Code
Complete R source code for this analysis
Mediation & Moderation Analysis
Does a predictor relate to an outcome through a third variable (mediation), or does its relationship with the outcome depend on a third variable (moderation)? Map a predictor, an outcome, and either a mediator or a moderator (plus optional covariates).
Mediation decomposes the total association into a direct path and an indirect (mediated) path: the a path (predictor to mediator), the b path (mediator to outcome, holding the predictor fixed), the direct path c-prime, and the total path c. The indirect effect ab gets a percentile bootstrap confidence interval computed by resampling rows and refitting both models — the product ab is markedly non-normal, so a bootstrap interval is preferred over the Sobel normal-theory test, which is reported alongside for completeness.
Moderation fits the interaction model and reports the conditional (simple) slope of the predictor at low, mean, and high values of the moderator (or within each moderator level when the moderator is categorical).
Statistical honesty
On observational cross-sectional data this analysis does NOT establish that the predictor causes the outcome through the mediator. Every narrative it emits states the assumption set the causal reading rests on — no unmeasured confounding of any of the three paths, correct causal ordering, and no reverse causation — in plain language, using the user's own column names.
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {predictor, outcome, mediator | moderator, covariate_1..N}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# mode $ "mediation" | "moderation"
# initial_rows/final_rows/rows_removed/n $ row accounting
# x_h / y_h / m_h / w_h $ humanized user names of the mapped columns
# cov_labels $ humanized covariate labels actually in the model
# dropped_notes $ character — every column dropped and why
# a/b/c/c_prime + *_se/*_p $ mediation path estimates
# indirect / ind_lo / ind_hi $ indirect effect + percentile bootstrap CI
# prop_med / prop_med_ok $ proportion mediated + whether interpretable
# sobel_z / sobel_p $ Sobel normal-theory test (reported, not preferred)
# inter_coef/inter_se/inter_p $ moderation interaction estimate
# slopes_df $ conditional (simple) slopes with CIs
# boot_B / boot_valid / boot_seed $ bootstrap accounting
# path_df / effect_df / draws_df / assume_df / methods_df $ card datasets
# assumption_text / headline $ computed prose blocks reused across cards
# metrics / json_output
# === /SHARED EXPORTS ===Step 1: Resolve the mapped columns and decide the analysis mode
initial_rows <- nrow(df)
x_h <- humanize_semantic("predictor", col_map)
y_h <- humanize_semantic("outcome", col_map)
m_h <- humanize_semantic("mediator", col_map)
w_h <- humanize_semantic("moderator", col_map)
if (!("predictor" %in% names(df))) {
stop(sprintf("No predictor column was mapped. Map the column you think does the driving(requested as '%s').", x_h))
}
if (!("outcome" %in% names(df))) {
stop(sprintf("No outcome column was mapped. Map the numeric result you want explained(requested as '%s').", y_h))
}
has_m <- "mediator" %in% names(df)
has_w <- "moderator" %in% names(df)
if (!has_m && !has_w) {
stop(sprintf("Neither a mediator nor a moderator was mapped, so there is no third variable to analyse. Map a mediator(the step '%s' is supposed to work through on its way to '%s') for a mediation analysis, or a moderator (the condition under which the link between '%s' and '%s' changes) for a moderation analysis.",
x_h, y_h, x_h, y_h))
}
mode <- if (has_m) "mediation" else "moderation"
both_mapped <- has_m && has_wStep 2: Coerce the modelled columns (95% numeric-coercion rule)
dropped_notes <- character(0)
as_num_95 <- function(v) {
if (is.numeric(v)) return(v)
ch <- as.character(v)
non_blank <- !is.na(ch) & trimws(ch) != ""
conv <- suppressWarnings(as.numeric(ch))
if (sum(non_blank) == 0) return(NULL)
if (sum(!is.na(conv[non_blank])) < 0.95 * sum(non_blank)) return(NULL)
conv
}The predictor may legitimately be a two-level treatment flag; anything else non-numeric is refused by name rather than silently coerced.
xv <- as_num_95(df$predictor)
predictor_encoding <- ""
if (is.null(xv)) {
ch <- as.character(df$predictor)
lv <- sort(unique(trimws(ch[!is.na(ch) & trimws(ch) != ""])))
if (length(lv) == 2) {
xv <- ifelse(trimws(ch) == lv[2], 1, ifelse(trimws(ch) == lv[1], 0, NA_real_))
predictor_encoding <- sprintf("'%s' has exactly two values and was encoded as 0 for '%s' and 1 for '%s', so its coefficient is the difference between those two groups. ",
x_h, lv[1], lv[2])
} else {
stop(sprintf("The predictor column '%s' is neither numeric nor a two-value indicator (it has %d distinct values), so it cannot be used as a predictor here. Map a numeric column, or a column with exactly two values.",
x_h, length(lv)))
}
}
yv <- as_num_95(df$outcome)
if (is.null(yv)) {
stop(sprintf("The outcome column '%s' does not look numeric — fewer than 95%% of its values parse as numbers. Mediation and moderation here need a numeric outcome.", y_h))
}
mv <- NULL
if (mode == "mediation") {
mv <- as_num_95(df$mediator)
if (is.null(mv)) {
stop(sprintf("The mediator column '%s' does not look numeric — fewer than 95%% of its values parse as numbers. The mediator must be a numeric quantity that '%s' can move and that can in turn move '%s'.",
m_h, x_h, y_h))
}
}Step 3: Prepare the moderator (numeric, or a lumped categorical)
wv <- NULL; w_levels <- character(0); w_is_numeric <- TRUE
if (mode == "moderation") {
wv <- as_num_95(df$moderator)
if (is.null(wv)) {
w_is_numeric <- FALSE
ch <- trimws(as.character(df$moderator))
ch[is.na(ch) | ch == ""] <- "Missing"
tab <- sort(table(ch), decreasing = TRUE)
if (length(tab) > 8) {
keep_lv <- names(tab)[seq_len(8)]
ch[!(ch %in% keep_lv)] <- "Other"
dropped_notes <- c(dropped_notes, sprintf(
"'%s' had %d distinct values; the 8 most common were kept and the rest pooled into 'Other'.",
w_h, length(tab)))
}
if (length(unique(ch)) < 2) {
stop(sprintf("The moderator column '%s' has only one distinct value, so there is nothing for the link between '%s' and '%s' to vary across.",
w_h, x_h, y_h))
}
wv <- ch
}
}Step 4: Covariates — coerce, dummy-code, drop the unusable by name
cov_keys <- grep("^covariate_[0-9]+$", names(df), value = TRUE)
cov_keys <- cov_keys[order(as.integer(sub("^covariate_", "", cov_keys)))]
cov_mat_list <- list(); cov_labels <- character(0)
n_rows_all <- nrow(df)
for (ck in cov_keys) {
ch_lab <- humanize_semantic(ck, col_map)
v <- df[[ck]]
num <- as_num_95(v)
if (!is.null(num)) {
if (all(is.na(num))) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: every value is missing.", ch_lab))
next
}
if (isTRUE(stats::var(num, na.rm = TRUE) == 0)) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it is constant, so it cannot explain any variation.", ch_lab))
next
}
cm <- matrix(num, ncol = 1); colnames(cm) <- ch_lab
cov_mat_list[[length(cov_mat_list) + 1]] <- cm
cov_labels <- c(cov_labels, ch_lab)
next
}
ch <- trimws(as.character(v))
ch[is.na(ch) | ch == ""] <- "Missing"
u <- unique(ch)
if (length(u) < 2) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it is constant, so it cannot explain any variation.", ch_lab))
next
}
if (length(u) > 0.9 * n_rows_all) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it has a near-unique value on almost every row, which makes it an identifier rather than a covariate.", ch_lab))
next
}
tab <- sort(table(ch), decreasing = TRUE)
if (length(tab) > 8) {
keep_lv <- names(tab)[seq_len(8)]
ch[!(ch %in% keep_lv)] <- "Other"
dropped_notes <- c(dropped_notes, sprintf("'%s' had %d categories; the 8 most common were kept and the rest pooled into 'Other'.", ch_lab, length(tab)))
}
lv <- sort(unique(ch))
ref <- lv[1]
for (l in lv[-1]) {
cm <- matrix(as.numeric(ch == l), ncol = 1)
colnames(cm) <- sprintf("%s: %s vs %s", ch_lab, l, ref)
cov_mat_list[[length(cov_mat_list) + 1]] <- cm
cov_labels <- c(cov_labels, colnames(cm))
}
}
C <- if (length(cov_mat_list) > 0) do.call(cbind, cov_mat_list) else NULLStep 5: Row-wise complete cases across everything the models use
ok <- !is.na(xv) & !is.na(yv)
if (mode == "mediation") ok <- ok & !is.na(mv)
if (mode == "moderation") ok <- ok & (if (w_is_numeric) !is.na(wv) else !is.na(wv) & wv != "")
if (!is.null(C)) ok <- ok & stats::complete.cases(C)
n_dropped <- sum(!ok)
xv <- xv[ok]; yv <- yv[ok]
if (mode == "mediation") mv <- mv[ok]
if (mode == "moderation") wv <- if (w_is_numeric) wv[ok] else wv[ok]
if (!is.null(C)) C <- C[ok, , drop = FALSE]
n <- length(xv)
final_rows <- n
rows_removed <- initial_rows - final_rows
MIN_ROWS <- 30
if (n < MIN_ROWS) {
stop(sprintf("Only %d rows have a value for every mapped column('%s', '%s'%s) — at least %d complete rows are needed before a mediation or moderation model can be fitted, and a percentile bootstrap on fewer rows would be meaningless.",
n, x_h, y_h,
if (mode == "mediation") sprintf(", '%s'", m_h) else sprintf(", '%s'", w_h),
MIN_ROWS))
}Step 6: Degenerate-input guards, each naming the user's own column
if (isTRUE(stats::var(xv) == 0)) {
stop(sprintf("The predictor column '%s' is constant — every row has the same value — so it cannot be related to anything.", x_h))
}
if (isTRUE(stats::var(yv) == 0)) {
stop(sprintf("The outcome column '%s' is constant — every row has the same value — so there is no variation to explain.", y_h))
}
if (mode == "mediation" && isTRUE(stats::var(mv) == 0)) {
stop(sprintf("The mediator column '%s' is constant — every row has the same value — so nothing can travel through it from '%s' to '%s'.",
m_h, x_h, y_h))
}
if (mode == "moderation" && w_is_numeric && isTRUE(stats::var(wv) == 0)) {
stop(sprintf("The moderator column '%s' is constant — every row has the same value — so the link between '%s' and '%s' has nothing to vary across.",
w_h, x_h, y_h))
}Drop covariate columns that are collinear with the rest of the design.
if (!is.null(C)) {
keep_c <- rep(TRUE, ncol(C))
for (j in seq_len(ncol(C))) {
if (isTRUE(stats::var(C[, j]) == 0)) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it is constant on the rows that survived cleaning.", colnames(C)[j]))
keep_c[j] <- FALSE
}
}
C <- C[, keep_c, drop = FALSE]
if (ncol(C) > 0) {
base_cols <- cbind(1, xv, if (mode == "mediation") mv else NULL)
probe <- cbind(base_cols, C)
qp <- qr(probe)
if (qp$rank < ncol(probe)) {
keep_idx <- sort(qp$pivot[seq_len(qp$rank)])
drop_idx <- setdiff(seq_len(ncol(probe)), keep_idx)
drop_cov <- drop_idx[drop_idx > ncol(base_cols)] - ncol(base_cols)
for (j in drop_cov) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it is an exact linear combination of the other columns in the model.", colnames(C)[j]))
}
C <- C[, setdiff(seq_len(ncol(C)), drop_cov), drop = FALSE]
}
}
if (ncol(C) == 0) C <- NULL
}
cov_labels <- if (is.null(C)) character(0) else colnames(C)Step 7: Bootstrap settings — deterministic and disclosed
boot_seed <- 20260728L
boot_B <- suppressWarnings(as.integer(params$bootstrap_samples %||% 2000L))
if (is.na(boot_B)) boot_B <- 2000L
boot_B <- max(500L, min(5000L, boot_B))
if (n > 5000) boot_B <- min(boot_B, 1000L)
a <- b <- c_tot <- c_prime <- NA_real_
a_se <- b_se <- c_se <- cp_se <- NA_real_
a_p <- b_p <- c_p <- cp_p <- NA_real_
indirect <- ind_lo <- ind_hi <- NA_real_
prop_med <- NA_real_; prop_med_ok <- FALSE
sobel_z <- sobel_p <- NA_real_
inter_coef <- inter_se <- inter_p <- NA_real_
slopes_df <- NULL
boot_valid <- 0L
boot_draws <- numeric(0)
boot_label <- ""
identity_gap <- NA_real_
if (mode == "mediation") {Step 8 (mediation): the three regressions behind a, b, c and c-prime
Am <- cbind(1, xv); if (!is.null(C)) Am <- cbind(Am, C) # M ~ X + covariates
Ay <- cbind(1, xv, mv); if (!is.null(C)) Ay <- cbind(Ay, C) # Y ~ X + M + covariates
At <- cbind(1, xv); if (!is.null(C)) At <- cbind(At, C) # Y ~ X + covariates
fit_m <- ols_fit(Am, mv)
fit_y <- ols_fit(Ay, yv)
fit_t <- ols_fit(At, yv)
if (is.null(fit_m) || is.null(fit_y) || is.null(fit_t)) {
stop(sprintf("The mediation models for '%s', '%s' and '%s' could not be fitted — the mapped columns (or the covariates) are collinear, so the paths are not separately identified.",
x_h, m_h, y_h))
}
a <- fit_m$coef[2]; a_se <- fit_m$se[2]; a_p <- fit_m$p[2]
c_prime <- fit_y$coef[2]; cp_se <- fit_y$se[2]; cp_p <- fit_y$p[2]
b <- fit_y$coef[3]; b_se <- fit_y$se[3]; b_p <- fit_y$p[3]
c_tot <- fit_t$coef[2]; c_se <- fit_t$se[2]; c_p <- fit_t$p[2]
indirect <- a * bWith OLS on one sample and one covariate set, c = c-prime + a*b exactly.
identity_gap <- abs(c_tot - (c_prime + indirect))Step 9 (mediation): percentile bootstrap for the indirect effect
set.seed(boot_seed)
boot_draws <- replicate(boot_B, {
idx <- sample.int(n, n, replace = TRUE)
fm <- .lm.fit(Am[idx, , drop = FALSE], mv[idx])
fy <- .lm.fit(Ay[idx, , drop = FALSE], yv[idx])
if (fm$rank < ncol(Am) || fy$rank < ncol(Ay)) return(NA_real_)
ca <- fm$coefficients[2]; cb <- fy$coefficients[3]
if (is.na(ca) || is.na(cb)) return(NA_real_)
ca * cb
})
boot_label <- sprintf("indirect effect of '%s' on '%s' through '%s'", x_h, y_h, m_h)
good <- boot_draws[is.finite(boot_draws)]
boot_valid <- length(good)
if (boot_valid >= max(200L, as.integer(0.8 * boot_B))) {
qs <- stats::quantile(good, c(0.025, 0.975), names = FALSE, type = 7)
ind_lo <- qs[1]; ind_hi <- qs[2]
}Step 10 (mediation): Sobel test — reported, but not the preferred CI
if (!is.na(a_se) && !is.na(b_se)) {
denom <- sqrt(b^2 * a_se^2 + a^2 * b_se^2)
if (is.finite(denom) && denom > 0) {
sobel_z <- indirect / denom
sobel_p <- 2 * stats::pnorm(-abs(sobel_z))
}
}Proportion mediated is only interpretable when the total path is non-trivial and the direct and indirect paths point the same way.
if (!is.na(c_tot) && abs(c_tot) > 1e-9) {
prop_med <- indirect / c_tot
prop_med_ok <- is.finite(prop_med) && prop_med > 0 && prop_med <= 1 &&
!is.na(c_p) && c_p < 0.05
}
} else {Step 8 (moderation): the interaction model
if (w_is_numeric) {
Aw <- cbind(1, xv, wv, xv * wv)
lab_w <- c("(Intercept)", x_h, w_h, sprintf("%s × %s", x_h, w_h))
} else {
lv <- sort(unique(wv)); ref <- lv[1]
D <- sapply(lv[-1], function(l) as.numeric(wv == l))
D <- matrix(as.numeric(D), nrow = n)
XD <- D * xv
Aw <- cbind(1, xv, D, XD)
lab_w <- c("(Intercept)", x_h,
sprintf("%s: %s vs %s", w_h, lv[-1], ref),
sprintf("%s × %s: %s vs %s", x_h, w_h, lv[-1], ref))
}
if (!is.null(C)) { Aw <- cbind(Aw, C); lab_w <- c(lab_w, cov_labels) }
fit_w <- ols_fit(Aw, yv)
if (is.null(fit_w)) {
stop(sprintf("The moderation model for '%s', '%s' and '%s' could not be fitted — the mapped columns (or the covariates) are collinear, so the interaction is not identified.",
x_h, w_h, y_h))
}Step 9 (moderation): conditional (simple) slopes with their own CIs
tq <- stats::qt(0.975, df = fit_w$df)
simple_slope <- function(k) {
# k is a contrast vector over the coefficient vector
est <- sum(k * fit_w$coef)
va <- as.numeric(t(k) %*% fit_w$vcov %*% k)
se <- sqrt(max(va, 0))
tv <- if (se > 0) est / se else NA_real_
pv <- if (is.na(tv)) NA_real_ else 2 * stats::pt(-abs(tv), df = fit_w$df)
c(est = est, se = se, lo = est - tq * se, hi = est + tq * se, p = pv)
}
p_all <- length(fit_w$coef)
if (w_is_numeric) {
mw <- mean(wv); sw <- stats::sd(wv)
w_points <- c(mw - sw, mw, mw + sw)
w_names <- c(sprintf("Low %s(%s)", w_h, r2(w_points[1])),
sprintf("Mean %s(%s)", w_h, r2(w_points[2])),
sprintf("High %s(%s)", w_h, r2(w_points[3])))
rows <- lapply(seq_along(w_points), function(i) {
k <- numeric(p_all); k[2] <- 1; k[4] <- w_points[i]
simple_slope(k)
})
inter_coef <- fit_w$coef[4]; inter_se <- fit_w$se[4]; inter_p <- fit_w$p[4]
} else {
lv <- sort(unique(wv))
w_names <- sprintf("%s = %s", w_h, lv)
rows <- lapply(seq_along(lv), function(i) {
k <- numeric(p_all); k[2] <- 1
if (i > 1) k[2 + length(lv) - 1 + (i - 1)] <- 1
simple_slope(k)
})With a categorical moderator the "interaction" is a set of terms; the headline number is the spread between the strongest and weakest slope.
ests <- sapply(rows, function(r) r[["est"]])
hi_i <- which(ests == max(ests))[1]; lo_i <- which(ests == min(ests))[1]
k <- numeric(p_all)
if (hi_i > 1) k[2 + length(lv) - 1 + (hi_i - 1)] <- 1
if (lo_i > 1) k[2 + length(lv) - 1 + (lo_i - 1)] <- k[2 + length(lv) - 1 + (lo_i - 1)] - 1
sp <- simple_slope(k)
inter_coef <- sp[["est"]]; inter_se <- sp[["se"]]; inter_p <- sp[["p"]]
}
slopes_df <- data.frame(
condition = w_names,
slope = round(sapply(rows, function(r) r[["est"]]), 4),
std_error = round(sapply(rows, function(r) r[["se"]]), 4),
ci_low = round(sapply(rows, function(r) r[["lo"]]), 4),
ci_high = round(sapply(rows, function(r) r[["hi"]]), 4),
p_value = sapply(rows, function(r) fmt_p(r[["p"]])),
stringsAsFactors = FALSE
)Step 10 (moderation): bootstrap the moderation quantity itself
inter_idx <- if (w_is_numeric) 4L else NA_integer_
kv <- if (w_is_numeric) NULL else {
lv <- sort(unique(wv)); ests <- slopes_df$slope
hi_i <- which(ests == max(ests))[1]; lo_i <- which(ests == min(ests))[1]
k <- numeric(p_all)
if (hi_i > 1) k[2 + length(lv) - 1 + (hi_i - 1)] <- 1
if (lo_i > 1) k[2 + length(lv) - 1 + (lo_i - 1)] <- k[2 + length(lv) - 1 + (lo_i - 1)] - 1
k
}
set.seed(boot_seed)
boot_draws <- replicate(boot_B, {
idx <- sample.int(n, n, replace = TRUE)
fw <- .lm.fit(Aw[idx, , drop = FALSE], yv[idx])
if (fw$rank < ncol(Aw)) return(NA_real_)
cf <- fw$coefficients
if (any(is.na(cf))) return(NA_real_)
if (!is.na(inter_idx)) cf[inter_idx] else sum(kv * cf)
})
boot_label <- if (w_is_numeric)
sprintf("interaction between '%s' and '%s'", x_h, w_h)
else
sprintf("gap between the strongest and weakest '%s' slope across levels of '%s'", x_h, w_h)
good <- boot_draws[is.finite(boot_draws)]
boot_valid <- length(good)
if (boot_valid >= max(200L, as.integer(0.8 * boot_B))) {
qs <- stats::quantile(good, c(0.025, 0.975), names = FALSE, type = 7)
ind_lo <- qs[1]; ind_hi <- qs[2]
}
indirect <- inter_coef
}
boot_sig <- !is.na(ind_lo) && !is.na(ind_hi) && (ind_lo > 0 || ind_hi < 0)Step 11: The assumption block — computed, and printed every time
assumption_text <- if (mode == "mediation") {
paste0(
"Reading this as \"", x_h, " changes ", m_h, ", which in turn changes ", y_h,
"\" requires three assumptions that no amount of arithmetic on this table can check: ",
"(1) no unmeasured variable causes both ", x_h, " and ", m_h, ", or both ", x_h,
" and ", y_h, ", or both ", m_h, " and ", y_h, "; ",
"(2) the causal order really is ", x_h, " then ", m_h, " then ", y_h,
", and not some other ordering of the same three columns; ",
"(3) no reverse causation — ", y_h, " does not feed back into ", m_h, ". ",
"These columns were measured together on the same rows, so the ordering comes from ",
"your knowledge of the process, not from the data. If any of the three fails, the ",
"numbers here are still an accurate description of the associations, but they are ",
"not causal effects."
)
} else {
paste0(
"Reading this as \"", w_h, " changes how ", x_h, " affects ", y_h,
"\" requires assumptions this data cannot check: ",
"(1) no unmeasured variable causes both ", x_h, " and ", y_h,
" in a way that happens to vary with ", w_h, "; ",
"(2) ", w_h, " is a condition that precedes the ", x_h, "-to-", y_h,
" relationship rather than a consequence of it; ",
"(3) no reverse causation — ", y_h, " does not feed back into ", x_h, " or ", w_h, ". ",
"These columns were measured together on the same rows, so the ordering comes from ",
"your knowledge of the process, not from the data. An interaction is a statement ",
"about how an association varies, not proof that either variable causes the other."
)
}
boot_pref_text <- paste0(
"The confidence interval above is a percentile bootstrap: the rows were resampled with ",
"replacement ", format(boot_B, big.mark = ","), " times, the model refitted on each ",
"resample, and the middle 95% of the resulting estimates reported. This is used in ",
"preference to the Sobel test because ",
if (mode == "mediation")
"the product of two coefficients is skewed and not normally distributed, so a normal-theory interval on it is too narrow and mis-centred, especially in smaller samples. The Sobel test is reported alongside for completeness, not as the verdict."
else
"the sampling distribution of a conditional effect can be skewed, and the bootstrap does not require it to be normal."
)Step 13: Metrics and the computed one-paragraph answer
metrics <- if (mode == "mediation") {
list(
`Rows Analysed` = n,
`Path a` = round(a, 4),
`Path b` = round(b, 4),
`Direct Effect` = round(c_prime, 4),
`Total Effect` = round(c_tot, 4),
`Indirect Effect` = round(indirect, 4),
`Indirect 95% CI` = paste0(r3(ind_lo), " to ", r3(ind_hi)),
`Indirect CI Excludes Zero` = if (boot_sig) "yes" else "no",
`Proportion Mediated` = if (prop_med_ok) pct1(100 * prop_med) else "not interpretable",
`Bootstrap Resamples` = as.integer(boot_B)
)
} else {
list(
`Rows Analysed` = n,
`Moderation Estimate` = round(inter_coef, 4),
`Moderation 95% CI` = paste0(r3(ind_lo), " to ", r3(ind_hi)),
`Moderation CI Excludes Zero` = if (boot_sig) "yes" else "no",
`Lowest Conditional Slope` = round(min(slopes_df$slope), 4),
`Highest Conditional Slope` = round(max(slopes_df$slope), 4),
`Bootstrap Resamples` = as.integer(boot_B)
)
}
headline <- if (mode == "mediation") {
paste0(
"Across ", format(n, big.mark = ","), " rows, ", x_h, " is associated with ", m_h,
" (a = ", r3(a), ", ", fmt_pp(a_p), "), and ", m_h, " is associated with ", y_h,
" among rows with the same ", x_h, " (b = ", r3(b), ", ", fmt_pp(b_p), "). ",
"The indirect path a×b is ", r3(indirect), " with a percentile bootstrap 95% CI of ",
r3(ind_lo), " to ", r3(ind_hi), ", which ",
if (boot_sig) "excludes zero" else "includes zero",
"; the direct path c-prime is ", r3(c_prime), " and the total is ", r3(c_tot), ". ",
if (prop_med_ok)
paste0(pct1(100 * prop_med), " of the total association runs through ", m_h, ". ")
else
paste0("A proportion-mediated figure is not reported: the total association(",
r3(c_tot), ", ", fmt_pp(c_p),
") is not solid enough for that ratio to mean anything. ")
)
} else {
paste0(
"Across ", format(n, big.mark = ","), " rows, the association between ", x_h,
" and ", y_h, " varies with ", w_h, ": the conditional slope runs from ",
r3(min(slopes_df$slope)), " to ", r3(max(slopes_df$slope)),
" across the reported conditions. The moderation estimate is ", r3(inter_coef),
" (bootstrap 95% CI ", r3(ind_lo), " to ", r3(ind_hi), "), which ",
if (boot_sig) "excludes zero" else "includes zero", ". "
)
}
json_output <- list(
answer = paste0(
if (mode == "mediation") "Mediation analysis. " else "Moderation analysis. ",
headline,
"These are regression decompositions of associations, not established causal effects: ",
assumption_text
),
cards = lapply(
c("tldr", "overview", "preprocessing", "effect_paths", "effect_decomposition",
"bootstrap_distribution", "assumptions", "methods"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
mode = mode, both_mapped = both_mapped,
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed, n = n, n_dropped = n_dropped,
x_h = x_h, y_h = y_h, m_h = m_h, w_h = w_h,
w_is_numeric = w_is_numeric,
cov_labels = cov_labels, dropped_notes = dropped_notes,
predictor_encoding = predictor_encoding,
a = a, b = b, c_tot = c_tot, c_prime = c_prime,
a_se = a_se, b_se = b_se, c_se = c_se, cp_se = cp_se,
a_p = a_p, b_p = b_p, c_p = c_p, cp_p = cp_p,
indirect = indirect, ind_lo = ind_lo, ind_hi = ind_hi, boot_sig = boot_sig,
prop_med = prop_med, prop_med_ok = prop_med_ok,
sobel_z = sobel_z, sobel_p = sobel_p, identity_gap = identity_gap,
inter_coef = inter_coef, inter_se = inter_se, inter_p = inter_p,
slopes_df = slopes_df,
boot_B = boot_B, boot_valid = boot_valid, boot_seed = boot_seed,
boot_label = boot_label, boot_pref_text = boot_pref_text,
assumption_text = assumption_text, headline = headline,
path_df = path_df, effect_df = effect_df, draws_df = draws_df,
assume_df = assume_df, methods_df = methods_df,
metrics = metrics, json_output = json_output
)
}