Executive Summary
Multiple-testing correction across 40 tests
Of your 40 tests, 11 were significant uncorrected (p < 0.05). After applying Holm-Bonferroni correction—which controls the strictest family-wise error rate—10 survive. Benjamini-Hochberg, which tolerates a controlled fraction of false discoveries, also retains 10. One uncorrected finding (likely Metric 22, raw p = 0.0466) fails to survive either correction, flagged as a probable multiple-testing artifact. Your core discoveries are robust.
Analysis Overview
Holm-Bonferroni and Benjamini-Hochberg corrections applied to 40 p-values from Test Name.
Running 40 independent tests at α = 0.05 inflates false positives: by chance alone, you'd expect ~2 spurious hits. Holm-Bonferroni controls the family-wise error rate—the probability of even ONE false positive across all 40 tests—making it the strictest standard. Benjamini-Hochberg controls the false discovery rate, the expected fraction of false positives among your discoveries, permitting slightly more findings in exchange for accepting a small proportion of errors. Here, 11 tests reached raw significance, but Holm retained only 10 and BH also retained 10. The two methods converge, suggesting the surviving 10 findings are genuinely robust against multiple-testing inflation.
Data Quality
P-value validation and exclusions.
All 40 rows loaded from your data were valid: 0 p-values were missing and 0 fell outside the valid range [0, 1]. All 40 tests entered the correction analysis. Data quality was complete, so no adjustment to the correction factors was needed. The family size remained 40 throughout, ensuring both Holm and Benjamini-Hochberg thresholds reflect the full scope of your testing burden.
Adjusted Results
Every test with raw, Holm, and BH p-values and significance verdicts.
| Test | Raw P | Holm P | Bh P | Sig Raw | Sig Holm | Sig Bh |
|---|---|---|---|---|---|---|
| Metric 9 | 0.0001 | 0.003 | 0.0019 | Yes | Yes | Yes |
| Metric 7 | 0.0001 | 0.0045 | 0.0019 | Yes | Yes | Yes |
| Metric 4 | 0.0001 | 0.0055 | 0.0019 | Yes | Yes | Yes |
| Metric 2 | 0.0003 | 0.0112 | 0.003 | Yes | Yes | Yes |
| Metric 1 | 0.0006 | 0.0233 | 0.0048 | Yes | Yes | Yes |
| Metric 6 | 0.0007 | 0.0256 | 0.0048 | Yes | Yes | Yes |
| Metric 10 | 0.0009 | 0.0295 | 0.0048 | Yes | Yes | Yes |
| Metric 8 | 0.001 | 0.0335 | 0.0048 | Yes | Yes | Yes |
| Metric 5 | 0.0011 | 0.0343 | 0.0048 | Yes | Yes | Yes |
| Metric 3 | 0.0013 | 0.0404 | 0.0052 | Yes | Yes | Yes |
| Metric 22 | 0.0466 | 1 | 0.1694 | Yes | No | No |
| Metric 35 | 0.0596 | 1 | 0.1932 | No | No | No |
| Metric 34 | 0.0628 | 1 | 0.1932 | No | No | No |
| Metric 11 | 0.0699 | 1 | 0.1996 | No | No | No |
| Metric 12 | 0.0907 | 1 | 0.2419 | No | No | No |
| Metric 26 | 0.1178 | 1 | 0.2913 | No | No | No |
| Metric 15 | 0.1238 | 1 | 0.2913 | No | No | No |
| Metric 25 | 0.1443 | 1 | 0.3206 | No | No | No |
| Metric 29 | 0.1807 | 1 | 0.3805 | No | No | No |
| Metric 36 | 0.206 | 1 | 0.4119 | No | No | No |
| Metric 16 | 0.2232 | 1 | 0.4252 | No | No | No |
| Metric 24 | 0.2896 | 1 | 0.5236 | No | No | No |
| Metric 27 | 0.3085 | 1 | 0.5236 | No | No | No |
| Metric 39 | 0.3141 | 1 | 0.5236 | No | No | No |
| Metric 32 | 0.3724 | 1 | 0.5958 | No | No | No |
| Metric 20 | 0.3967 | 1 | 0.6103 | No | No | No |
| Metric 13 | 0.4245 | 1 | 0.6108 | No | No | No |
| Metric 38 | 0.4276 | 1 | 0.6108 | No | No | No |
| Metric 33 | 0.5477 | 1 | 0.732 | No | No | No |
| Metric 19 | 0.5771 | 1 | 0.732 | No | No | No |
| Metric 30 | 0.5816 | 1 | 0.732 | No | No | No |
| Metric 40 | 0.5856 | 1 | 0.732 | No | No | No |
| Metric 17 | 0.6274 | 1 | 0.7517 | No | No | No |
| Metric 31 | 0.6389 | 1 | 0.7517 | No | No | No |
| Metric 37 | 0.6804 | 1 | 0.7776 | No | No | No |
| Metric 28 | 0.8161 | 1 | 0.8939 | No | No | No |
| Metric 14 | 0.8269 | 1 | 0.8939 | No | No | No |
| Metric 23 | 0.8585 | 1 | 0.9037 | No | No | No |
| Metric 18 | 0.9477 | 1 | 0.972 | No | No | No |
| Metric 21 | 0.9763 | 1 | 0.9763 | No | No | No |
The sorted results show your strongest finding is Metric 9 (raw p = 0.0001, Holm p = 0.003, BH p = 0.0019)—significant under all three standards. The top 10 tests all pass both corrections: Metric 9, 7, 4, 2, 1, 6, 10, 8, 5, and 3, with raw p-values ranging from 0.0001 to 0.0013. The 11th significant uncorrected test, Metric 22 (raw p = 0.0466), fails both corrections (Holm p = 1, BH p = 0.1694). This boundary case exemplifies how correction penalizes marginal raw p-values in a large family; it does not meet the stricter thresholds and is likely a false positive.
What Survives Correction
Significant test counts: uncorrected vs Holm vs BH.
Of your 40 tests, the significance counts under each standard are: 11 uncorrected, 10 Holm-Bonferroni, and 10 Benjamini-Hochberg. The loss of 1 finding (from 11 to 10) under both corrections is modest, reflecting that your strongest signals are well below the noise floor. Holm and BH agree exactly here, both removing Metric 22 as a plausible artifact. This tight agreement suggests the surviving 10 findings are solid; the single casualty was borderline (raw p = 0.0466) and unlikely to be a true effect.
P-Value Distribution
10-bin histogram of raw p-values — the real-effects-vs-noise diagnostic.
The raw p-value histogram reveals a strong signal: the lowest bin (0.0–0.1) contains 15 of 40 p-values, roughly 3.8 times the 4 expected under a uniform null distribution. This spike is the diagnostic hallmark of genuine non-null effects mixed with noise. The remaining bins (0.1–1.0) hold 25 p-values distributed relatively flatly, consistent with null tests. The enrichment in the left tail supports the reality of your top 10 surviving findings and explains why the corrections remove only 1 uncorrected result—most of the significant p-values are far enough below the noise threshold to survive stringent adjustment.
Methodology
Statistical methodology and diagnostics for Multiple Comparisons Correction
Statistical Method
Standard-library analysis: you ran many statistical tests — which results are still significant after correction? Upload your p-values and get Holm-Bonferroni and Benjamini-Hochberg adjusted p-values side by side, a count of what survives each method, and a p-value distribution diagnostic that hints whether your significant results reflect real effects or multiple-testing luck.
- P-values come from valid individual tests (the correction adjusts them; it cannot fix a broken test)
- Tests are treated as one family; BH's FDR guarantee assumes independent or positively dependent tests
- Corrections reduce power — a true effect with a middling p-value may not survive, especially under Holm
- The choice of family (which tests to correct together) is a judgment call the tool cannot make for you
- P-values outside 0-1 or missing are dropped, which shrinks the family and changes the adjustment
Analysis Code
Complete R source code for this analysis
Multiple Comparisons Correction — What Survives?
Takes a set of raw p-values from many tests, applies Holm-Bonferroni and Benjamini-Hochberg corrections, and reports which results are still significant under each method at alpha = 0.05 — plus the classic p-value distribution diagnostic.
Why This Method?
Running many tests inflates false positives: at alpha = 0.05, 20 null tests yield one "significant" result by chance on average. Holm-Bonferroni controls the family-wise error rate (the chance of ANY false positive); Benjamini-Hochberg controls the false discovery rate (the expected fraction of false positives among discoveries). Reporting both, per test, is the standard defensible answer to "which results are real?"
What This Analysis Covers
- Holm and BH adjusted p-values for every test, with significance verdicts
- Significant counts under no correction / Holm / BH
- A 10-bin p-value histogram diagnosing real effects vs mostly-null noise
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {test_label, p_value}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting (rows_removed = invalid p dropped)
# p_name / label_name $ humanized user column names
# n_tests $ integer — valid tests entering the correction
# n_na / n_out_of_range $ integer — dropped: missing vs outside [0,1]
# alpha $ 0.05
# results_all $ data.frame(test, raw_p, holm_p, bh_p, sig_raw, sig_holm, sig_bh) sorted by raw_p
# adjusted_results_df $ results_all capped at 50 rows
# results_truncated $ logical — TRUE when results_all > 50 rows
# n_sig_raw/n_sig_holm/n_sig_bh $ integer — significant counts per method
# significance_summary_df $ data.frame(method, significant_count)
# p_value_distribution_df $ data.frame(p_range, count) — 10 bins of width 0.1
# dist_diagnosis $ character — computed narrative for the histogram shape
# metrics / json_output
# === /SHARED EXPORTS ===Step 1: Locate the mapped columns
initial_rows <- nrow(df)
p_name <- humanize_semantic("p_value", col_map)
label_name <- humanize_semantic("test_label", col_map)
if (!("p_value" %in% names(df))) {
stop(sprintf("column_mapping must map p_value to your p-value column('%s' not found).", p_name))
}
if (!("test_label" %in% names(df))) {
stop(sprintf("column_mapping must map test_label to your test-name column('%s' not found).", label_name))
}Step 2: Labels — character, blanks filled with a positional name
lab <- as.character(df$test_label)
blank <- is.na(lab) | trimws(lab) == ""
if (any(blank)) lab[blank] <- paste0("Test ", which(blank))Step 3: Coerce p-values (95% rule) and validate to [0, 1]
v <- df$p_value
if (!is.numeric(v)) {
conv <- suppressWarnings(as.numeric(as.character(v)))
n_orig <- sum(!is.na(v) & as.character(v) != "")
if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
v <- conv
} else {
stop(sprintf("The mapped p-value column('%s') is not numeric — map a column of raw p-values between 0 and 1.", p_name))
}
}
n_na <- sum(is.na(v))
n_out_of_range <- sum(!is.na(v) & (v < 0 | v > 1))
keep <- !is.na(v) & v >= 0 & v <= 1
p <- as.numeric(v[keep])
labels <- lab[keep]
final_rows <- length(p)
rows_removed <- initial_rows - final_rows
if (final_rows < 5) {
stop(sprintf(
"Only %d valid p-values remained in '%s' after dropping %d invalid (missing or outside 0-1) — need at least 5 tests to correct.",
final_rows, p_name, rows_removed))
}
n_tests <- final_rows
alpha <- 0.05Step 4: Adjust — Holm (family-wise) + Benjamini-Hochberg (FDR)
holm <- p.adjust(p, method = "holm")
bh <- p.adjust(p, method = "BH")
sig_raw <- p < alpha
sig_holm <- holm < alpha
sig_bh <- bh < alpha
n_sig_raw <- sum(sig_raw)
n_sig_holm <- sum(sig_holm)
n_sig_bh <- sum(sig_bh)
results_all <- data.frame(
test = labels,
raw_p = signif(p, 4),
holm_p = signif(holm, 4),
bh_p = signif(bh, 4),
sig_raw = ifelse(sig_raw, "Yes", "No"),
sig_holm = ifelse(sig_holm, "Yes", "No"),
sig_bh = ifelse(sig_bh, "Yes", "No"),
stringsAsFactors = FALSE
)
results_all <- results_all[order(results_all$raw_p), , drop = FALSE]
rownames(results_all) <- NULL
results_truncated <- nrow(results_all) > 50
adjusted_results_df <- head(results_all, 50)
significance_summary_df <- data.frame(
method = c("Uncorrected", "Holm-Bonferroni", "Benjamini-Hochberg"),
significant_count = c(n_sig_raw, n_sig_holm, n_sig_bh),
stringsAsFactors = FALSE
)Step 5: P-value distribution — 10 bins of width 0.1
breaks <- seq(0, 1, by = 0.1)
bin_idx <- pmin(pmax(findInterval(p, breaks, rightmost.closed = TRUE), 1), 10)
bin_labels <- paste0(format(breaks[-11], nsmall = 1), "-", format(breaks[-1], nsmall = 1))
counts <- as.integer(table(factor(bin_idx, levels = 1:10)))
p_value_distribution_df <- data.frame(
p_range = bin_labels,
count = counts,
stringsAsFactors = FALSE
)Classic diagnostic, narrated from the computed bin counts: a spike in the lowest bin suggests true effects; a near-uniform histogram suggests mostly null effects.
uniform_expected <- n_tests / 10
first_bin <- counts[1]
dist_diagnosis <- if (first_bin >= 2 * uniform_expected && first_bin >= 3) {
sprintf(paste0(
"The lowest bin(0.0-0.1) holds %d of %d p-values — about %.1fx the ~%.0f a ",
"uniform(all-null) distribution would put there. That spike near zero is the ",
"classic signature of genuinely non-null effects among your tests."),
first_bin, n_tests, first_bin / max(uniform_expected, 1e-9), uniform_expected)
} else if (first_bin > uniform_expected) {
sprintf(paste0(
"The lowest bin(0.0-0.1) holds %d of %d p-values, modestly above the ~%.0f a ",
"uniform(all-null) distribution would produce — weak evidence of a few true ",
"effects mixed into mostly null tests."),
first_bin, n_tests, uniform_expected)
} else {
sprintf(paste0(
"The histogram is close to uniform(lowest bin: %d of %d p-values vs ~%.0f ",
"expected under all-null) — consistent with mostly null effects, so treat any ",
"uncorrected significant results with suspicion."),
first_bin, n_tests, uniform_expected)
}
metrics <- list(
`Tests Evaluated` = n_tests,
`Invalid P-Values Dropped` = as.integer(rows_removed),
`Significant(Uncorrected)` = as.integer(n_sig_raw),
`Significant(Holm)` = as.integer(n_sig_holm),
`Significant(BH)` = as.integer(n_sig_bh),
`Smallest P-Value` = signif(min(p), 3)
)
json_output <- list(
answer = paste0(
"Of ", n_tests, " tests, ", n_sig_raw, " were significant uncorrected(p < 0.05); ",
"Holm-Bonferroni keeps ", n_sig_holm, " (strict family-wise control) and ",
"Benjamini-Hochberg keeps ", n_sig_bh, " (false-discovery-rate control). ",
dist_diagnosis
),
cards = lapply(
c("tldr", "overview", "preprocessing", "adjusted_results",
"significance_summary", "p_value_distribution"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed,
p_name = p_name, label_name = label_name,
n_tests = n_tests, n_na = n_na, n_out_of_range = n_out_of_range,
alpha = alpha,
results_all = results_all,
adjusted_results_df = adjusted_results_df,
results_truncated = results_truncated,
n_sig_raw = n_sig_raw, n_sig_holm = n_sig_holm, n_sig_bh = n_sig_bh,
significance_summary_df = significance_summary_df,
p_value_distribution_df = p_value_distribution_df,
dist_diagnosis = dist_diagnosis,
metrics = metrics, json_output = json_output
)
}
# Card: tldr (tldr)
card_tldr <- function(shared, df, params) {
survival_word <- if (shared$n_sig_raw == 0) {
"No result was significant even before correction."
} else if (shared$n_sig_holm == shared$n_sig_raw) {
"Every uncorrected finding survived even the strictest correction — these results are robust to multiple testing."
} else if (shared$n_sig_holm == 0 && shared$n_sig_bh == 0) {
"No result survived either correction — the uncorrected 'significant' findings are consistent with multiple-testing luck."
} else {
sprintf("%d uncorrected finding(s) did not survive Holm — likely multiple-testing artifacts.",
shared$n_sig_raw - shared$n_sig_holm)
}
text <- paste0(
"Of ", shared$n_tests, " tests, ", shared$n_sig_raw,
" were significant uncorrected(p < ", shared$alpha, "); Holm-Bonferroni keeps ",
shared$n_sig_holm, " (strict family-wise control) and Benjamini-Hochberg keeps ",
shared$n_sig_bh, " (false-discovery-rate control). ", survival_word
)
list(
title = "Executive Summary",
description = paste0("Multiple-testing correction across ", shared$n_tests, " tests"),
metrics = shared$metrics,
text = text
)
}