Standard Multiple Comparisons
Executive Summary

Executive Summary

Multiple-testing correction across 40 tests

Tests Evaluated
40
Invalid P-Values Dropped
0
Significant (Uncorrected)
11
Significant (Holm)
10
Significant (BH)
10
Smallest P-Value
0.0001
Of 40 tests, 11 were significant uncorrected (p < 0.05); Holm-Bonferroni keeps 10 (strict family-wise control) and Benjamini-Hochberg keeps 10 (false-discovery-rate control). 1 uncorrected finding(s) did not survive Holm — likely multiple-testing artifacts.
Interpretation

Of your 40 tests, 11 were significant uncorrected (p < 0.05). After applying Holm-Bonferroni correction—which controls the strictest family-wise error rate—10 survive. Benjamini-Hochberg, which tolerates a controlled fraction of false discoveries, also retains 10. One uncorrected finding (likely Metric 22, raw p = 0.0466) fails to survive either correction, flagged as a probable multiple-testing artifact. Your core discoveries are robust.

Overview

Analysis Overview

Holm-Bonferroni and Benjamini-Hochberg corrections applied to 40 p-values from Test Name.

N Tests40
Alpha0.05
N Significant Raw11
N Significant Holm10
N Significant Bh10
Interpretation

Running 40 independent tests at α = 0.05 inflates false positives: by chance alone, you'd expect ~2 spurious hits. Holm-Bonferroni controls the family-wise error rate—the probability of even ONE false positive across all 40 tests—making it the strictest standard. Benjamini-Hochberg controls the false discovery rate, the expected fraction of false positives among your discoveries, permitting slightly more findings in exchange for accepting a small proportion of errors. Here, 11 tests reached raw significance, but Holm retained only 10 and BH also retained 10. The two methods converge, suggesting the surviving 10 findings are genuinely robust against multiple-testing inflation.

Data Preparation

Data Quality

P-value validation and exclusions.

Initial Rows40
Final Rows40
Rows Removed0
N Missing0
N Out Of Range0
Interpretation

All 40 rows loaded from your data were valid: 0 p-values were missing and 0 fell outside the valid range [0, 1]. All 40 tests entered the correction analysis. Data quality was complete, so no adjustment to the correction factors was needed. The family size remained 40 throughout, ensuring both Holm and Benjamini-Hochberg thresholds reflect the full scope of your testing burden.

Data Table

Adjusted Results

Every test with raw, Holm, and BH p-values and significance verdicts.

TestRaw PHolm PBh PSig RawSig HolmSig Bh
Metric 90.00010.0030.0019YesYesYes
Metric 70.00010.00450.0019YesYesYes
Metric 40.00010.00550.0019YesYesYes
Metric 20.00030.01120.003YesYesYes
Metric 10.00060.02330.0048YesYesYes
Metric 60.00070.02560.0048YesYesYes
Metric 100.00090.02950.0048YesYesYes
Metric 80.0010.03350.0048YesYesYes
Metric 50.00110.03430.0048YesYesYes
Metric 30.00130.04040.0052YesYesYes
Metric 220.046610.1694YesNoNo
Metric 350.059610.1932NoNoNo
Metric 340.062810.1932NoNoNo
Metric 110.069910.1996NoNoNo
Metric 120.090710.2419NoNoNo
Metric 260.117810.2913NoNoNo
Metric 150.123810.2913NoNoNo
Metric 250.144310.3206NoNoNo
Metric 290.180710.3805NoNoNo
Metric 360.20610.4119NoNoNo
Metric 160.223210.4252NoNoNo
Metric 240.289610.5236NoNoNo
Metric 270.308510.5236NoNoNo
Metric 390.314110.5236NoNoNo
Metric 320.372410.5958NoNoNo
Metric 200.396710.6103NoNoNo
Metric 130.424510.6108NoNoNo
Metric 380.427610.6108NoNoNo
Metric 330.547710.732NoNoNo
Metric 190.577110.732NoNoNo
Metric 300.581610.732NoNoNo
Metric 400.585610.732NoNoNo
Metric 170.627410.7517NoNoNo
Metric 310.638910.7517NoNoNo
Metric 370.680410.7776NoNoNo
Metric 280.816110.8939NoNoNo
Metric 140.826910.8939NoNoNo
Metric 230.858510.9037NoNoNo
Metric 180.947710.972NoNoNo
Metric 210.976310.9763NoNoNo
Interpretation

The sorted results show your strongest finding is Metric 9 (raw p = 0.0001, Holm p = 0.003, BH p = 0.0019)—significant under all three standards. The top 10 tests all pass both corrections: Metric 9, 7, 4, 2, 1, 6, 10, 8, 5, and 3, with raw p-values ranging from 0.0001 to 0.0013. The 11th significant uncorrected test, Metric 22 (raw p = 0.0466), fails both corrections (Holm p = 1, BH p = 0.1694). This boundary case exemplifies how correction penalizes marginal raw p-values in a large family; it does not meet the stricter thresholds and is likely a false positive.

Visualization

What Survives Correction

Significant test counts: uncorrected vs Holm vs BH.

Interpretation

Of your 40 tests, the significance counts under each standard are: 11 uncorrected, 10 Holm-Bonferroni, and 10 Benjamini-Hochberg. The loss of 1 finding (from 11 to 10) under both corrections is modest, reflecting that your strongest signals are well below the noise floor. Holm and BH agree exactly here, both removing Metric 22 as a plausible artifact. This tight agreement suggests the surviving 10 findings are solid; the single casualty was borderline (raw p = 0.0466) and unlikely to be a true effect.

Visualization

P-Value Distribution

10-bin histogram of raw p-values — the real-effects-vs-noise diagnostic.

Interpretation

The raw p-value histogram reveals a strong signal: the lowest bin (0.0–0.1) contains 15 of 40 p-values, roughly 3.8 times the 4 expected under a uniform null distribution. This spike is the diagnostic hallmark of genuine non-null effects mixed with noise. The remaining bins (0.1–1.0) hold 25 p-values distributed relatively flatly, consistent with null tests. The enrichment in the left tail supports the reality of your top 10 surviving findings and explains why the corrections remove only 1 uncorrected result—most of the significant p-values are far enough below the noise threshold to survive stringent adjustment.

Methodology

Methodology

Statistical methodology and diagnostics for Multiple Comparisons Correction

Statistical Method

Multiple Comparisons Correction

Standard-library analysis: you ran many statistical tests — which results are still significant after correction? Upload your p-values and get Holm-Bonferroni and Benjamini-Hochberg adjusted p-values side by side, a count of what survives each method, and a p-value distribution diagnostic that hints whether your significant results reflect real effects or multiple-testing luck.

Data
N = 40 observations
Assumptions
  • P-values come from valid individual tests (the correction adjusts them; it cannot fix a broken test)
  • Tests are treated as one family; BH's FDR guarantee assumes independent or positively dependent tests
Limitations
  • Corrections reduce power — a true effect with a middling p-value may not survive, especially under Holm
  • The choice of family (which tests to correct together) is a judgment call the tool cannot make for you
  • P-values outside 0-1 or missing are dropped, which shrinks the family and changes the adjustment
Software & Citation
MCP Analytics · mcpanalytics.ai
Code Appendix

Analysis Code

Complete R source code for this analysis

Multiple Comparisons Correction — What Survives?

Takes a set of raw p-values from many tests, applies Holm-Bonferroni and Benjamini-Hochberg corrections, and reports which results are still significant under each method at alpha = 0.05 — plus the classic p-value distribution diagnostic.

Why This Method?

Running many tests inflates false positives: at alpha = 0.05, 20 null tests yield one "significant" result by chance on average. Holm-Bonferroni controls the family-wise error rate (the chance of ANY false positive); Benjamini-Hochberg controls the false discovery rate (the expected fraction of false positives among discoveries). Reporting both, per test, is the standard defensible answer to "which results are real?"

What This Analysis Covers

  • Holm and BH adjusted p-values for every test, with significance verdicts
  • Significant counts under no correction / Holm / BH
  • A 10-bin p-value histogram diagnosing real effects vs mostly-null noise

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {test_label, p_value}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed  $ row accounting (rows_removed = invalid p dropped)
  #   p_name / label_name    $ humanized user column names
  #   n_tests                $ integer — valid tests entering the correction
  #   n_na / n_out_of_range  $ integer — dropped: missing vs outside [0,1]
  #   alpha                  $ 0.05
  #   results_all            $ data.frame(test, raw_p, holm_p, bh_p, sig_raw, sig_holm, sig_bh) sorted by raw_p
  #   adjusted_results_df    $ results_all capped at 50 rows
  #   results_truncated      $ logical — TRUE when results_all > 50 rows
  #   n_sig_raw/n_sig_holm/n_sig_bh $ integer — significant counts per method
  #   significance_summary_df $ data.frame(method, significant_count)
  #   p_value_distribution_df $ data.frame(p_range, count) — 10 bins of width 0.1
  #   dist_diagnosis          $ character — computed narrative for the histogram shape
  #   metrics / json_output
  # === /SHARED EXPORTS ===

Step 1: Locate the mapped columns

initial_rows <- nrow(df)
  p_name <- humanize_semantic("p_value", col_map)
  label_name <- humanize_semantic("test_label", col_map)
  if (!("p_value" %in% names(df))) {
    stop(sprintf("column_mapping must map p_value to your p-value column(&#x27;%s' not found).", p_name))
  }
  if (!("test_label" %in% names(df))) {
    stop(sprintf("column_mapping must map test_label to your test-name column(&#x27;%s' not found).", label_name))
  }

Step 2: Labels — character, blanks filled with a positional name

lab <- as.character(df$test_label)
  blank <- is.na(lab) | trimws(lab) == ""
  if (any(blank)) lab[blank] <- paste0("Test ", which(blank))

Step 3: Coerce p-values (95% rule) and validate to [0, 1]

v <- df$p_value
  if (!is.numeric(v)) {
    conv <- suppressWarnings(as.numeric(as.character(v)))
    n_orig <- sum(!is.na(v) & as.character(v) != "")
    if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
      v <- conv
    } else {
      stop(sprintf("The mapped p-value column(&#x27;%s') is not numeric — map a column of raw p-values between 0 and 1.", p_name))
    }
  }
  n_na <- sum(is.na(v))
  n_out_of_range <- sum(!is.na(v) & (v < 0 | v > 1))
  keep <- !is.na(v) & v >= 0 & v <= 1
  p <- as.numeric(v[keep])
  labels <- lab[keep]
  final_rows <- length(p)
  rows_removed <- initial_rows - final_rows
  if (final_rows < 5) {
    stop(sprintf(
      "Only %d valid p-values remained in &#x27;%s' after dropping %d invalid (missing or outside 0-1) — need at least 5 tests to correct.",
      final_rows, p_name, rows_removed))
  }
  n_tests <- final_rows
  alpha <- 0.05

Step 4: Adjust — Holm (family-wise) + Benjamini-Hochberg (FDR)

holm <- p.adjust(p, method = "holm")
  bh   <- p.adjust(p, method = "BH")
  sig_raw  <- p    < alpha
  sig_holm <- holm < alpha
  sig_bh   <- bh   < alpha
  n_sig_raw  <- sum(sig_raw)
  n_sig_holm <- sum(sig_holm)
  n_sig_bh   <- sum(sig_bh)

  results_all <- data.frame(
    test     = labels,
    raw_p    = signif(p, 4),
    holm_p   = signif(holm, 4),
    bh_p     = signif(bh, 4),
    sig_raw  = ifelse(sig_raw,  "Yes", "No"),
    sig_holm = ifelse(sig_holm, "Yes", "No"),
    sig_bh   = ifelse(sig_bh,   "Yes", "No"),
    stringsAsFactors = FALSE
  )
  results_all <- results_all[order(results_all$raw_p), , drop = FALSE]
  rownames(results_all) <- NULL
  results_truncated <- nrow(results_all) > 50
  adjusted_results_df <- head(results_all, 50)

  significance_summary_df <- data.frame(
    method = c("Uncorrected", "Holm-Bonferroni", "Benjamini-Hochberg"),
    significant_count = c(n_sig_raw, n_sig_holm, n_sig_bh),
    stringsAsFactors = FALSE
  )

Step 5: P-value distribution — 10 bins of width 0.1

breaks <- seq(0, 1, by = 0.1)
  bin_idx <- pmin(pmax(findInterval(p, breaks, rightmost.closed = TRUE), 1), 10)
  bin_labels <- paste0(format(breaks[-11], nsmall = 1), "-", format(breaks[-1], nsmall = 1))
  counts <- as.integer(table(factor(bin_idx, levels = 1:10)))
  p_value_distribution_df <- data.frame(
    p_range = bin_labels,
    count = counts,
    stringsAsFactors = FALSE
  )

Classic diagnostic, narrated from the computed bin counts: a spike in the lowest bin suggests true effects; a near-uniform histogram suggests mostly null effects.

uniform_expected <- n_tests / 10
  first_bin <- counts[1]
  dist_diagnosis <- if (first_bin >= 2 * uniform_expected && first_bin >= 3) {
    sprintf(paste0(
      "The lowest bin(0.0-0.1) holds %d of %d p-values — about %.1fx the ~%.0f a ",
      "uniform(all-null) distribution would put there. That spike near zero is the ",
      "classic signature of genuinely non-null effects among your tests."),
      first_bin, n_tests, first_bin / max(uniform_expected, 1e-9), uniform_expected)
  } else if (first_bin > uniform_expected) {
    sprintf(paste0(
      "The lowest bin(0.0-0.1) holds %d of %d p-values, modestly above the ~%.0f a ",
      "uniform(all-null) distribution would produce — weak evidence of a few true ",
      "effects mixed into mostly null tests."),
      first_bin, n_tests, uniform_expected)
  } else {
    sprintf(paste0(
      "The histogram is close to uniform(lowest bin: %d of %d p-values vs ~%.0f ",
      "expected under all-null) — consistent with mostly null effects, so treat any ",
      "uncorrected significant results with suspicion."),
      first_bin, n_tests, uniform_expected)
  }

  metrics <- list(
    `Tests Evaluated`         = n_tests,
    `Invalid P-Values Dropped` = as.integer(rows_removed),
    `Significant(Uncorrected)` = as.integer(n_sig_raw),
    `Significant(Holm)`      = as.integer(n_sig_holm),
    `Significant(BH)`        = as.integer(n_sig_bh),
    `Smallest P-Value`        = signif(min(p), 3)
  )

  json_output <- list(
    answer = paste0(
      "Of ", n_tests, " tests, ", n_sig_raw, " were significant uncorrected(p < 0.05); ",
      "Holm-Bonferroni keeps ", n_sig_holm, " (strict family-wise control) and ",
      "Benjamini-Hochberg keeps ", n_sig_bh, " (false-discovery-rate control). ",
      dist_diagnosis
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "adjusted_results",
        "significance_summary", "p_value_distribution"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed,
    p_name = p_name, label_name = label_name,
    n_tests = n_tests, n_na = n_na, n_out_of_range = n_out_of_range,
    alpha = alpha,
    results_all = results_all,
    adjusted_results_df = adjusted_results_df,
    results_truncated = results_truncated,
    n_sig_raw = n_sig_raw, n_sig_holm = n_sig_holm, n_sig_bh = n_sig_bh,
    significance_summary_df = significance_summary_df,
    p_value_distribution_df = p_value_distribution_df,
    dist_diagnosis = dist_diagnosis,
    metrics = metrics, json_output = json_output
  )
}

# Card: tldr (tldr)
card_tldr <- function(shared, df, params) {
  survival_word <- if (shared$n_sig_raw == 0) {
    "No result was significant even before correction."
  } else if (shared$n_sig_holm == shared$n_sig_raw) {
    "Every uncorrected finding survived even the strictest correction — these results are robust to multiple testing."
  } else if (shared$n_sig_holm == 0 && shared$n_sig_bh == 0) {
    "No result survived either correction — the uncorrected &#x27;significant' findings are consistent with multiple-testing luck."
  } else {
    sprintf("%d uncorrected finding(s) did not survive Holm — likely multiple-testing artifacts.",
            shared$n_sig_raw - shared$n_sig_holm)
  }
  text <- paste0(
    "Of ", shared$n_tests, " tests, ", shared$n_sig_raw,
    " were significant uncorrected(p < ", shared$alpha, "); Holm-Bonferroni keeps ",
    shared$n_sig_holm, " (strict family-wise control) and Benjamini-Hochberg keeps ",
    shared$n_sig_bh, " (false-discovery-rate control). ", survival_word
  )
  list(
    title = "Executive Summary",
    description = paste0("Multiple-testing correction across ", shared$n_tests, " tests"),
    metrics = shared$metrics,
    text = text
  )
}
Your data has more stories to tell. Run any analysis on your own data — 60+ validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No Signup Sign Up Free

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing