Standard Equivalence
Executive Summary

Executive Summary

Is Fill Volume ml practically the same across Production Line?

Observations
120
Mean Difference
-1.038
Equivalence Margin
5
Verdict
Equivalent
TOST p-value
< 0.001
Standard t-test p
0.173
Fill Volume ml looks practically the same in the two Production Line groups: the TOST rules out any difference beyond 5 in either direction (p = < 0.001). Observed difference: Line A averages 1.038 lower than Line B (90% CI -2.294 to 0.218, margin 5).
Suggested Interpretation

Yes: your two bottling lines are filling to practically the same volume within 5 ml. The TOST rules out any difference beyond that margin with p < 0.001. Line A averages 250.093 ml and Line B averages 251.131 ml—a gap of only 1.038 ml, well inside your tolerance. The 90% confidence interval (−2.294 to 0.218 ml) sits entirely within the ±5 ml band, the visual proof of equivalence. A standard t-test would only say "not significantly different" (p = 0.173); the TOST goes further and actively demonstrates equivalence.

Overview

Analysis Overview

Equivalence (TOST) analysis of Fill Volume ml between the Production Line groups Line A and Line B (120 observations).

N Observations120
N Groups2
Margin5
Mean Difference-1.038
Suggested Interpretation

The analysis directly answers your question by inverting the usual hypothesis test. Instead of asking "is there a difference?", it asks "can we rule out any difference larger than 5 ml?" This is a stronger claim than merely failing to find a difference. The TOST procedure runs two one-sided tests, both of which must pass to declare equivalence. With 120 observations across Line A and Line B, the analysis has the power to make a positive statement: the lines are filling to practically the same volume within your 5 ml tolerance.

Data Preparation

Data Quality

Row and group cleaning applied before testing.

Initial Rows120
Final Rows120
Rows Removed0
Groups Dropped0
Suggested Interpretation

All 120 rows loaded were used; no missing Fill Volume values were dropped. Line A contributed 56 observations and Line B contributed 64, both well above the minimum of 3 per group. The data were clean and complete, with no groups excluded. This two-group structure exactly matches the equivalence comparison design.

Visualization

Difference vs the Equivalence Margin

Mean difference in Fill Volume ml with 90% and 95% CIs against the margin band.

Suggested Interpretation

The confidence interval plot is the decision picture. The observed difference (Line A minus Line B) is −1.038 ml, shown with its 90% TOST-consistent interval of −2.294 to 0.218 ml. Both bounds lie comfortably inside the margin band (−5 to +5 ml marked by the horizontal reference lines), which is exactly what equivalence requires. The wider 95% interval (−2.538 to 0.462 ml) is also inside the band but is shown for conventional reference. The point estimate sits close to zero, indicating the lines are nearly balanced; the full interval confirms no practically meaningful divergence exists.

Data Table

The Two One-Sided Tests

TOST breakdown for Fill Volume ml, with the standard t-test for contrast.

TestStatisticDfP ValueInterpretation
TOST lower one-sided (rules out a large deficit)5.231117.6< 0.001Tests whether Line A minus Line B is above the lower margin of -5; a deficit beyond the margin is ruled out.
TOST upper one-sided (rules out a large excess)-7.972117.6< 0.001Tests whether Line A minus Line B is below the upper margin of +5; an excess beyond the margin is ruled out.
TOST combined (equivalence verdict)117.6< 0.001The larger of the two one-sided p-values; equivalence requires BOTH to pass. Both pass, so the groups are statistically equivalent within the margin of 5.
Standard Welch t-test (difference verdict, for contrast)-1.371117.60.173Asks the OPPOSITE question — is there evidence of ANY difference? It finds no significant difference. A non-significant t-test alone never demonstrates equivalence; only the TOST above can.
Suggested Interpretation

Two one-sided tests carry the verdict. The lower test (t = 5.231, df = 117.6, p < 0.001) rules out Line A being far below Line B. The upper test (t = −7.972, df = 117.6, p < 0.001) rules out Line A being far above Line B. Both reject, so the TOST combined p-value is < 0.001 and equivalence is declared. For contrast, the standard Welch t-test (t = −1.371, p = 0.173) only says there is no significant difference—a much weaker statement. Here both verdicts agree, but for different reasons: the TOST adds affirmative evidence that any difference larger than 5 ml is ruled out.

Data Table

Group Statistics

n, mean, spread, median and 95% CI of Fill Volume ml per Production Line group.

GroupNMeanSDMedianCI LowCI High
Line A56250.13.739249.7249.1251.1
Line B64251.14.553250.9250252.3
Suggested Interpretation

Line A (n=56) fills to a mean of 250.093 ml with SD 3.739 ml; Line B (n=64) fills to 251.131 ml with SD 4.553 ml. The raw gap is 1.038 ml, which is only about one-fifth of your 5 ml tolerance. The two groups have comparable spread (pooled SD = 4.193 ml), and both medians (249.685 and 250.92 ml) are close to their respective means, indicating roughly symmetric distributions. Neither group shows outlier-driven drift.

Data Table

Methods & Margin Disclosure

How the TOST verdict is computed and where the margin came from.

ItemDetail
MethodTwo one-sided t-tests (TOST) on the Welch statistic
ComparisonDifference in mean Fill Volume ml: Line A minus Line B = -1.038 (standard error 0.757)
Equivalence margin5 (absolute, in the units of Fill Volume ml)
Margin sourceuser-specified, in the units of Fill Volume ml
Significance levelalpha = 0.05 per one-sided test
Confidence intervals90% CI is the TOST-consistent interval (two 5% one-sided tests); the conventional 95% CI is shown for comparison and is always wider
Degrees of freedomWelch-Satterthwaite: 117.6
Verdict ruleEquivalent only if BOTH one-sided tests reject, i.e. the whole 90% CI lies inside the margin band
AssumptionsRoughly normal group means (t-based), independent observations, unequal variances allowed (Welch)
Suggested Interpretation

The TOST runs two one-sided Welch t-tests at α = 0.05 each, testing whether the mean difference falls outside the margin band (−5 to +5 ml). Equivalence is concluded only when both tests reject, which is mathematically equivalent to the 90% confidence interval lying entirely inside the band—hence the 90% interval is the decision interval here, not the familiar 95%. Your 5 ml margin was user-specified in the units of Fill Volume ml, the gold standard for margin choice because it encodes what difference actually matters operationally. Welch degrees of freedom (117.6) allow unequal variances. Equivalence is always relative to its margin: the claim is "equivalent within 5 ml," never "identical."

Methodology

Methodology

Statistical methodology and diagnostics for Equivalence & Non-Inferiority (TOST)

Statistical Method

Equivalence & Non-Inferiority (TOST)

Standard-library analysis: are these two things practically the SAME within a margin? A standard t-test can never prove sameness — 'not significantly different' may just mean too little data. This tool runs the classical two one-sided tests (TOST): the mean difference with its 90% (TOST-consistent) and 95% confidence intervals plotted against your equivalence margin band, both one-sided tests spelled out, the standard t-test alongside for contrast, and a one-sided non-inferiority variant when only a deficit would matter. Supply a margin in your outcome's units, or let the tool use a clearly-labeled default of 0.2 x pooled SD.

Data
N = 120 observations
Assumptions
  • The outcome is numeric (or cleanly convertible) and the group column has exactly two usable levels
  • Observations are independent; group means are approximately t-distributed (Welch machinery, unequal variances allowed)
  • The equivalence margin is meaningful in the outcome's units — ideally chosen from practical stakes, not statistics
Limitations
  • An equivalence verdict is always relative to its margin — 'equivalent within the margin' is the full claim, never 'identical'
  • When no margin is supplied the tool defaults to 0.2 x pooled SD and says so; a domain-justified margin makes the verdict decision-ready
  • Small samples rarely establish equivalence even when groups truly match — the intervals are simply too wide
  • Exactly two groups are compared; use the group comparison tool for many-group questions
Software & Citation
MCP Analytics · mcpanalytics.ai
Code Appendix

Analysis Code

Complete R source code for this analysis

Equivalence & Non-Inferiority (TOST) — Practically the Same?

Tests whether a numeric outcome is practically EQUIVALENT between two groups, within a stated margin — the question a standard t-test cannot answer. Runs the classical two one-sided tests (TOST) on the Welch statistic, shows the TOST-consistent 90% confidence interval next to the familiar 95% interval, and reports the standard t-test alongside with a computed explanation of why "not significantly different" is not the same claim as "equivalent". A one-sided non-inferiority variant is available via module parameters.

Why This Method?

A non-significant t-test only says the data failed to prove a difference — it never proves sameness. TOST reverses the burden of proof: equivalence is concluded only when the data actively rule out a difference larger than the margin in BOTH directions. It is the standard approach in bioequivalence, method validation, and "did the change break anything" testing.

What This Analysis Covers

  • Mean difference with 90% (TOST-consistent) and 95% confidence intervals

plotted against the equivalence margin band

  • Both one-sided tests, the combined TOST verdict, and the standard

Welch t-test side by side

  • Per-group statistics and a full methods disclosure (including how the

margin was chosen)

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {outcome, group}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed/n_na_outcome  $ row accounting
  #   outcome_h / group_h   $ humanized user names for the mapped columns
  #   l1 / l2               $ the two group levels (factor order)
  #   n1/n2, m1/m2, s1/s2   $ per-group n, mean, sd
  #   d / se / df_w         $ mean difference (l1 - l2), Welch SE + df
  #   pooled_sd             $ pooled standard deviation
  #   margin / margin_is_default / margin_source  $ the equivalence margin
  #   test_type             $ "equivalence" | "non_inferiority"
  #   t_lower/p_lower, t_upper/p_upper, p_tost, equivalent  $ TOST results
  #   ci90_low/ci90_high, ci95_low/ci95_high  $ the two intervals on d
  #   t_std / p_std / std_sig                 $ standard Welch t-test on d
  #   ref_level/new_level/higher_is_better/d_ni/t_ni/p_ni/ni_bound/noninferior
  #                                           $ non-inferiority variant (NI mode)
  #   verdict_label / verdict_phrase / quadrant_text  $ computed conclusions
  #   dropped_groups_df     $ groups dropped for n < 3
  #   ci_margin_df / tost_df / group_summary_df / methods_df  $ card datasets
  #   metrics / json_output
  # === /SHARED EXPORTS ===

Step 1: Resolve mapped columns (humanized for all prose)

initial_rows <- nrow(df)
  outcome_h <- humanize_semantic("outcome", col_map)
  group_h   <- humanize_semantic("group", col_map)
  if (!("outcome" %in% names(df)) || !("group" %in% names(df))) {
    stop(sprintf("Equivalence testing needs both &#x27;%s' (the numeric outcome) and '%s' (the two groups) mapped.",
                 outcome_h, group_h))
  }

Step 2: Coerce the outcome to numeric (95% rule); drop NA-outcome rows

v <- df$outcome
  if (!is.numeric(v)) {
    ch <- as.character(v)
    non_blank <- !is.na(ch) & trimws(ch) != ""
    conv <- suppressWarnings(as.numeric(ch))
    if (sum(non_blank) == 0 ||
        sum(!is.na(conv[non_blank])) < 0.95 * sum(non_blank)) {
      stop(sprintf("The outcome column &#x27;%s' does not look numeric — fewer than 95%% of its values parse as numbers. Pick a numeric column to test for equivalence.",
                   outcome_h))
    }
    v <- conv
  }
  df$outcome <- v

  g <- as.character(df$group)
  g[is.na(g) | trimws(g) == ""] <- "Missing"

  keep <- !is.na(df$outcome)
  n_na_outcome <- sum(!keep)
  df <- df[keep, , drop = FALSE]
  g  <- g[keep]
  if (nrow(df) == 0) {
    stop(sprintf("No rows with a usable numeric value in &#x27;%s' remained after cleaning.", outcome_h))
  }

Step 3: Clean the groups — drop n<3 (reported), require exactly 2 levels

tab <- table(g)
  small <- names(tab)[tab < 3]
  dropped_groups_df <- data.frame(group = character(0), n = integer(0),
                                  stringsAsFactors = FALSE)
  if (length(small) > 0) {
    dropped_groups_df <- data.frame(group = small, n = as.integer(tab[small]),
                                    stringsAsFactors = FALSE)
    sel <- !(g %in% small)
    df <- df[sel, , drop = FALSE]
    g  <- g[sel]
  }

  gf <- factor(g)
  k  <- nlevels(gf)
  if (k < 2) {
    stop(sprintf("Equivalence testing needs exactly 2 groups in &#x27;%s' with 3 or more rows each; only %d usable group(s) remained after cleaning. Check that '%s' really splits the data into two groups.",
                 group_h, k, group_h))
  }
  if (k > 2) {
    stop(sprintf("Equivalence testing compares exactly two groups, but &#x27;%s' has %d usable levels (%s). Filter the data to two groups, or use the group comparison tool for a many-group question.",
                 group_h, k, paste(levels(gf), collapse = ", ")))
  }
  y <- df$outcome
  final_rows <- length(y)
  rows_removed <- initial_rows - final_rows
  if (final_rows < 10) {
    stop(sprintf("Only %d usable rows remained — at least 10 are needed to test &#x27;%s' for equivalence.", final_rows, outcome_h))
  }
  if (isTRUE(stats::var(y) == 0)) {
    stop(sprintf("The outcome &#x27;%s' has no variation at all (every value is identical) — an equivalence margin cannot be assessed.", outcome_h))
  }

Step 4: Per-group statistics (Welch machinery)

l1 <- levels(gf)[1]; l2 <- levels(gf)[2]
  x1 <- y[gf == l1]; x2 <- y[gf == l2]
  n1 <- length(x1); n2 <- length(x2)
  m1 <- mean(x1); m2 <- mean(x2)
  s1 <- stats::sd(x1); s2 <- stats::sd(x2)
  pooled_sd <- sqrt(((n1 - 1) * s1^2 + (n2 - 1) * s2^2) / (n1 + n2 - 2))
  d  <- m1 - m2
  se <- sqrt(s1^2 / n1 + s2^2 / n2)
  if (is.na(se) || se == 0) {
    stop(sprintf("The outcome &#x27;%s' is essentially constant within each '%s' group — the standard error is zero, so no equivalence test can be computed.",
                 outcome_h, group_h))
  }
  df_w <- (s1^2 / n1 + s2^2 / n2)^2 /
    ((s1^2 / n1)^2 / (n1 - 1) + (s2^2 / n2)^2 / (n2 - 1))

Step 5: The equivalence margin — user-specified or a labeled default

raw_margin <- params$equivalence_margin %||% params$margin %||% NULL
  margin_is_default <- is.null(raw_margin)
  if (margin_is_default) {
    margin <- 0.2 * pooled_sd
    margin_source <- "default: 0.2 x pooled SD(no margin was supplied)"
  } else {
    margin <- suppressWarnings(as.numeric(raw_margin))
    if (is.na(margin) || margin <= 0) {
      stop(sprintf("The equivalence_margin parameter must be a positive number in the same units as &#x27;%s' (got '%s').",
                   outcome_h, as.character(raw_margin)))
    }
    margin_source <- sprintf("user-specified, in the units of %s", outcome_h)
  }

  test_type <- tolower(as.character(params$test_type %||% "equivalence"))
  if (!(test_type %in% c("equivalence", "non_inferiority"))) {
    stop(sprintf("test_type must be &#x27;equivalence' or 'non_inferiority' (got '%s').", test_type))
  }

Step 6: TOST — two one-sided Welch t-tests, hand-rolled on pt()

Lower test: H0 d <= -margin vs H1 d > -margin (rules out "much lower"). Upper test: H0 d >= +margin vs H1 d < +margin (rules out "much higher").

t_lower <- (d + margin) / se
  p_lower <- stats::pt(t_lower, df_w, lower.tail = FALSE)
  t_upper <- (d - margin) / se
  p_upper <- stats::pt(t_upper, df_w, lower.tail = TRUE)
  p_tost  <- max(p_lower, p_upper)
  equivalent <- !is.na(p_tost) && p_tost < 0.05

The TOST-consistent 90% interval and the familiar 95% interval

ci90_half <- stats::qt(0.95, df_w) * se
  ci95_half <- stats::qt(0.975, df_w) * se
  ci90_low <- d - ci90_half; ci90_high <- d + ci90_half
  ci95_low <- d - ci95_half; ci95_high <- d + ci95_half

Step 7: The standard Welch t-test alongside (the pedagogical foil)

t_std <- d / se
  p_std <- 2 * stats::pt(abs(t_std), df_w, lower.tail = FALSE)
  std_sig <- !is.na(p_std) && p_std < 0.05

Step 8: Non-inferiority variant (one-sided margin), when requested

ref_level <- NULL; new_level <- NULL; higher_is_better <- TRUE
  d_ni <- NA_real_; t_ni <- NA_real_; p_ni <- NA_real_
  ni_bound <- NA_real_; noninferior <- NA
  if (test_type == "non_inferiority") {
    ref_level <- as.character(params$reference_group %||% l1)
    if (!(ref_level %in% c(l1, l2))) {
      stop(sprintf("reference_group &#x27;%s' is not a level of '%s' — the two groups are '%s' and '%s'.",
                   ref_level, group_h, l1, l2))
    }
    new_level <- if (ref_level == l1) l2 else l1
    higher_is_better <- !isFALSE(params$higher_is_better)
    m_ref <- if (ref_level == l1) m1 else m2
    m_new <- if (new_level == l1) m1 else m2
    d_ni <- m_new - m_ref
    if (higher_is_better) {

H0: new is worse by at least the margin (d_ni <= -margin)

t_ni <- (d_ni + margin) / se
      p_ni <- stats::pt(t_ni, df_w, lower.tail = FALSE)
      ni_bound <- d_ni - stats::qt(0.95, df_w) * se
      noninferior <- !is.na(p_ni) && p_ni < 0.05
    } else {

Lower outcome is better: H0 new is worse by at least the margin (d_ni >= +margin)

t_ni <- (d_ni - margin) / se
      p_ni <- stats::pt(t_ni, df_w, lower.tail = TRUE)
      ni_bound <- d_ni + stats::qt(0.95, df_w) * se
      noninferior <- !is.na(p_ni) && p_ni < 0.05
    }
  }

Step 9: Verdicts + the computed t-test-vs-TOST explanation

if (test_type == "non_inferiority") {
    verdict_label <- if (isTRUE(noninferior)) "Non-inferior" else "Not established"
    verdict_phrase <- if (isTRUE(noninferior)) {
      sprintf("the data support that %s is not meaningfully worse than %s on %s(by more than %s)",
              new_level, ref_level, outcome_h, fmt_num(margin))
    } else {
      sprintf("the data do NOT rule out that %s is worse than %s on %s by more than %s — non-inferiority is not established",
              new_level, ref_level, outcome_h, fmt_num(margin))
    }
  } else {
    verdict_label <- if (equivalent) "Equivalent" else "Not established"
    verdict_phrase <- if (equivalent) {
      sprintf("the two groups are statistically equivalent on %s within a margin of %s — differences larger than the margin are ruled out in both directions",
              outcome_h, fmt_num(margin))
    } else {
      sprintf("equivalence within a margin of %s could not be established for %s — the data leave room for a difference larger than the margin",
              fmt_num(margin), outcome_h)
    }
  }

The four-quadrant teaching text: standard t-test verdict x TOST verdict.

quadrant_text <- if (!std_sig && equivalent) {
    paste0(
      "Here the two verdicts agree for the right reason: the standard t-test finds no ",
      "significant difference(p = ", fmt_p(p_std), ") AND the TOST actively demonstrates ",
      "equivalence(p = ", fmt_p(p_tost), "). Note these are different claims — the t-test ",
      "alone would only say the data failed to prove a difference, which can also happen ",
      "with too little data. The TOST adds the positive evidence: any difference larger ",
      "than ", fmt_num(margin), " is ruled out."
    )
  } else if (!std_sig && !equivalent) {
    paste0(
      "This is the case the tool exists for: the standard t-test finds no significant ",
      "difference(p = ", fmt_p(p_std), "), which is often misread as &#x27;the groups are the ",
      "same&#x27;. But the TOST does NOT conclude equivalence (p = ", fmt_p(p_tost), ") — the ",
      "90% confidence interval(", fmt_num(ci90_low), " to ", fmt_num(ci90_high),
      ") extends beyond the margin of ", fmt_num(margin), ", so a practically meaningful ",
      "difference has not been ruled out. &#x27;Not significantly different' here reflects ",
      "limited evidence, not demonstrated sameness — more data or a wider(justified) ",
      "margin would be needed to claim equivalence."
    )
  } else if (std_sig && equivalent) {
    paste0(
      "An instructive combination: the standard t-test says the difference is statistically ",
      "significant(p = ", fmt_p(p_std), "), yet the TOST still concludes equivalence ",
      "(p = ", fmt_p(p_tost), "). Both are correct — the difference is real but small: the ",
      "entire 90% confidence interval(", fmt_num(ci90_low), " to ", fmt_num(ci90_high),
      ") sits inside the margin of ", fmt_num(margin), ". Statistical significance measures ",
      "detectability, not practical importance."
    )
  } else {
    paste0(
      "Here the verdicts agree that the groups differ: the standard t-test is significant ",
      "(p = ", fmt_p(p_std), ") and the TOST fails to conclude equivalence(p = ",
      fmt_p(p_tost), ") — the observed difference of ", fmt_num(d), " is not compatible ",
      "with sameness within the margin of ", fmt_num(margin), ". Note the two tests ask ",
      "different questions; they simply reach compatible answers on this data."
    )
  }

Step 11: Metrics + JSON answer

metrics <- list(
    `Observations`       = final_rows,
    `Mean Difference`    = round(d, 3),
    `Equivalence Margin` = round(margin, 3),
    `Verdict`            = verdict_label,
    `TOST p-value`       = fmt_p(if (test_type == "non_inferiority") p_ni else p_tost),
    `Standard t-test p`  = fmt_p(p_std)
  )

  json_output <- list(
    answer = paste0(
      "TOST equivalence analysis of ", outcome_h, " between the two ", group_h,
      " groups ", l1, " and ", l2, " (", format(final_rows, big.mark = ","),
      " rows): the mean difference is ", fmt_num(d), " (90% CI ", fmt_num(ci90_low),
      " to ", fmt_num(ci90_high), ") against an equivalence margin of ",
      fmt_num(margin),
      if (margin_is_default) " (defaulted to 0.2 x pooled SD because no margin was supplied)" else "",
      ". Verdict: ", verdict_phrase, " (TOST p = ",
      fmt_p(if (test_type == "non_inferiority") p_ni else p_tost),
      "). For contrast, the standard Welch t-test p = ", fmt_p(p_std),
      " — note that a non-significant t-test alone would not demonstrate equivalence."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "equivalence_plot",
        "tost_table", "group_summary", "methods"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed, n_na_outcome = n_na_outcome,
    outcome_h = outcome_h, group_h = group_h,
    l1 = l1, l2 = l2, n1 = n1, n2 = n2, m1 = m1, m2 = m2, s1 = s1, s2 = s2,
    d = d, se = se, df_w = df_w, pooled_sd = pooled_sd,
    margin = margin, margin_is_default = margin_is_default,
    margin_source = margin_source, test_type = test_type,
    t_lower = t_lower, p_lower = p_lower,
    t_upper = t_upper, p_upper = p_upper,
    p_tost = p_tost, equivalent = equivalent,
    ci90_low = ci90_low, ci90_high = ci90_high,
    ci95_low = ci95_low, ci95_high = ci95_high,
    t_std = t_std, p_std = p_std, std_sig = std_sig,
    ref_level = ref_level, new_level = new_level,
    higher_is_better = higher_is_better,
    d_ni = d_ni, t_ni = t_ni, p_ni = p_ni, ni_bound = ni_bound,
    noninferior = noninferior,
    verdict_label = verdict_label, verdict_phrase = verdict_phrase,
    quadrant_text = quadrant_text,
    dropped_groups_df = dropped_groups_df,
    ci_margin_df = ci_margin_df, tost_df = tost_df,
    group_summary_df = group_summary_df, methods_df = methods_df,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell. Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No Signup Sign Up Free

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing