Can I Stop?
Stop — a winner has separated
The short answer
You can stop for a winner: Variant B has separated from Control by 3.84 pp, and the always-valid confidence sequence (1.34 pp to 6.34 pp) excludes zero—even though the experiment was re-tested after every one of its 6,000 observations.
The detail
The boundary was first crossed at observation 1,671. A fixed-horizon test on the same data would have shown p below 0.05 at 929 points where the always-valid boundary was not crossed, across 12 separate excursions. The always-valid p-value is < 0.001 and the fixed-horizon p-value is also < 0.001 at this sample size. The guarantee costs precision: the always-valid interval is 1.56 times wider than the fixed-horizon interval (2.50 pp vs 1.60 pp half-width), and reaching today's fixed-horizon tightness would require about 15,512 observations (2.59 times the current sample).
What this can't tell you
The guarantee covers this one comparison as specified. It does not protect a metric, segment, or arm picked after seeing the results.
Analysis Overview
Always-valid sequential monitoring of Variant B against Control across 6,000 observations.
The short answer
A mixture sequential probability ratio test replaces the ordinary 5% test with one valid at every look. The fixed-horizon test earns its 5% false-positive rate only if you decide your sample size before seeing data and look once. Here, 6,000 observations were ordered by Signup Date and re-tested after every observation; the control arm was identified by name, not outcome. The cumulative effect and both boundaries—always-valid and fixed-horizon—are computed on your data so you can compare them yourself.
The detail
The analysis monitored 6,000 observations (n_control = 2,975; n_treatment = 3,025) at 5,973 analysis points. The effect is the difference in mean Converted between Variant B and Control. A mixture sequential probability ratio test with a normal mixture over the true effect produces a boundary valid at all simultaneous looks. An ordinary Welch t-test is also run at every point so both corridors can be displayed on the same data.
What this can't tell you
This boundary protects this one comparison as specified. It does not protect a metric, segment, or arm chosen after seeing results.
Data Quality
Row accounting, arm sizes, and how the ordering was resolved.
The short answer
All 6,000 rows were analysed; none were incomplete. The outcome Converted is already binary (0 or 1), so the effect is a difference in conversion rates. Control has 2,975 rows and Variant B has 3,025. When 5,850 rows tied on arrival time, they kept their original file order and stepped through the monitoring path one at a time.
The detail
6,000 rows loaded; 6,000 analysed; 0 rows removed. The Converted column holds only 0 and 1 values, making the effect a difference in proportions. Control contributed 2,975 observations and Variant B contributed 3,025. Tied arrival values (5,850 rows) were resolved by preserving the original file order, so each observation enters the monitoring sequence individually.
What this can't tell you
This analysis cannot verify from the file that assignment to Assigned Group was actually random, or that each row represents a distinct person. Both are inherited assumptions about how the experiment was run.
Effect and Boundaries Over the Experiment
The cumulative effect with the always-valid and fixed-horizon boundaries drawn over it.
The short answer
The cumulative effect traces an upward pattern from observation 1,671 onward, leaving the always-valid corridor at that point and remaining outside it through the final observation. Both corridors narrow as data accumulates, but the always-valid boundary (outer pair) narrows more slowly because it pays for every look you might take.
The detail
The cumulative effect estimate (Variant B minus Control) starts at 4.4444 pp at observation 28 and drifts downward through observation 154 (−3.8961 pp), then climbs steadily from observation 218 onward. At observation 1,671 the effect exits the always-valid corridor on the upper side and stays outside. The inner corridor (fixed-horizon 0.05 boundary) is narrower everywhere; this is why watching it makes teams stop too early. Both corridors narrow as n increases, but the always-valid corridor widens more slowly because it is protecting against all possible looks, not just one.
What this can't tell you
The pattern shown is descriptive of the monitoring path as data arrived in calendar order. It does not explain why the effect moved as it did.
The Two p-Values Side by Side
The fixed-horizon p-value and the always-valid p-value at every analysis point.
The short answer
The fixed-horizon p-value wanders below 0.05 and back above it across 929 analysis points where the always-valid p-value does not cross that threshold. The always-valid p-value only ever moves down and is a legitimate stopping criterion at any point.
The detail
Both p-values were computed at all 5,973 analysis points on the same cumulative data. The fixed-horizon p-value fell below 0.05 at 5,219 points in 12 separate excursions; the always-valid p-value (the running minimum of the inverse mixture likelihood ratio) crossed below 0.05 at 4,290 points. The 929 points where only the fixed-horizon test fired are moments this experiment could have been stopped on evidence that had not actually arrived. The fixed-horizon series is free to wander because it is recomputed from scratch at each point; the always-valid series is monotonic because it records the best evidence seen so far.
What this can't tell you
These counts assume a look after every observation—the most aggressive peeking schedule. A weekly review cadence would see fewer opportunities to be misled, but not zero.
The Decision, Line by Line
Everything the stopping decision rests on.
| Item | Detail |
|---|---|
| Decision | Stop — a winner has separated |
| Always-valid boundary crossed | yes, first at observation 1,671 |
| Observations analysed | 6,000 |
| Effect (Variant B minus Control) | 3.84 pp |
| Always-valid 95% confidence sequence | 1.34 pp to 6.34 pp |
| Always-valid p-value | < 0.001 |
| Fixed-horizon p-value at this sample size | < 0.001 |
| Smallest effect this test was set up to find | 3.18 pp |
| Control arm | Control (n = 2,975) |
| Treatment arm | Variant B (n = 3,025) |
The short answer
Stop—a winner has separated. The always-valid confidence sequence (1.34 pp to 6.34 pp) excludes zero, and the boundary was crossed at observation 1,671.
The detail
The effect estimate is 3.84 pp (Variant B minus Control). The always-valid 95% confidence sequence is 1.34 pp to 6.34 pp; the always-valid p-value is < 0.001 and the fixed-horizon p-value at this sample size is also < 0.001. The smallest effect this test was set up to find is 3.18 pp (tau). Control has n = 2,975 and Variant B has n = 3,025. The decision rule has three exits and one refusal: stop for a winner when the confidence sequence excludes zero (true here), stop for futility when it lies entirely within 3.18 pp (false), keep running when neither is true but both are reachable (not applicable), or refuse when the interval is wider than 3.18 pp and no boundary is reachable (not applicable).
What this can't tell you
Nothing in this table settles whether assignment was random or whether each row is a distinct person.
What Peeking Would Have Cost You
The fixed-horizon test and the always-valid boundary, counted over the same analysis points.
| Check | Value |
|---|---|
| Analysis points examined (one after every observation) | 5,973 |
| Points where the fixed-horizon p-value was below 0.05 | 5,219 |
| Separate excursions below 0.05 by the fixed-horizon p-value | 12 |
| Points where the always-valid boundary was crossed | 4,290 |
| Points where the fixed-horizon test said stop but the always-valid boundary did not | 929 |
| First point the fixed-horizon test would have stopped | observation 701 |
| First point the always-valid boundary was crossed | observation 1,671 |
The short answer
Over 5,973 analysis points, the fixed-horizon test would have stopped at 929 moments where the always-valid boundary had not yet been crossed. The fixed-horizon p-value fell below 0.05 at 5,219 points in 12 separate excursions; the always-valid boundary crossed at 4,290.
The detail
The fixed-horizon test would have first stopped at observation 701; the always-valid boundary was first crossed at observation 1,671—a gap of 970 observations. This assumes the most aggressive schedule possible: a look after every observation. Both tests were evaluated at all 5,973 analysis points on the same cumulative data. The 929 points where only the fixed-horizon test fired represent the concrete cost of watching a dashboard that is not always-valid.
What this can't tell you
A weekly or monthly review cadence would see fewer opportunities to be misled by the fixed-horizon test, but the gap between the two stopping points (observation 701 vs 1,671) would still exist.
The Price of Being Allowed to Peek
How much extra data the always-valid guarantee costs on this experiment.
| Measure | Value |
|---|---|
| Always-valid half-width at this sample size | 2.50 pp |
| Fixed-horizon half-width at this sample size | 1.60 pp |
| How much wider the always-valid interval is | 1.56 times |
| Observations needed for the always-valid interval to be as tight as the fixed-horizon one is now | 15,512 |
| That as a multiple of the current sample size | 2.59 times |
The short answer
The always-valid interval is 1.56 times wider than the fixed-horizon interval at this sample size. Reaching today's fixed-horizon precision under the always-valid boundary would require about 15,512 observations—2.59 times what you have.
The detail
The always-valid half-width is 2.50 pp; the fixed-horizon half-width is 1.60 pp. The ratio is 1.56. If you genuinely can commit to a single analysis at a sample size fixed in advance, the fixed-horizon test is the more efficient choice. This tool is for the common case where that commitment is not realistic and you need the right to look whenever you like.
What this can't tell you
These calculations assume the same power and type-I error rate for both methods. The trade-off is inherent to always-valid inference and cannot be eliminated.
Does the Verdict Survive a Different Alternative?
The same monitoring path re-run against three pre-specified alternatives.
| Alternative | Alternative Size | Half Width | Decision | Ever Crossed |
|---|---|---|---|---|
| 0.05 x the outcome's standard deviation | 1.59 | 2.53 | stop for a winner | yes |
| 0.10 x the outcome's standard deviation | 3.18 | 2.50 | stop for a winner | yes |
| 0.20 x the outcome's standard deviation | 6.35 | 2.62 | stop for a winner | yes |
The short answer
The boundary's tuning constant (tau) was set to 0.10 times the outcome's standard deviation. Re-running the entire path at half and double that value (0.05 and 0.20 times the standard deviation) yields the same verdict—stop for a winner—under all three. The conclusion does not hinge on the choice of tau.
The detail
Tau defines the effect size the test is aimed at, set here to 0.10 × outcome standard deviation (3.18 pp). At tau = 0.05 × SD (1.59 pp), the half-width is 2.53 pp and the verdict is stop for a winner. At tau = 0.10 × SD (3.18 pp), the half-width is 2.50 pp and the verdict is stop for a winner. At tau = 0.20 × SD (6.35 pp), the half-width is 2.62 pp and the verdict is stop for a winner. All three alternatives crossed the boundary. Tau should strictly be fixed before the experiment; it is read from the outcome's spread here rather than from the observed gap between arms, so it does not borrow the answer.
What this can't tell you
Tau sensitivity does not test robustness to violations of the normal-mixture assumption or to departures from the Welch t-test model. The choice of tau before the experiment is a design decision, not a data-driven one.
Methods and Disclosure
Every formula, constant and assumption behind the decision.
| Item | Detail |
|---|---|
| Test | Mixture sequential probability ratio test (mSPRT) on the difference in means between the two arms, implemented directly from the normal mixture — no group-sequential package is used. |
| Boundary | Robbins' normal-mixture boundary. The mixture likelihood ratio is Lambda = sqrt(V/(V+tau^2)) exp(delta^2 tau^2 / (2 V (V+tau^2))) with V the sampling variance of the effect; Ville's inequality bounds the chance that Lambda ever reaches 1/alpha at alpha, which is what makes continuous monitoring safe. |
| Confidence sequence | Inverting the boundary gives effect +/- h with h = sqrt(2 V (V+tau^2)/tau^2 log(sqrt((V+tau^2)/V)/alpha)); this interval covers the true effect at every sample size at once, not just at one pre-chosen sample size. |
| Pre-specified alternative (tau) | 3.18 pp — set at 0.10 times the outcome's standard deviation (31.76). It is the effect size the boundary is tuned to find, and it is read from how much the outcome varies, not from the difference between the arms. |
| Significance level | alpha = 0.05, applied to the whole monitoring path rather than to a single look. |
| Analysis schedule | The test is evaluated after every single observation from the point where both arms reach 10 rows — the most aggressive peeking possible — giving 5,973 analysis points. |
| Fixed-horizon comparison | At each of those points a Welch two-sample t-test is computed on the same cumulative data, purely so the two boundaries can be compared on this experiment's own numbers. |
| Effect definition | Mean of Variant B minus mean of Control, both being conversion rates, reported in percentage points. |
| Control arm chosen by | Control — its name identifies it as the control. |
| Arrival order read as | calendar dates in the column 'Signup Date'; 5,850 row(s) share an arrival value with an earlier row and keep their original file order. |
| What this cannot do | The guarantee is about the stopping rule, not about everything else. It holds for this one pre-specified comparison of these two arms on this one outcome. It does not cover choosing the metric, the segment or the arm after seeing the results, it cannot detect repeat visits by the same person, and it does not know whether the assignment was actually random. |
The short answer
The test is a mixture sequential probability ratio test (mSPRT) using Robbins' normal-mixture boundary. Ville's inequality guarantees that the chance the mixture likelihood ratio ever reaches 1/0.05 is at most 0.05, no matter how many times you look. The confidence sequence inverts this boundary and covers the true effect at every sample size simultaneously.
The detail
The mixture likelihood ratio is Lambda = sqrt(V/(V+tau²)) exp(δ²τ²/(2V(V+τ²))) with V the sampling variance of the effect and tau = 3.18 pp. The confidence sequence is effect ± h where h = sqrt(2V(V+τ²)/τ² log(sqrt((V+τ²)/V)/α)) with α = 0.05. The analysis evaluates the test after every observation from the point both arms reach 10 rows, giving 5,973 analysis points. A Welch two-sample t-test is computed at each point for comparison. The control arm was identified by name; arrival order was read as calendar dates from Signup Date, with 5,850 ties resolved by preserving file order.
What this can't tell you
The normal mixture is asymptotic and improves with sample size. Variance is estimated from the same data being tested, so coverage at very small samples is optimistic. The guarantee covers the stopping rule and this one pre-specified comparison only—it does not protect against choosing the metric, segment, or arm after seeing results, detecting repeat visits, or verifying random assignment.
Methodology
Statistical methodology and diagnostics for Sequential A/B Test — Peek Safely
Statistical Method
Standard-library analysis: can I stop this experiment early? Upload an ordered experiment log — the variant each row saw, what happened, and the date or sequence the observation arrived — and get inference that stays valid under continuous monitoring. The analysis re-tests after every single observation using a mixture sequential probability ratio test, and returns the always-valid confidence sequence over time, the current decision (stop for a winner, stop for futility, keep running, or not enough data for any boundary to be reachable), the boundary that was crossed and when, and the cumulative effect charted with both boundaries drawn over it. The naive fixed-horizon p-value is computed alongside at the same points, so the report can show — from your own data, not from an assertion — how many times it would have crossed 0.05 while the always-valid boundary did not. The report also states what the guarantee costs in power, in this experiment's own numbers.
- Each row is one independent observation assigned to one of exactly two arms
- Assignment to the arms was random, so the arms are comparable on everything except the variant
- The arrival column reflects the order the observations actually arrived — sequential inference is entirely about that order
- The outcome is numeric or a 0/1 flag, and the two arms' means are the quantity of interest
- The comparison — these two arms, this outcome — was decided before the data was inspected
- Always-valid inference buys the right to peek by widening the interval, so it needs more data than a fixed-horizon test to reach the same conclusion; the report states that cost in your own numbers rather than presenting peeking as free
- The guarantee covers the stopping rule only. It does not protect against choosing the metric, the segment or the arm after seeing the data, and no single file can reveal that this happened
- The boundary has one tuning constant, the effect size it is aimed at. It should strictly be fixed before the experiment; here it is read from the outcome's spread, and the report re-runs the verdict at half and double that value
- The mixture boundary uses a normal approximation with the variance estimated from the same data, so at small samples the stated coverage is optimistic
Analysis Code
Complete R source code for this analysis
Sequential A/B Test — Peek Safely
Reads an ordered experiment log (who saw which variant, what happened, and when the observation arrived) and answers the question a running experiment actually raises: given everything so far, can I stop?
Why This Method?
A fixed-horizon p-value is only valid if you look once, at a sample size fixed before the data existed. Look after every visitor and the chance of seeing p < 0.05 somewhere along the way is far above 5%. A mixture sequential probability ratio test (mSPRT) replaces the fixed critical value with a boundary that is valid at every sample size simultaneously, so monitoring the experiment continuously does not inflate the false-positive rate. The price is power, and this module states that price in the user's own numbers rather than selling peeking as free.
What This Analysis Covers
- The always-valid confidence sequence for the treatment effect over time
- The current decision: stop for a winner, stop for futility, keep running,
or not enough data for any boundary to be reachable
- The naive fixed-horizon p-value computed alongside, and how many times it
would have crossed 0.05 while the always-valid boundary did not
- The cost of the guarantee: how much wider the always-valid interval is and
how many observations it needs to match today's fixed-horizon precision
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {variant, outcome, arrival}. The mSPRT, the confidence sequence and the boundary crossing are implemented directly from the normal mixture — no group-sequential package is used.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
Step 1: Confirm the mapping and humanize the user's column names
variant_h <- humanize_semantic("variant", col_map)
outcome_h <- humanize_semantic("outcome", col_map)
arrival_h <- humanize_semantic("arrival", col_map)
needed <- c("variant", "outcome", "arrival")
missing_keys <- needed[!(needed %in% names(df))]
if (length(missing_keys) > 0) {
labels <- c(variant = variant_h, outcome = outcome_h, arrival = arrival_h)
stop(sprintf(
paste0("Sequential testing needs three columns mapped: the variant each ",
"row saw, the outcome, and the order or date the observation ",
"arrived. Missing: %s."),
paste(labels[missing_keys], collapse = ", ")))
}Step 2: Read the arrival column as dates, or as a numeric order
arr_ch <- trimws(as.character(df$arrival))
arr_ch[arr_ch == "NA"] <- NA_character_
n_nonblank_arr <- sum(!is.na(arr_ch) & arr_ch != "")
parse_dates_vec <- function(ch) {
ok <- function(dd) sum(!is.na(dd)) >= 0.95 * sum(!is.na(ch) & trimws(ch) != "")
d <- suppressWarnings(as.Date(ch, format = "%Y-%m-%d"))
if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::ymd(ch, quiet = TRUE)))
if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::mdy(ch, quiet = TRUE)))
if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::dmy(ch, quiet = TRUE)))
if (ok(d)) d else NULL
}
arrival_kind <- ""
ord_key <- NULL
if (n_nonblank_arr > 0) {
dts <- parse_dates_vec(arr_ch)
if (!is.null(dts)) {
ord_key <- as.numeric(dts)
arrival_kind <- "calendar dates"
} else {
nv <- suppressWarnings(as.numeric(arr_ch))
if (sum(!is.na(nv)) >= 0.95 * n_nonblank_arr) {
ord_key <- nv
arrival_kind <- "a numeric arrival order"
}
}
}
if (is.null(ord_key)) {
stop(sprintf(
paste0("The arrival column '%s' could not be read either as dates or as ",
"a numeric order, so the observations cannot be put in the order ",
"they arrived. Sequential testing is entirely about that order — ",
"map a date, a timestamp, or a running sequence number."),
arrival_h))
}Step 3: Read the outcome as a number (95% rule, then yes/no words)
y_raw <- df$outcome
outcome_read <- ""
if (is.numeric(y_raw)) {
y <- as.numeric(y_raw)
outcome_read <- "already numeric"
} else {
ch <- trimws(as.character(y_raw))
ch[ch == "NA"] <- NA_character_
nonblank <- !is.na(ch) & ch != ""
conv <- suppressWarnings(as.numeric(ch))
if (sum(nonblank) > 0 && sum(!is.na(conv)) >= 0.95 * sum(nonblank)) {
y <- conv
outcome_read <- "read as numbers"
} else {
lut <- c("yes" = 1, "no" = 0, "true" = 1, "false" = 0, "y" = 1, "n" = 0,
"t" = 1, "f" = 0, "converted" = 1, "not converted" = 0,
"1" = 1, "0" = 0, "success" = 1, "failure" = 0)
conv2 <- unname(lut[tolower(ch)])
if (sum(nonblank) > 0 && sum(!is.na(conv2)) >= 0.95 * sum(nonblank)) {
y <- conv2
outcome_read <- "read as yes/no labels turned into 1 and 0"
} else {
stop(sprintf(
paste0("The outcome column '%s' is not numeric and is not a yes/no ",
"flag — fewer than 95%% of its values could be turned into ",
"numbers. Map a numeric outcome(revenue, time, a 0/1 ",
"conversion flag) so the two arms can be compared."),
outcome_h))
}
}
}Step 4: Keep complete rows only
g_ch <- trimws(as.character(df$variant))
g_ch[g_ch == "" | g_ch == "NA"] <- NA_character_
keep <- !is.na(g_ch) & !is.na(y) & !is.na(ord_key)
n_dropped <- sum(!keep)
g_k <- g_ch[keep]; y_k <- y[keep]; o_k <- ord_key[keep]
final_rows <- length(y_k)
rows_removed <- initial_rows - final_rows
if (final_rows < 30) {
stop(sprintf(
paste0("Only %d complete row(s) remained after dropping rows missing ",
"%s, %s or %s. Sequential monitoring needs at least 30 ",
"observations before any boundary means anything."),
final_rows, variant_h, outcome_h, arrival_h))
}Step 5: Resolve the two arms — never chosen after seeing the outcome
tab <- sort(table(g_k), decreasing = TRUE)
tiny_arms <- names(tab)[tab < MIN_ARM_KEEP]
if (length(tiny_arms) > 0) {
keep2 <- !(g_k %in% tiny_arms)
g_k <- g_k[keep2]; y_k <- y_k[keep2]; o_k <- o_k[keep2]
final_rows <- length(y_k)
rows_removed <- initial_rows - final_rows
tab <- sort(table(g_k), decreasing = TRUE)
}
arms <- names(tab)
if (length(arms) < 2) {
stop(sprintf(
paste0("The variant column '%s' holds only %d group%s with at least %d ",
"observation%s(%s). A sequential A/B test compares exactly two ",
"arms."),
variant_h, length(arms), if (length(arms) == 1) "" else "s",
MIN_ARM_KEEP, "s", paste(arms, collapse = ", ")))
}
if (length(arms) > 2) {
stop(sprintf(
paste0("The variant column '%s' holds %d arms (%s). This analysis will ",
"not pick two of them for you: the always-valid guarantee covers ",
"the comparison you specified before looking at the data, and ",
"choosing which arms to compare after seeing the results is ",
"exactly what it does not protect against. Filter the file to the ",
"two arms you want to compare and re-run."),
variant_h, length(arms), paste(arms, collapse = ", ")))
}Control is chosen by NAME, not by outcome — a fixed, stated rule.
ctrl_pat <- "^(control|baseline|a|original|holdout|current|old)$"
ctrl_hit <- arms[grepl(ctrl_pat, tolower(trimws(arms)))]
control_arm <- if (length(ctrl_hit) >= 1) ctrl_hit[1] else sort(arms)[1]
control_rule <- if (length(ctrl_hit) >= 1)
"its name identifies it as the control" else
"no arm was named as a control, so the alphabetically first label was used"
treat_arm <- setdiff(arms, control_arm)[1]Step 6: Order by arrival (ties keep their original file order)
ordv <- order(o_k, seq_along(o_k))
g_s <- g_k[ordv]; y_s <- y_k[ordv]; o_s <- o_k[ordv]
n_tied <- sum(duplicated(o_s))
n_total <- length(y_s)
sd_all <- stats::sd(y_s)
if (!is.finite(sd_all) || sd_all <= 0) {
stop(sprintf(
paste0("Every retained row has the same value in the outcome column ",
"'%s', so there is no difference between the arms to test and no ",
"boundary to cross."),
outcome_h))
}
tau <- TAU_MULT * sd_all
t2 <- tau^2
is_binary <- all(y_s %in% c(0, 1))
unit <- if (is_binary) "percentage points" else "units of the outcome"fmt_eff already carries "pp" for a binary outcome, so the spelled-out unit is appended only when it would not be a duplicate.
unit_suffix <- if (is_binary) "" else paste0(" ", unit)
to_disp <- function(x) if (is_binary) 100 * x else x
disp_digits <- if (is_binary) 2 else 4
fmt_eff <- function(x) trimws(paste0(.fmt_n(to_disp(x), disp_digits),
if (is_binary) " pp" else ""))Step 7: Cumulative statistics after every single observation
is_c <- as.integer(g_s == control_arm)
is_t <- 1L - is_c
nA <- cumsum(is_c); nB <- cumsum(is_t)
sA <- cumsum(y_s * is_c); sB <- cumsum(y_s * is_t)
qA <- cumsum(y_s^2 * is_c); qB <- cumsum(y_s^2 * is_t)
mA <- ifelse(nA > 0, sA / nA, NA_real_)
mB <- ifelse(nB > 0, sB / nB, NA_real_)
vA <- ifelse(nA > 1, (qA - nA * mA^2) / (nA - 1), NA_real_)
vB <- ifelse(nB > 1, (qB - nB * mB^2) / (nB - 1), NA_real_)
vA[!is.na(vA) & vA < 0] <- 0
vB[!is.na(vB) & vB < 0] <- 0
V <- vA / nA + vB / nB
delta <- mB - mA
valid <- nA >= MIN_PER_ARM & nB >= MIN_PER_ARM &
is.finite(V) & V > 0 & is.finite(delta)
idx <- which(valid)
if (length(idx) == 0) {
stop(sprintf(
paste0("Neither arm of '%s' reached %d observations with any variation ",
"in '%s', so the experiment cannot be monitored yet. %s has %d ",
"row(s) and %s has %d."),
variant_h, MIN_PER_ARM, outcome_h,
control_arm, sum(is_c), treat_arm, sum(is_t)))
}
h_path <- .av_halfwidth(V[idx], t2, ALPHA)
d_path <- delta[idx]
lam_path <- .av_lambda(d_path, V[idx], t2)
se_path <- sqrt(V[idx])
dfw <- (V[idx])^2 / ((vA[idx] / nA[idx])^2 / pmax(nA[idx] - 1, 1) +
(vB[idx] / nB[idx])^2 / pmax(nB[idx] - 1, 1))
dfw[!is.finite(dfw) | dfw < 1] <- 1
t_path <- d_path / se_path
p_fh_path <- 2 * stats::pt(-abs(t_path), df = dfw)
fh_hw_path <- stats::qt(1 - ALPHA / 2, df = dfw) * se_path
p_av_path <- cummin(pmin(1, 1 / lam_path))
crossed_av <- is.finite(h_path) & abs(d_path) > h_path
naive_sig <- is.finite(p_fh_path) & p_fh_path < ALPHA
n_points <- length(idx)
av_points <- sum(crossed_av)
naive_points <- sum(naive_sig)
naive_excursions <- sum(naive_sig & !c(FALSE, head(naive_sig, -1)))
naive_only <- sum(naive_sig & !crossed_av)
cross_idx <- if (av_points > 0) which(crossed_av)[1] else NA_integer_
cross_n <- if (!is.na(cross_idx)) idx[cross_idx] else NA_integer_
naive_first_idx <- if (naive_points > 0) which(naive_sig)[1] else NA_integer_
naive_first_n <- if (!is.na(naive_first_idx)) idx[naive_first_idx] else NA_integer_Step 8: The current decision
last <- length(idx)
effect <- d_path[last]
h_final <- h_path[last]
se_final <- se_path[last]
fh_hw_final <- fh_hw_path[last]
p_av <- p_av_path[last]
p_fh <- p_fh_path[last]
av_low <- effect - h_final
av_high <- effect + h_final
fh_low <- effect - fh_hw_final
fh_high <- effect + fh_hw_final
decide <- function(eff, hw) {
if (!is.finite(hw)) return("insufficient")
if (abs(eff) > hw) return("winner")
if (hw > tau) return("insufficient")
if (abs(eff) + hw <= tau) return("futility")
"continue"
}
decision <- decide(effect, h_final)
decision_label <- switch(
decision,
winner = "Stop — a winner has separated",
futility = "Stop for futility — no effect worth acting on remains possible",
continue = "Keep running — no boundary has been crossed yet",
insufficient = "Not enough data yet — no boundary is reachable at this sample size")Step 9: The price of the guarantee, in this experiment's own numbers
frac_c <- sum(is_c) / n_total
frac_c <- min(max(frac_c, 1e-6), 1 - 1e-6)
vA_f <- vA[idx][last]; vB_f <- vB[idx][last]
h_at_n <- function(n) {
Vn <- vA_f / (n * frac_c) + vB_f / (n * (1 - frac_c))
.av_halfwidth(Vn, t2, ALPHA)
}
grid <- unique(round(exp(seq(log(n_total), log(n_total * 5000), length.out = 6000))))
grid <- grid[is.finite(grid) & grid >= n_total]
hg <- h_at_n(grid)
first_at_or_below <- function(target) {
okv <- which(is.finite(hg) & hg <= target)
if (length(okv) == 0) NA_real_ else grid[okv[1]]
}
n_equiv <- first_at_or_below(fh_hw_final)
n_needed_futility <- if (decision == "insufficient")
first_at_or_below(tau) else NA_real_
width_ratio <- if (is.finite(fh_hw_final) && fh_hw_final > 0)
h_final / fh_hw_final else NA_real_Step 10: Sensitivity to the pre-specified alternative
tau_mults <- c(0.05, 0.10, 0.20)
tau_rows <- lapply(tau_mults, function(mm) {
tt <- mm * sd_all
hh <- .av_halfwidth(V[idx], tt^2, ALPHA)
dd <- decide(effect, hh[last])
data.frame(
alternative = paste0(.fmt_n(mm, 2), " x the outcome's standard deviation"),
alternative_size = .fmt_n(to_disp(tt), disp_digits),
half_width = .fmt_n(to_disp(hh[last]), disp_digits),
decision = switch(dd, winner = "stop for a winner",
futility = "stop for futility",
continue = "keep running",
insufficient = "not enough data"),
ever_crossed = if (sum(is.finite(hh) & abs(d_path) > hh) > 0) "yes" else "no",
stringsAsFactors = FALSE)
})
tau_df <- do.call(rbind, tau_rows)
tau_agreement <- length(unique(tau_df$decision)) == 1Step 12: One computed paragraph for the JSON answer
contrast_txt <- if (naive_only > 0) {
paste0("Along the way the fixed-horizon p-value fell below 0.05 at ",
format(naive_only, big.mark = ","), " analysis point(s) where the ",
"always-valid boundary was not crossed, in ",
format(naive_excursions, big.mark = ","),
" separate excursion(s) — that gap is the false-positive risk ",
"peeking at a fixed-horizon test would have exposed you to here.")
} else if (naive_points > 0) {
paste0("The fixed-horizon p-value was below 0.05 at ",
format(naive_points, big.mark = ","),
" analysis point(s), all of them points where the always-valid ",
"boundary had also been crossed, so on this data the two tests ",
"never disagreed.")
} else {
paste0("The fixed-horizon p-value never fell below 0.05 at any of the ",
format(n_points, big.mark = ","),
" analysis points, so on this data peeking would not have misled ",
"you — which is luck, not protection.")
}
json_output <- list(
answer = paste0(
decision_label, ". After ", format(n_total, big.mark = ","),
" observations the effect of ", treat_arm, " over ", control_arm,
" is ", fmt_eff(effect), " with an always-valid 95% confidence sequence of ",
fmt_eff(av_low), " to ", fmt_eff(av_high), unit_suffix,
" (always-valid p-value ", .fmt_p(p_av), "). ", contrast_txt,
" The guarantee costs power: the always-valid interval is ",
if (is.finite(width_ratio)) paste0(.fmt_n(width_ratio, 2), " times") else "wider than",
" the fixed-horizon interval at this sample size."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "boundary_path", "pvalue_paths",
"decision", "peek_comparison", "power_cost", "tau_sensitivity", "methods"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed, n_dropped = n_dropped, n_tied = n_tied,
variant_h = variant_h, outcome_h = outcome_h, arrival_h = arrival_h,
arrival_kind = arrival_kind, outcome_read = outcome_read,
tiny_arms = tiny_arms,
control_arm = control_arm, treat_arm = treat_arm, control_rule = control_rule,
n_control = sum(is_c), n_treat = sum(is_t), n_total = n_total,
is_binary = is_binary, unit = unit, unit_suffix = unit_suffix,
to_disp = to_disp, fmt_eff = fmt_eff,
sd_all = sd_all, tau = tau, alpha = ALPHA, tau_mult = TAU_MULT,
min_per_arm = MIN_PER_ARM,
decision = decision, decision_label = decision_label,
effect = effect, av_low = av_low, av_high = av_high,
fh_low = fh_low, fh_high = fh_high,
h_final = h_final, fh_hw_final = fh_hw_final, se_final = se_final,
p_av = p_av, p_fh = p_fh,
cross_idx = cross_idx, cross_n = cross_n,
naive_first_idx = naive_first_idx, naive_first_n = naive_first_n,
n_points = n_points, naive_points = naive_points,
naive_excursions = naive_excursions, av_points = av_points,
naive_only = naive_only,
width_ratio = width_ratio, n_equiv = n_equiv,
n_needed_futility = n_needed_futility, grid_max = max(grid),
tau_df = tau_df, tau_agreement = tau_agreement,
boundary_path_df = boundary_path_df, pvalue_paths_df = pvalue_paths_df,
decision_df = decision_df, peek_df = peek_df, power_df = power_df,
methods_df = methods_df,
metrics = metrics, json_output = json_output
)
}