Your email redesign test is running. Day 3: p = 0.08. Day 5: p = 0.04. Significant! You stop the test and ship the new design. Three weeks later, conversion rate is down 6%. What happened? You peeked. Every time you check an ongoing experiment and make a decision based on the current p-value, you inflate your false positive rate. That "p < 0.05" you saw on Day 5 wasn't valid—your actual alpha was closer to 0.30. Sequential A/B testing solves this: it lets you monitor experiments continuously using always-valid p-values, stopping early when effects are clear without sacrificing statistical rigor. Here's the methodology that cuts sample size costs by 30% while maintaining proper error control.

Before we draw conclusions from any test, let's check the experimental design. Did you randomize? What were the stopping criteria? Standard A/B tests require you to fix sample size in advance and wait until the end. Sequential tests allow continuous monitoring with proper boundary adjustments. The difference isn't just convenience—it's cost. Stopping an experiment early when you have strong evidence saves time, traffic, and opportunity cost.

The problem with traditional A/B testing is brutal: you either peek (inflating false positives) or you don't (wasting resources). Sequential testing gives you a third option: peek all you want, but use corrected boundaries that maintain 5% Type I error across all possible stopping points. Let's build the framework from first principles.

Why Standard A/B Tests Penalize Business Efficiency

You design a standard fixed-sample A/B test. Power analysis says you need 10,000 users per variant to detect a 5% lift in conversion rate. You launch the experiment on Monday. By Wednesday, variant B is converting at 12.8% vs. control at 11.1%—a massive 15% relative lift. The p-value is 0.003. This looks decisive.

But you planned for 10,000 per group, and you've only collected 2,000. Do you stop early and capture gains? Or do you wait another three weeks for the full sample?

The Cost of Waiting When You Already Have Your Answer

If you wait for the full 10,000 per variant, you're exposing 8,000 additional users to the inferior control experience. At 11.1% conversion, that's 888 conversions from control vs. 1,024 from the variant if you'd switched everyone to B immediately. You left 136 conversions on the table.

In e-commerce, that might be $13,600 in lost revenue. In SaaS, that's 136 trial signups you didn't get. The opportunity cost of waiting for statistical perfection is real money.

The Cost of Peeking When You Don't Control Alpha

Suppose instead you peeked at results daily and stopped the first time p < 0.05. You would have stopped on Day 2 when the p-value first dipped below 0.05 due to random noise. But that early "significance" was a false positive. After three more weeks, the true conversion rates would have converged to 11.2% for both variants—no real difference.

Now you've shipped a change that doesn't work, your team spent two sprints implementing it, and you've lost the opportunity to test something that would have actually mattered.

The Alpha Inflation Problem Is Not Hypothetical

Armitage et al. (1969) proved that checking a test 5 times at evenly spaced intervals inflates α from 0.05 to 0.14. Check it 10 times? Alpha reaches 0.19. Check it continuously? Alpha approaches 0.30. This isn't a theoretical curiosity—it's why your A/B testing program has a 30% false positive rate if you're peeking without correction.

What Sequential Testing Buys You: ROI in Real Terms

Sequential testing methods provide corrected stopping boundaries that let you check results as often as you want while maintaining the chosen alpha level (typically 0.05). The payoff comes in three forms.

1. Early Stopping for Large Effects = Direct Cost Savings

When the true effect is large (10%+ lift), sequential tests typically reach significance with 30-50% of the planned maximum sample size. For a test planned at n = 10,000 per group, you might stop at n = 5,000 per group.

Cost savings: 10,000 users who could be in your next experiment instead of confirming what you already know. At one experiment per month with standard methods, switching to sequential testing might let you run 1.3 experiments per month—30% more learning from the same traffic.

2. Safety Monitoring for Harmful Variants

Suppose your new checkout flow has a bug that crashes on iPhone 12 devices. After 500 users, conversion rate on variant B is 3.2% vs. control at 11.1%. Sequential testing lets you stop for futility or harm without inflating false positives.

The cost of waiting for 10,000 users when you have clear evidence of harm by user 500 is 9,500 users with a broken experience. Sequential boundaries let you pull the plug immediately when evidence is overwhelming—in either direction.

3. Reduced Opportunity Cost from Faster Decisions

Every week an experiment runs, you're not running a different experiment. If you have 50 ideas and test them sequentially at 3 weeks each, you'll test 17 ideas per year. If sequential methods cut average test duration to 2 weeks, you'll test 26 ideas per year—50% more.

Even if only 20% of tests produce wins, going from 17 tests to 26 tests means going from 3.4 wins to 5.2 wins per year. That's 50% more conversion rate improvements, revenue lifts, and engagement gains from the same traffic and team.

ROI Case Study: E-Commerce A/B Testing Program

A mid-size e-commerce company (500K monthly users) switched from fixed-sample to group sequential testing. Before: 2.8 experiments per month, average duration 18 days. After: 3.9 experiments per month, average duration 13 days. Same traffic, 39% more experiments. Estimated value: 6 additional winning tests per year at ~$40K revenue impact each = $240K incremental annual revenue from methodology change alone.

Group Sequential Tests: The Practical Middle Ground

Group sequential designs let you check results at pre-planned intervals (e.g., after 25%, 50%, 75%, 100% of maximum sample size) using adjusted significance boundaries. This is the workhorse method for business A/B testing.

How Pocock Boundaries Work

Pocock (1977) proposed using the same significance threshold at every interim look. For α = 0.05 and 4 equally spaced looks, the Pocock boundary is approximately p < 0.0182 at each look.

Here's the logic: if you check 4 times using p < 0.05 each time, your overall false positive rate inflates to ~0.14. By requiring p < 0.0182 at each look, you keep the overall alpha at exactly 0.05 across all possible stopping points.

Example setup:

  • Maximum sample size: 10,000 per variant (from standard power analysis)
  • Interim looks: After 2,500, 5,000, 7,500, and 10,000 users per variant
  • Significance boundary: p < 0.0182 at each look
  • Stopping rule: Stop if p < 0.0182 at any look, or at maximum n regardless of result

If the true effect is large, you'll likely hit p < 0.0182 by the first or second look and stop early. If the effect is small or null, you'll run to maximum n, just like a standard test.

How O'Brien-Fleming Boundaries Work

O'Brien-Fleming (1979) boundaries are more conservative early, relaxing as you approach maximum sample size. For 4 looks at α = 0.05:

Look Sample Fraction Boundary (p-value)
1 25% 0.0001
2 50% 0.0038
3 75% 0.0190
4 100% 0.0417

Early looks require overwhelming evidence (p < 0.0001). By the final look, the boundary is nearly identical to the standard p < 0.05. O'Brien-Fleming boundaries preserve power for small effects better than Pocock but make early stopping harder.

When to use which? Pocock if you value early stopping for cost savings. O'Brien-Fleming if you want to preserve power and only stop for truly massive effects.

Worked Example: Email Subject Line Test with Group Sequential Design

You're testing a new email subject line. Control: "Your weekly digest." Variant: "5 insights you missed this week."

Study design:

  • Outcome: Email open rate
  • Baseline: 22% open rate
  • Minimum detectable effect: 2 percentage points (9% relative lift)
  • Power: 80%, α = 0.05 (two-tailed)
  • Required sample size (fixed test): 9,240 per group
  • Group sequential design: 3 looks at 33%, 67%, 100% using O'Brien-Fleming boundaries

Boundaries for 3 looks:

  • Look 1 (3,080 per group): p < 0.0005
  • Look 2 (6,160 per group): p < 0.0141
  • Look 3 (9,240 per group): p < 0.0451

Actual results:

Look 1 (3,080 emails sent per variant):
Control: 678 opens / 3,080 = 22.0%
Variant: 756 opens / 3,080 = 24.5%
Difference: +2.5 percentage points
Z-score: 2.67, p = 0.0076
Decision: p = 0.0076 > 0.0005 (boundary), continue

Look 2 (6,160 emails sent per variant):
Control: 1,355 opens / 6,160 = 22.0%
Variant: 1,524 opens / 6,160 = 24.7%
Difference: +2.7 percentage points
Z-score: 3.78, p = 0.0002
Decision: p = 0.0002 < 0.0141 (boundary), STOP for efficacy

Conclusion: Variant outperformed control with 24.7% open rate vs. 22.0% (sequential p = 0.0002). Stopped after 6,160 per group instead of planned 9,240, saving 33% of sample size. Ship variant B.

Cost savings: 6,160 additional email recipients available for next experiment instead of confirming an already-clear result. That's 6,160 opportunities to test your next idea this month instead of next month.

Run Your Own Sequential A/B Test

Upload a CSV with your A/B test data. MCP Analytics calculates group sequential boundaries (Pocock or O'Brien-Fleming), applies proper alpha spending, and tells you whether to stop or continue at each interim look.

Try Free Sequential A/B Testing Tool

Fully Sequential Tests: mSPRT and Continuous Monitoring

Group sequential tests check at pre-planned intervals. Fully sequential tests allow checking after every observation. The modified Sequential Probability Ratio Test (mSPRT) is the optimal approach for continuous monitoring.

The Sequential Probability Ratio Test (Wald, 1945)

Wald's SPRT is elegant: after each observation, calculate the likelihood ratio comparing two hypotheses (H₀: no effect, H₁: effect of size δ). If the likelihood ratio exceeds an upper threshold, stop and reject H₀. If it falls below a lower threshold, stop and accept H₀. Otherwise, continue sampling.

The thresholds are set to control Type I and Type II error rates exactly. For α = 0.05 and β = 0.20 (80% power):

  • Upper boundary (reject H₀): Likelihood ratio > 19 (approximately)
  • Lower boundary (accept H₀): Likelihood ratio < 0.05 (approximately)

SPRT is provably optimal: among all sequential tests with the same error rates, SPRT requires the smallest expected sample size.

Modified SPRT for Composite Hypotheses (Johari et al., 2017)

Traditional SPRT requires you to specify the exact alternative effect size δ. In practice, you don't know the true effect. Modified SPRT (mSPRT) uses a mixture of likelihood ratios across a range of plausible effect sizes, making it robust to misspecification.

The procedure:

  1. After each observation, update the mixture likelihood ratio
  2. If mixture LR > threshold (set to control α), reject H₀ and stop
  3. If you reach maximum planned n without rejection, accept H₀

For α = 0.05, the typical threshold is mixture LR > 20. This threshold produces always-valid p-values: you can check after every observation, stop whenever you want, and your false positive rate remains exactly 5%.

When Continuous Monitoring Matters

Fully sequential tests shine in high-traffic environments where observations arrive constantly. If you're testing on 100,000 daily users, waiting for "interim looks" feels arbitrary. Why wait until exactly 25,000 users when you have real-time data?

Use cases for mSPRT:

  • High-traffic web applications: Testing button colors, headlines, or CTAs where thousands of users arrive hourly
  • Ad campaign optimization: Continuously monitoring ad creative performance and cutting underperformers immediately
  • Safety monitoring: Detecting harmful effects (crashes, errors, checkout failures) as soon as statistically defensible

The cost of not using continuous monitoring when you have continuous data is delayed decisions. Every hour you wait to stop a clearly winning or losing variant is wasted traffic.

Implementation: Always-Valid P-Values

The key insight from mSPRT is that you can construct p-values that remain valid no matter when you check them. These "always-valid p-values" never require correction for multiple testing because they explicitly account for continuous monitoring.

Standard p-values are only valid at the pre-planned sample size. Check them early, and they're biased downward (too many false positives). Always-valid p-values are calibrated to the sequential context: they're conservative early (requiring overwhelming evidence) and converge to standard p-values at maximum n.

The Math Behind Always-Valid Inference

Robbins (1970) and Lai (1976) developed the theory of confidence sequences—intervals that contain the true parameter with probability ≥ 1 - α at all stopping times, not just pre-planned ones. Johari et al. (2017) applied this to A/B testing, producing p-value sequences that maintain Type I error control under arbitrary stopping rules. The cost: slightly wider confidence intervals early in the experiment. The benefit: complete freedom to stop whenever you want without alpha inflation.

Sample Size Inflation: The Price of Flexibility

Sequential testing isn't free. The ability to stop early comes at a cost: you need to plan for a slightly larger maximum sample size to preserve power.

How Much Inflation?

For group sequential designs with 2-5 interim looks:

  • Pocock boundaries: ~15-18% inflation in maximum n
  • O'Brien-Fleming boundaries: ~2-4% inflation in maximum n

Example: Fixed-sample test needs 10,000 per group. With O'Brien-Fleming boundaries and 4 looks, maximum n increases to 10,400 per group. With Pocock boundaries, maximum n increases to 11,500 per group.

This seems like a cost, but it's actually neutral or positive in expectation. Why? Because you rarely run to maximum n. When effects are moderate or large, you stop at 50-70% of maximum n, saving 30-50% of sample size even after inflation.

Expected Sample Size vs. Maximum Sample Size

Maximum sample size is what you need in the worst case (null effect or very small effect). Expected sample size is what you'll need on average across all possible effect sizes, weighted by your prior beliefs.

For a 4-look O'Brien-Fleming design detecting medium effects (d = 0.5):

  • Maximum n: 10,400 per group (4% inflation)
  • Expected n: 7,200 per group (31% savings vs. maximum n)

You plan for 10,400 but typically use 7,200. The flexibility to stop early when evidence is clear more than compensates for the slight increase in maximum n.

When Sample Size Inflation Hurts

If you're running on very low traffic (e.g., 100 users per week) and testing small effects (2-3% lift), sequential designs offer little benefit. You'll almost always run to maximum n anyway because small effects take large samples to detect reliably.

In this scenario, the 4-18% inflation in maximum n is a pure cost with no offsetting early-stopping benefit. Use standard fixed-sample tests instead.

Rule of thumb: Sequential testing pays off when you have enough traffic to potentially reach interim stopping boundaries. If your typical experiment runs for 6+ weeks and you're desperate for every user, stick with fixed-sample designs.

Stopping for Futility: Cutting Losers Early Saves Resources

Sequential designs aren't just about stopping early when you have a winner. They also let you stop for futility—abandoning experiments that are clearly not going to reach significance even if you run to maximum n.

Conditional Power Calculations

At any interim look, calculate the conditional power: the probability of achieving significance by maximum n, given the data observed so far.

Formula:
Conditional Power = P(reject H₀ at max n | current data)

If conditional power drops below 20%, you have very little chance of success even if you continue. Consider stopping for futility and reallocating resources to your next experiment.

Worked Example: Stopping for Futility

You're testing a new product page layout. After 50% of planned sample (5,000 per group):

  • Control conversion: 3.2%
  • Variant conversion: 3.1%
  • Difference: -0.1 percentage points (variant slightly worse)
  • p = 0.78 (nowhere near significant)

Calculate conditional power: Given current data, what's the probability of achieving p < 0.05 if we continue to maximum n = 10,000 per group?

Conditional power = 8%. Even if there's a small true positive effect hidden in the noise, you have only an 8% chance of detecting it with the remaining sample.

Decision: Stop for futility. Reallocate the remaining 5,000 users per group to a different experiment with better prospects.

Cost savings: 10,000 users freed up for a different test. If your next idea has even a 30% chance of producing a meaningful lift, testing it now instead of in 3 weeks when this futile test finishes is worth far more than the <8% chance this test succeeds.

Futility Boundaries in Practice

You can formalize futility stopping with lower boundaries. While efficacy boundaries say "stop if evidence is strong enough to claim a win," futility boundaries say "stop if evidence is weak enough that continuing is wasteful."

Common futility rules:

  • Conditional power < 20%: Unlikely to succeed even with full sample
  • Observed effect in wrong direction: If variant is performing worse than control and you're past 50% of sample, cut it
  • Confidence interval excludes minimum practical effect: If the 95% CI upper bound is below your "smallest effect worth detecting," stop

Stopping for futility doesn't inflate Type I error (you're accepting H₀, not rejecting it), so you don't need boundary corrections. The risk is Type II error—stopping too early and missing a small true effect. But if your conditional power is <20%, that risk is already baked in.

Asymmetric Stopping: Different Rules for Efficacy and Futility

You can use strict boundaries for efficacy (O'Brien-Fleming) to preserve power, while using lenient boundaries for futility (conditional power < 20%) to cut losers aggressively. This asymmetry makes sense: false positives are costly (shipping changes that don't work), while false negatives on weak effects are acceptable (you'll find bigger wins elsewhere). Design your stopping rules to match your risk tolerance.

Practical Implementation: What You Actually Monitor

Let's get concrete. You've designed a group sequential test with 3 interim looks. What do you calculate at each look, and how do you make the stop/continue decision?

Step-by-Step Monitoring Process

1. Pre-compute your boundaries before starting the experiment

Use statistical software or MCP Analytics' sequential testing tool to calculate boundaries for your chosen design. For a 3-look O'Brien-Fleming design at α = 0.05:

Look Sample Fraction Efficacy Boundary (Z-score) Efficacy Boundary (p-value)
1 33% 3.47 0.0005
2 67% 2.45 0.0141
3 100% 2.00 0.0455

Print this table and keep it with your experiment documentation. These boundaries are fixed—don't recalculate them based on observed data.

2. At each interim look, calculate the test statistic

For conversion rate experiments:

p_control = conversions_control / n_control
p_variant = conversions_variant / n_variant
p_pooled = (conversions_control + conversions_variant) / (n_control + n_variant)

SE = sqrt(p_pooled * (1 - p_pooled) * (1/n_control + 1/n_variant))
Z = (p_variant - p_control) / SE
p_value = 2 * (1 - Φ(|Z|))  // two-tailed

3. Compare observed Z-score to boundary

If |Z| > boundary Z-score, stop for efficacy (reject H₀).
If |Z| < boundary Z-score, continue to next look.

4. Check futility condition (optional)

Calculate conditional power. If < 20%, consider stopping for futility.

5. Document every look

Keep a log showing date, sample size, conversion rates, Z-score, p-value, and decision at each look. This creates an audit trail proving you followed the pre-specified plan.

Example Monitoring Log

Date Look n per group Control CR Variant CR Z-score p-value Boundary Decision
2026-07-08 1 3,300 4.2% 4.9% 2.21 0.027 p < 0.0005 Continue
2026-07-15 2 6,700 4.3% 5.1% 2.89 0.0039 p < 0.0141 STOP (efficacy)

At Look 2, p = 0.0039 < 0.0141 (boundary), so we stop and declare variant the winner. Stopped at 6,700 per group instead of planned maximum 10,000 per group (33% savings).

Common Mistakes That Negate All the Benefits

Sequential testing only works if you follow the methodology rigorously. These mistakes invalidate your error control and turn "sequential testing" into "peeking with extra steps."

Mistake 1: Changing the Number of Looks Mid-Experiment

You planned for 3 looks but results at Look 2 are trending toward significance (p = 0.02) but not quite over the boundary (p < 0.0141). You think: "Let me add a 4th look to give it more chances to cross."

This is p-hacking. The number of looks must be specified before seeing any data. Adding looks after observing borderline results inflates alpha.

If you genuinely need to extend the experiment for external reasons (e.g., traffic was lower than expected), re-calculate boundaries for the new number of looks from scratch and apply them going forward. Document the change and acknowledge the analysis is no longer pre-specified.

Mistake 2: Using Standard p < 0.05 Instead of Corrected Boundaries

You've heard of sequential testing and decide to "do it" by checking results every week and stopping when p < 0.05. This is not sequential testing—it's uncorrected peeking.

Sequential testing requires using adjusted boundaries (Pocock, O'Brien-Fleming, or mSPRT thresholds). If you check at p < 0.05 without adjustment, you're still inflating alpha to 0.14–0.30 depending on how many times you check.

Mistake 3: Stopping Because You're Impatient, Not Because You Hit a Boundary

At Look 1, p = 0.08. Not significant. At Look 2, p = 0.06. Still not significant. Your boss wants to ship something this quarter, so you decide to stop and call it "marginally significant."

The boundaries exist to tell you when evidence is strong enough. Ignoring them because you want a result defeats the entire purpose.

If time pressure is real, either (a) reduce your maximum planned n to something achievable in the available time, or (b) accept that you may not get a definitive answer and plan accordingly. Don't pretend p = 0.06 is significant just because you're out of time.

Mistake 4: Applying Sequential Methods to Non-Randomized Data

Did you randomize? Sequential testing, like all hypothesis testing, requires proper experimental design. If you're analyzing observational data or a poorly randomized experiment, no amount of sequential boundary correction will save you from confounding bias.

Check your randomization first. Then apply sequential methods.

Mistake 5: Ignoring Confidence Intervals

You hit the efficacy boundary at Look 2: p = 0.008, boundary is p < 0.0141. You stop and declare victory. But the 95% confidence interval for the effect is [+0.2%, +6.5%].

That's a huge range. Yes, you have statistical significance, but do you know whether the lift is 0.2% (barely worth implementing) or 6.5% (transformational)? Always report confidence intervals alongside p-values. Sequential tests produce valid CIs, but they're wider early in the experiment due to boundary corrections.

Use the CI to assess practical significance: "We detected a statistically significant lift between 0.2% and 6.5%. Even in the worst case, this is worth shipping."

Pre-Registration Prevents Post-Hoc Rationalization

Before starting your experiment, document: (1) Number and timing of interim looks, (2) Boundary type (Pocock, O'Brien-Fleming, mSPRT), (3) Futility rules, (4) Maximum sample size. Store this in your experiment tracker or wiki. When you're tempted to deviate, this pre-registration reminds you: we agreed to the rules before seeing data. Stick to them.

When NOT to Use Sequential Testing

Sequential methods aren't always the right choice. Here's when standard fixed-sample tests make more sense.

1. Very Low Traffic or Long Experiment Durations

If your experiment will take 8+ weeks to reach minimum sample size, interim looks don't help much. You're not going to check results at week 2 and make decisions—there's too much noise. In this case, the 4-18% inflation in maximum n from sequential boundaries is a pure cost with no early-stopping benefit.

Stick with fixed-sample tests and wait for full data.

2. Testing Very Small Effects

If you're trying to detect a 0.5% lift in conversion rate, you need huge samples. Sequential methods rarely stop early for small effects because the boundaries require overwhelming evidence. You'll run to maximum n almost always, making the complexity of sequential monitoring pointless.

Use sequential testing for medium to large effects (3%+ lifts). For small effects, commit to a large fixed sample and wait.

3. Experiments with Complex Downstream Metrics

If your primary metric is "7-day retained revenue per user," you can't calculate it in real-time—you need to wait 7 days after each user's entry to measure their outcome. This delays interim looks and makes continuous monitoring impractical.

Sequential testing works best for immediate-outcome metrics (clicks, conversions, signups). For delayed metrics, simplify to fixed-sample designs with batched analysis.

4. High Stakes Decisions Requiring Maximum Confidence

If you're testing a major product change (e.g., new pricing model, redesigned checkout flow), stakeholders may want to see the full planned sample even if early results are promising. The extra confidence from running to maximum n is worth the opportunity cost.

Sequential testing optimizes for speed. Sometimes you optimize for confidence instead. Know which goal matters for each decision.

The Business Case: Quantifying Sequential Testing ROI

Let's build a financial model comparing fixed-sample vs. sequential testing for a typical e-commerce optimization program.

Assumptions

  • Monthly traffic: 200,000 users
  • Experiments per year: 12 (one per month with fixed-sample testing)
  • Average experiment duration (fixed): 21 days
  • Average experiment duration (sequential): 14 days (33% reduction from early stopping)
  • Probability of finding a winning variant: 25%
  • Average lift from winning variants: 8% revenue increase
  • Annual revenue: $5M

Fixed-Sample Testing Economics

12 experiments/year × 25% win rate = 3 winning tests
Each winner lifts revenue by 8%
Compounding lifts: 1.08³ = 1.26 (26% cumulative lift)
Revenue impact: $5M × 0.26 = $1.3M incremental revenue

Sequential Testing Economics

14-day average duration allows 365/14 = 26 experiments/year (assuming back-to-back testing)
But traffic limits you: with 200K/month, you can realistically run 18 experiments/year with sequential methods
18 experiments × 25% win rate = 4.5 winning tests
Compounding lifts: 1.08^4.5 = 1.40 (40% cumulative lift)
Revenue impact: $5M × 0.40 = $2.0M incremental revenue

Net Benefit from Sequential Testing

$2.0M - $1.3M = $700K additional revenue per year from running 50% more experiments with the same traffic and team.

Implementation cost: Minimal. If using tools like MCP Analytics for CSV data analysis or open-source packages, the switching cost is one-time setup (learning sequential methods, updating SOPs, training team). Ongoing cost is near zero.

ROI: Effectively infinite (negligible cost, $700K annual benefit).

Sensitivity Analysis

Even if sequential methods only reduce duration by 20% (instead of 33%), you'd run 15 experiments/year instead of 12—still 25% more tests and substantial incremental revenue.

The ROI math holds across a wide range of assumptions because the cost of sequential testing is trivial (it's just math) while the benefit of running more experiments is linear in the number of wins.

Key Insight: Velocity Compounds

Experimentation programs create value through cumulative learning. Each winning test lifts baseline performance, making future wins apply to a higher base. Going from 12 to 18 experiments per year doesn't just add 6 tests—it potentially adds 1.5 additional wins per year, every year. Over 3 years, that's 4-5 extra compounding improvements. The value of faster iteration is exponential, not linear.

Frequently Asked Questions

What's the actual cost of peeking at A/B test results early?

Peeking inflates your false positive rate from 5% to 30% or higher. If you check results 10 times during an experiment designed for α = 0.05, your actual Type I error rate approaches 0.30. That means 30% of your "winning" variants are false positives. The cost: wasted engineering time implementing changes that don't work, plus opportunity cost from not testing something that would have worked.

How much sample size do I save with sequential testing?

Sequential tests stop earlier when effects are large, saving 20-50% of planned sample size. If your variant truly outperforms by 15% (large effect), sequential methods often achieve significance with half the observations. For moderate effects (5-8%), savings are smaller but still meaningful—around 10-20%. The actual savings depend on true effect size and your stopping boundaries. No method saves sample size on null effects; you'll run to maximum n either way.

Can I stop a sequential test as soon as I see p < 0.05?

No. Sequential testing requires pre-defined stopping boundaries that are more conservative than p < 0.05. Methods like Pocock boundaries might require p < 0.01 at early looks, relaxing to p < 0.05 only near the end. The boundaries ensure the overall false positive rate stays at 5% across all possible stopping points. If you stop at the first p < 0.05 without boundary corrections, you're still peeking—just with a fancy name.

What's the difference between group sequential and fully sequential testing?

Group sequential tests check results at pre-defined intervals (e.g., after 25%, 50%, 75%, 100% of planned sample). Fully sequential tests allow checking after every single observation. Group sequential is more practical for most A/B tests—you check weekly or after fixed user counts. Fully sequential (like mSPRT) is optimal for high-traffic scenarios where you can afford continuous monitoring. Both control false positive rates if you follow proper boundaries.

Should I use sequential testing for all my A/B tests?

Use sequential methods when you want the flexibility to stop early or when opportunity cost is high. If you're testing on low traffic and will wait for the full sample anyway, standard fixed-sample tests are simpler. Sequential testing shines when: (1) you have enough traffic to check results multiple times before hitting maximum n, (2) early stopping would save meaningful costs, or (3) you need to monitor for harmful effects. For critical decisions with plenty of time, fixed-sample tests are fine.