Executive Summary
Does Power Output differ across the 3 Phase conditions?
The short answer
Power output increases across phases: Phase 1 averaged 99.55, Phase 2 averaged 103.188, and Phase 3 averaged 107.325. The Huynh-Feldt corrected test gives p < 0.001, and all 3 pairwise comparisons survive multiplicity correction. The effect is large (partial eta-squared 0.572 within subjects), but this is a repeated-measures design, not evidence that phase caused the change.
The detail
The repeated-measures ANOVA yields F(2.00, 94.00) = 62.723, p < 0.001 before sphericity correction. Mauchly's test rejects sphericity (W = 0.740, p < 0.001), but the Greenhouse-Geisser epsilon of 0.794 is at or above 0.75, so the Huynh-Feldt correction (epsilon 0.817) is preferred. After Huynh-Feldt correction, p < 0.001—the verdict does not change. The spread between extreme conditions is 7.78 (Phase 3 minus Phase 1). Partial eta-squared is 0.572; generalized eta-squared (comparable to a between-subjects study) is 0.147. All 3 of 3 pairwise comparisons are significant after Holm correction. The largest is Phase 1 vs Phase 3, a within-subject difference of −7.78 (95% CI −9.45 to −6.10, Cohen's dz −1.349).
What this can't tell you
The same athletes were measured throughout, so this is not a causal comparison. Order, practice, fatigue, and maturation remain confounded with phase.
Analysis Overview
Repeated-measures ANOVA of Power Output across 3 Phase conditions on 48 subjects.
The short answer
Yes, power output changes significantly across the three training phases within the same 48 athletes. A repeated-measures design isolates each athlete's stable differences from the noise, making the test more sensitive than comparing separate groups. This approach requires the assumption that differences between phase pairs have equal variance (sphericity), which is explicitly tested and corrected when violated.
The detail
The repeated-measures ANOVA fits Power Output with a separate error stratum for the 48 subjects across 3 Phase conditions. This removes stable person-to-person differences from the error term rather than leaving them in the noise—the mechanism that increases sensitivity. The design yields F = 62.723 with partial eta-squared = 0.572, a large within-subject effect. Sphericity is tested via Mauchly's test; when it fails, the degrees of freedom are corrected using either Greenhouse-Geisser or Huynh-Feldt adjustment, depending on the estimated epsilon.
What this can't tell you
The repeated-measures structure measures the same athletes across phases but does not establish causation. Differences in power output across phases reflect observed association within subjects, not proof that phase caused the change. Other time-dependent factors (fatigue, learning, seasonal variation) cannot be ruled out from the design alone.
Data Quality
Which subjects entered the model, and which could not.
The short answer
All 48 athletes with complete data across all 3 phases entered the model; none were dropped. The analysis uses every row of the 144 records loaded.
The detail
144 rows loaded; 144 rows used. All 48 values of Athlete had a complete set of all 3 Phase conditions (Phase 1, Phase 2, Phase 3 read in that order), so no subjects were excluded. Repeated-measures ANOVA requires a complete cell for every subject-condition pair; any subject with a missing condition must be dropped entirely. Here, subjects_dropped = 0 and rows_removed = 0, so the analysed sample is the same as the collected sample.
What this can't tell you
Listwise deletion—removing a subject because one condition is missing—introduces bias if the reason for missingness is related to power output itself. With zero dropouts here, that bias does not apply, but the principle remains: if data were missing, the remaining subjects would not be a neutral sample of the original set.
Condition Means
Mean Power Output under each Phase condition, with 95% confidence intervals.
The short answer
Phase 1 has the lowest mean at 99.55; Phase 3 has the highest at 107.325. The gap between them is 7.78. The error bars shown are between-subject confidence intervals and overlap substantially—a normal pattern in repeated-measures designs because the test operates on each athlete's own change, which is much smaller than the spread between athletes.
The detail
Condition means with 95% confidence intervals: Phase 1 = 99.55 (CI 97.449 to 101.652), Phase 2 = 103.188 (CI 101.089 to 105.286), Phase 3 = 107.325 (CI 104.817 to 109.833). The largest gap is 7.78 between Phase 1 and Phase 3. The error bars carry the full subject-to-subject spread—they reflect how different athletes are from each other, not how much each athlete changes. In a repeated-measures design, the test works on within-subject differences, which are much smaller than the between-subject spread, so overlapping bars are consistent with a significant test. For confidence intervals on the within-subject differences themselves, read the post-hoc table.
What this can't tell you
Overlapping error bars do not contradict the significance test. The test compares each athlete's own change across phases, not the overall distributions.
Per-Subject Profiles
Every subject's Power Output across the Phase conditions, one row per subject.
The short answer
All 48 athletes show the same directional pattern: power output increases from Phase 1 to Phase 3. Specifically, 93.75% of athletes move in the same direction as the group average, so the phase effect is consistent across individuals rather than driven by a small subset.
The detail
Each row is one athlete's profile across the three phases, ordered by average power output. A horizontal gradient within a row is the phase effect; a vertical gradient down the rows is the between-subject variation that the repeated-measures model removes. Example: ATH023 moves from 86.865 (Phase 1) to 88.31 (Phase 2) to 93.37 (Phase 3)—a steady increase. ATH048 moves from 84.394 to 93.629 to 101.133—also increasing. Of the 48 athletes, 93.75% move in the same direction from Phase 1 to Phase 3 as the group average does (upward), confirming that the effect is not driven by a minority of outliers.
What this can't tell you
Individual variation in the magnitude of change is visible here—some athletes gain more than others—but the analysis tests whether the average change is real, not whether every athlete responds identically.
Repeated-Measures ANOVA
Sources of variation, the F test, and the effect sizes.
| Source | Df | Sum Sq | F Statistic | P Value | Partial Eta Sq |
|---|---|---|---|---|---|
| Subjects (between subjects) | 47 | 7336 | — | n/a | — |
| Phase (within subjects) | 2 | 1453 | 62.72 | < 0.001 | 0.572 |
| Residual (within subjects) | 94 | 1089 | — | n/a | — |
The short answer
The phase effect is strong: F(2.00, 94.00) = 62.723, p < 0.001. Phase explains 57.2% of the within-subject variation in power output (partial eta-squared 0.572). When the between-subject variation is put back into the denominator for comparison to a between-subjects study, the effect size drops to 14.7% (generalized eta-squared 0.147).
The detail
The repeated-measures ANOVA splits variation into three strata: Subjects (between subjects, df 47, sum of squares 7336.378), Phase (within subjects, df 2, sum of squares 1452.688), and Residual (within subjects, df 94, sum of squares 1088.538). The Phase row carries the test: F(2.00, 94.00) = 62.723, p < 0.001. Partial eta-squared for Phase is 0.572—the proportion of within-subject variation explained by phase. Generalized eta-squared is 0.147, which restores the subject variance to the denominator and is the effect size comparable to a between-subjects study. The sums of squares were reproduced independently from the condition means as a verification check.
What this can't tell you
Partial eta-squared cannot be compared directly to effect sizes from between-subjects designs because it excludes subject variance from its denominator. Generalized eta-squared (0.147) is the correct figure for that comparison.
Sphericity & Corrections
Mauchly's test, both epsilon estimates, and the corrected p-values.
| Quantity | Value | Interpretation |
|---|---|---|
| Mauchly's W | 0.740 | Likelihood-ratio statistic for sphericity across the 3 conditions; 1 means the pairwise differences all have the same variance. |
| Mauchly p-value | < 0.001 | Sphericity is rejected — read a corrected p-value, not the uncorrected one. |
| Greenhouse-Geisser epsilon | 0.794 | How far the covariance departs from sphericity, on a scale from the lower bound 0.500 (worst case) to 1 (perfect). Degrees of freedom are multiplied by it. |
| Huynh-Feldt epsilon | 0.817 | A less conservative estimate of the same quantity; preferred when the Greenhouse-Geisser epsilon is at or above 0.75. |
| Uncorrected p-value | < 0.001 | The F test read at 2.00 and 94.00 degrees of freedom. |
| Greenhouse-Geisser p-value | < 0.001 | The same F read at 1.59 and 74.62 degrees of freedom. |
| Huynh-Feldt p-value | < 0.001 | The same F read at 1.63 and 76.80 degrees of freedom. |
The short answer
Sphericity is violated: Mauchly's test gives p < 0.001. However, the Greenhouse-Geisser epsilon (0.794) is at the threshold where the less conservative Huynh-Feldt correction (epsilon 0.817) is preferred. After correction, p < 0.001—the same verdict as the uncorrected test.
The detail
Mauchly's W = 0.740, p < 0.001, so the assumption that every pair of phases has the same variance of within-subject differences is rejected. Sphericity violation makes the uncorrected p-value too small (error toward false positives). The Greenhouse-Geisser epsilon of 0.794 (against a lower bound of 0.500 and a perfect value of 1.0) and Huynh-Feldt epsilon of 0.817 both estimate how badly sphericity is violated. When Greenhouse-Geisser epsilon is at or above 0.75, Huynh-Feldt is preferred. The uncorrected p-value is < 0.001; the Greenhouse-Geisser p-value is < 0.001; the Huynh-Feldt p-value is < 0.001. The correction did not change the verdict. Mauchly's test is itself sensitive to sample size and non-normality, so with 48 subjects a non-significant result would be weak evidence of sphericity rather than proof.
What this can't tell you
The sphericity violation is real and the correction is necessary, but in this case it does not alter the conclusion.
Which Conditions Differ
Pairwise comparisons across Phase, corrected for multiplicity.
| Comparison | Difference | CI Low | CI High | T Statistic | P Value | P Adjusted | Cohens Dz | Significant |
|---|---|---|---|---|---|---|---|---|
| Phase 1 vs Phase 3 | -7.775 | -9.448 | -6.101 | -9.347 | < 0.001 | < 0.001 | -1.349 | yes |
| Phase 2 vs Phase 3 | -4.137 | -5.545 | -2.73 | -5.913 | < 0.001 | < 0.001 | -0.853 | yes |
| Phase 1 vs Phase 2 | -3.638 | -4.675 | -2.6 | -7.051 | < 0.001 | < 0.001 | -1.018 | yes |
The short answer
All three phase pairs differ significantly. Phase 1 vs Phase 3 carries the largest effect: a within-subject difference of -7.78 (95% CI -9.45 to -6.10, Cohen's dz -1.35), a large effect that survives Holm correction with p < 0.001. The other two pairs also remain significant after multiplicity adjustment.
The detail
Pairwise comparisons use paired tests on each athlete's own difference, eliminating between-subject variation from the comparison. Phase 1 vs Phase 3: difference -7.775, t = -9.347, p < 0.001 (Holm-adjusted), Cohen's dz -1.349. Phase 2 vs Phase 3: difference -4.137, t = -5.913, p < 0.001 (Holm-adjusted), Cohen's dz -0.853. Phase 1 vs Phase 2: difference -3.638, t = -7.051, p < 0.001 (Holm-adjusted), Cohen's dz -1.018. All 3 of 3 comparisons survive the Holm correction. Confidence intervals are narrow because they rest on within-subject differences, not the spread across subjects.
What this can't tell you
These comparisons describe observed within-subject differences across phases; they do not establish that phase caused the change in power output.
Methods & Disclosure
The exact model, formulas, and what the design does not establish.
| Item | Detail |
|---|---|
| Model | aov(Power Output ~ Phase with Error(Athlete / Phase)) — the within-subject error stratum is separated from the subject stratum, which is what removes the stable subject-to-subject differences from the test. |
| Subjects analysed | 48 subject(s) with a complete set of all 3 conditions; 0 subject(s) dropped as incomplete. Replicate measurements in the same subject-condition cell were averaged (0 cell(s) affected). |
| Sphericity test | Mauchly's likelihood-ratio test computed directly from the covariance of the repeated measures: W = det(T) divided by the mean-of-eigenvalues to the power 2, where T is the covariance projected onto an orthonormal basis of the 2 contrasts orthogonal to the unit vector, with the standard chi-square approximation and its second-order refinement. |
| Epsilon corrections | Greenhouse-Geisser epsilon is the squared trace of T divided by 2 times the trace of T squared; Huynh-Feldt rescales it using the 48 subject(s) and 1 between-subjects group(s). Both multiply the numerator and denominator degrees of freedom before the F is read. |
| Effect sizes | Partial eta-squared is the effect's sum of squares over itself plus its own error stratum (0.572 here). Generalized eta-squared puts the subject variance back into the denominator (0.147), which is why it is much smaller and is the one comparable to a between-subjects study. |
| Post-hoc comparisons | All 3 pairs of conditions compared with a paired t-test on each subject's own difference, corrected for multiplicity with base p.adjust using the Holm method. |
| Sums-of-squares cross-check | The condition sum of squares from aov was reproduced from first principles (n times the squared deviations of the condition means); relative disagreement 0.00000000. |
| What this does not establish | A repeated-measures design removes stable differences between subjects, but it does not make the comparison causal. Anything else that changed alongside 'Phase' — practice, fatigue, maturation, the order the conditions were run in — remains confounded with it. |
The short answer
The model aov(Power Output ~ Phase with Error(Athlete / Phase)) separates the within-subject error stratum from the subject stratum, which removes stable athlete-to-athlete differences from the test and makes it more sensitive to phase changes. All 48 athletes with complete data entered the model; 0 were dropped. Sphericity was tested with Mauchly's test and corrected with Huynh-Feldt. Post-hoc comparisons used paired t-tests with Holm correction for multiplicity.
The detail
Model: aov(Power Output ~ Phase with Error(Athlete / Phase)). Subjects analysed: 48 complete, 0 dropped. Sphericity test: Mauchly's W computed directly from the covariance of the repeated measures, with standard chi-square approximation. Epsilon corrections: Greenhouse-Geisser epsilon = (trace of T)² ÷ (2 × trace of T²); Huynh-Feldt rescales using 48 subjects and 1 between-subjects group. Both multiply numerator and denominator degrees of freedom. Effect sizes: partial eta-squared (effect sum of squares ÷ effect sum of squares + error, = 0.572); generalized eta-squared restores subject variance to denominator (= 0.147). Post-hoc: all 3 pairs compared with paired t-tests, Holm correction. Sums-of-squares cross-check: condition sum of squares reproduced from condition means, relative disagreement 0.00000000.
What this can't tell you
A repeated-measures design removes stable differences between athletes and nothing else. Order, practice, fatigue, maturation, and any other change alongside Phase remains confounded with it. This is not a causal design.
Methodology
Statistical methodology and diagnostics for Repeated-Measures ANOVA
Statistical Method
Standard-library analysis for within-subject designs, with sphericity handled properly. Map a subject identifier, the condition or timepoint each measurement was taken under, and the measured value in long format, and get the repeated-measures ANOVA that removes the stable subject-to-subject differences from the error term: the F test with partial and generalized eta-squared, Mauchly's test of sphericity with the Greenhouse-Geisser and Huynh-Feldt corrected p-values, pairwise post-hoc comparisons corrected for multiplicity, condition means with confidence intervals, and a per-subject profile grid. An optional between-subjects group turns it into a mixed design with a group-by-condition interaction.
- Each row is one subject measured under one condition, in long format, and every analysed subject has a measurement under every condition
- The measurement is numeric or cleanly convertible
- The residuals are approximately normal within each condition
- Sphericity — that all pairwise differences between conditions have the same variance — which is tested explicitly and corrected for when rejected
- Subjects are independent of one another, even though their own repeated measurements are not
- Subjects without a measurement under every condition are dropped entirely rather than imputed, so the analysed set can differ systematically from the one that was collected
- At least 5 complete subjects and 3 conditions are required; the sphericity machinery additionally needs at least as many subjects as conditions
- A within-subject design removes stable differences between subjects and nothing else — order, practice, fatigue, and maturation remain confounded with the conditions, so a significant result is not a causal estimate
- Mauchly's test is sensitive to sample size and to non-normality, so a non-significant result is weak evidence that sphericity holds rather than proof of it
Analysis Code
Complete R source code for this analysis
Repeated-Measures ANOVA — The Same Subjects Across Several Conditions
Every subject is measured under three or more conditions (timepoints, treatments, tasks). Comparing those condition means with an ordinary ANOVA would be wrong: the measurements share a subject, so they are correlated. A repeated-measures ANOVA splits the stable subject-to-subject differences out of the error term, which is what makes the test both correct and far more powerful than the between-subjects version.
Why This Method?
Removing the between-subject variance shrinks the error term, so a real condition effect shows up with far fewer subjects. The price is a new assumption — sphericity, that all pairwise differences between conditions have the same variance. When it fails the F test is anti-conservative (the p-value is too small), so this module tests it (Mauchly) and reports the Greenhouse-Geisser and Huynh-Feldt corrected p-values, leading with the corrected result whenever the assumption is rejected.
What This Analysis Covers
- Condition means with confidence intervals (bar chart with error bars)
- Per-subject profiles across the conditions (profile grid)
- The repeated-measures ANOVA table with partial and generalized eta-squared
- Mauchly's test of sphericity plus both epsilon corrections
- Pairwise post-hoc comparisons with a multiplicity correction
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {subject, condition, value, group}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting
# subject_h/condition_h/value_h/group_h $ humanized user column names
# n_subjects / n_dropped_subjects $ complete subjects used / dropped
# n_conditions / cond_levels $ the within factor, in reading order
# n_replicate_cells $ subject-condition cells averaged
# f_stat/df1/df2/p_uncorr $ the within-subject F test
# mauchly_W/mauchly_p/sphericity_ok $ Mauchly's test of sphericity
# eps_gg/eps_hf/eps_lb $ epsilon corrections + lower bound
# p_gg/p_hf $ corrected p-values
# headline_label/headline_p $ the result the narrative leads with
# partial_eta2/generalized_eta2 $ effect sizes
# mixed / group_levels / inter_* $ optional between-subjects factor
# anova_df/sphericity_df/posthoc_df $ card tables
# condition_means_df/profiles_df $ chart datasets
# methods_df $ disclosure table
# metrics / json_output
# === /SHARED EXPORTS ===Step 1: Resolve the mapped columns (humanized for every sentence)
initial_rows <- nrow(df)
subject_h <- humanize_semantic("subject", col_map)
condition_h <- humanize_semantic("condition", col_map)
value_h <- humanize_semantic("value", col_map)
group_h <- humanize_semantic("group", col_map)
missing_keys <- setdiff(c("subject", "condition", "value"), names(df))
if (length(missing_keys) > 0) {
stop(sprintf("A repeated-measures analysis needs three mapped columns: the subject identifier('%s'), the within-subject condition ('%s'), and the measured value ('%s'). Missing: %s.",
subject_h, condition_h, value_h,
paste(humanize_semantic(missing_keys, col_map), collapse = ", ")))
}Step 2: Coerce the measurement to numeric with the 95% rule
v <- df$value
if (!is.numeric(v)) {
ch <- as.character(v)
non_blank <- !is.na(ch) & trimws(ch) != ""
conv <- suppressWarnings(as.numeric(ch))
if (sum(non_blank) == 0 ||
sum(!is.na(conv[non_blank])) < 0.95 * sum(non_blank)) {
stop(sprintf("The measurement column '%s' does not look numeric — fewer than 95%% of its values parse as numbers. Map a numeric column as the measurement.",
value_h))
}
v <- conv
}
df$value <- as.numeric(v)Step 3: Clean the identifiers; blank subject or condition is unusable
sid <- trimws(as.character(df$subject))
cnd <- trimws(as.character(df$condition))
bad_id <- is.na(sid) | sid == "" | is.na(cnd) | cnd == ""
n_blank_id <- sum(bad_id)
df <- df[!bad_id, , drop = FALSE]
sid <- sid[!bad_id]; cnd <- cnd[!bad_id]
if (nrow(df) == 0) {
stop(sprintf("Every row is missing either '%s' or '%s', so no repeated measurement could be assembled.",
subject_h, condition_h))
}Step 4: Sanity-check the within-subject factor before anything else
cond_levels_raw <- unique(cnd)
k_raw <- length(cond_levels_raw)
if (k_raw < 2) {
stop(sprintf("'%s' has only one distinct value ('%s'), so there is nothing to compare across conditions. A repeated-measures design needs three or more conditions measured on the same subjects.",
condition_h, cond_levels_raw[1]))
}
if (k_raw == 2) {
stop(sprintf("'%s' has exactly two levels ('%s' and '%s'). That is a two-condition design, not a repeated-measures ANOVA — use the paired comparison tool (Before vs After), which gives you the paired t-test, the Wilcoxon signed-rank cross-check, and a confidence interval on the change.",
condition_h, order_levels(cond_levels_raw)[1], order_levels(cond_levels_raw)[2]))
}
if (k_raw > 12) {
stop(sprintf("'%s' has %s distinct values. A repeated-measures factor is a handful of conditions or timepoints measured on every subject; %s values looks like an identifier or a continuous measurement rather than a condition. Map the column that names the condition.",
condition_h, fmt_n(k_raw), fmt_n(k_raw)))
}Step 5: Average replicate subject-condition cells (reported)
cell_key <- paste(sid, "\r", cnd)
n_replicate_cells <- sum(duplicated(cell_key))
n_rows_missing_value <- sum(is.na(df$value))
agg <- stats::aggregate(list(value = df$value),
by = list(subject = sid, condition = cnd),
FUN = function(x) if (all(is.na(x))) NA_real_ else mean(x, na.rm = TRUE))Carry the optional between-subjects group along (first non-blank per subject)
has_group_col <- "group" %in% names(df)
subj_group <- NULL
if (has_group_col) {
gvals <- trimws(as.character(df$group))
gvals[is.na(gvals) | gvals == ""] <- "Missing"
first_g <- tapply(gvals, sid, function(x) x[1])
subj_group <- setNames(as.character(first_g), names(first_g))
}Step 6: Completeness — a subject needs every condition, non-missing
cond_levels <- order_levels(cond_levels_raw)
k <- length(cond_levels)
ok_cells <- agg[!is.na(agg$value), , drop = FALSE]
per_subj <- table(ok_cells$subject)
n_subjects_seen <- length(unique(agg$subject))
complete_subj <- names(per_subj)[per_subj == k]
n_subjects <- length(complete_subj)
n_dropped_subjects <- n_subjects_seen - n_subjects
max_cells <- if (length(per_subj) > 0) max(as.integer(per_subj)) else 0L
if (n_subjects == 0) {
if (max_cells <= 1) {
stop(sprintf("No value of '%s' appears under more than one '%s' — every row is a different subject. That is a between-subjects design, not repeated measures: use the group comparison tool, which compares independent groups with ANOVA, Kruskal-Wallis, and pairwise tests.",
subject_h, condition_h))
}
stop(sprintf("Not one of the %s values of '%s' has a measurement under all %s levels of '%s'. Repeated-measures ANOVA needs a complete set of conditions per subject; the most any subject has here is %s.",
fmt_n(n_subjects_seen), subject_h, fmt_n(k), condition_h, fmt_n(max_cells)))
}
if (n_subjects < 5) {
stop(sprintf("Only %s value(s) of '%s' have a complete set of all %s '%s' conditions (%s were dropped as incomplete). At least 5 complete subjects are needed for a repeated-measures ANOVA.",
fmt_n(n_subjects), subject_h, fmt_n(k), condition_h,
fmt_n(n_dropped_subjects)))
}Step 7: Build the balanced subject-by-condition matrix
keep <- ok_cells$subject %in% complete_subj
long <- ok_cells[keep, , drop = FALSE]
long$subject <- factor(long$subject, levels = sort(complete_subj))
long$condition <- factor(long$condition, levels = cond_levels)
long <- long[order(long$subject, long$condition), , drop = FALSE]
Y <- matrix(long$value, nrow = n_subjects, ncol = k, byrow = TRUE,
dimnames = list(levels(long$subject), cond_levels))
final_rows <- nrow(long)
rows_removed <- initial_rows - final_rows
if (isTRUE(stats::var(as.numeric(Y)) == 0) || is.na(stats::var(as.numeric(Y)))) {
stop(sprintf("'%s' is constant — every one of the %s retained measurements is the same value (%s). There is no variation to attribute to '%s'.",
value_h, fmt_n(final_rows), r3(Y[1, 1]), condition_h))
}Step 8: The optional between-subjects factor (mixed design)
mixed <- FALSE
group_levels <- character(0)
group_note <- ""
gvec <- NULL
if (!is.null(subj_group)) {
gv <- unname(subj_group[levels(long$subject)])
gv[is.na(gv)] <- "Missing"
tabg <- sort(table(gv), decreasing = TRUE)
if (length(tabg) > 8) {
gv[gv %in% names(tabg)[-(1:8)]] <- "Other"
tabg <- sort(table(gv), decreasing = TRUE)
}
tiny <- names(tabg)[tabg < 2]
if (length(tiny) > 0) gv[gv %in% tiny] <- NA_character_
if (length(unique(stats::na.omit(gv))) >= 2 && !anyNA(gv)) {
gvec <- factor(gv, levels = order_levels(unique(gv)))
group_levels <- levels(gvec)
mixed <- TRUE
long$group <- rep(gvec, each = k)
} else {
group_note <- sprintf("The between-subjects column '%s' was mapped but did not split the complete subjects into two or more usable groups, so the analysis was run as a pure within-subject design. ",
group_h)
}
}
n_between_groups <- if (mixed) length(group_levels) else 1LStep 9: The repeated-measures ANOVA, via base aov with an Error stratum
fit <- if (mixed)
suppressWarnings(stats::aov(value ~ group * condition + Error(subject / condition),
data = long))
else
suppressWarnings(stats::aov(value ~ condition + Error(subject / condition),
data = long))
s <- summary(fit)
find_term <- function(s, term) {
for (st in names(s)) {
tb <- s[[st]][[1]]
rn <- trimws(rownames(tb))
if (term %in% rn) {
i <- which(rn == term)[1]
j <- which(rn == "Residuals")[1]
return(list(
df = tb[["Df"]][i], ss = tb[["Sum Sq"]][i],
f = tb[["F value"]][i], p = tb[["Pr(>F)"]][i],
res_df = if (!is.na(j)) tb[["Df"]][j] else NA_real_,
res_ss = if (!is.na(j)) tb[["Sum Sq"]][j] else NA_real_
))
}
}
NULL
}
cond_eff <- find_term(s, "condition")
if (is.null(cond_eff) || is.na(cond_eff$f)) {
stop(sprintf("The repeated-measures F test for '%s' could not be computed from the %s complete subjects — the design may be degenerate (for example every subject identical across conditions).",
condition_h, fmt_n(n_subjects)))
}Subject stratum residual = the between-subject variance the design removes
subj_res_ss <- {
tb <- s[[1]][[1]]
rn <- trimws(rownames(tb))
j <- which(rn == "Residuals")[1]
if (!is.na(j)) tb[["Sum Sq"]][j] else NA_real_
}
f_stat <- cond_eff$f
df1 <- cond_eff$df
df2 <- cond_eff$res_df
p_uncorr <- cond_eff$p
ss_cond <- cond_eff$ss
ss_err <- cond_eff$res_ssIndependent cross-check of the same sums of squares from first principles
grand <- mean(Y)
ss_cond_direct <- n_subjects * sum((colMeans(Y) - grand)^2)
ss_check <- if (!mixed && is.finite(ss_cond) && ss_cond > 0)
abs(ss_cond - ss_cond_direct) / ss_cond else NA_real_
partial_eta2 <- if (is.finite(ss_cond) && is.finite(ss_err) && (ss_cond + ss_err) > 0)
ss_cond / (ss_cond + ss_err) else NA_real_
gen_denom <- sum(c(ss_cond, subj_res_ss, ss_err), na.rm = TRUE)
generalized_eta2 <- if (is.finite(gen_denom) && gen_denom > 0)
ss_cond / gen_denom else NA_real_
inter_eff <- if (mixed) find_term(s, "group:condition") else NULL
group_eff <- if (mixed) find_term(s, "group") else NULL
inter_p <- if (!is.null(inter_eff)) inter_eff$p else NA_real_
inter_peta2 <- if (!is.null(inter_eff) && is.finite(inter_eff$ss) && is.finite(ss_err) &&
(inter_eff$ss + ss_err) > 0)
inter_eff$ss / (inter_eff$ss + ss_err) else NA_real_Step 10: Sphericity — Mauchly's W and the epsilon corrections,
computed directly from the covariance of the repeated measures. M is an orthonormal basis of the contrast space orthogonal to 1, T = M' SSCP M, and everything below is a function of T only.
p_dim <- k - 1
df_e_cov <- n_subjects - n_between_groups
mauchly_W <- NA_real_; mauchly_p <- NA_real_
eps_gg <- NA_real_; eps_hf <- NA_real_
eps_lb <- 1 / p_dim
sph_note <- ""
if (df_e_cov >= p_dim) {
SSCP <- if (mixed) {
Reduce(`+`, lapply(group_levels, function(l) {
Yi <- Y[gvec == l, , drop = FALSE]
if (nrow(Yi) < 2) matrix(0, k, k) else (nrow(Yi) - 1) * stats::cov(Yi)
}))
} else {
(n_subjects - 1) * stats::cov(Y)
}
M <- qr.Q(qr(matrix(1 / sqrt(k), nrow = k, ncol = 1)), complete = TRUE)[, -1, drop = FALSE]
Tm <- t(M) %*% SSCP %*% M
detT <- suppressWarnings(det(Tm))
trT <- sum(diag(Tm))
if (is.finite(detT) && detT > 0 && is.finite(trT) && trT > 0) {Mauchly's likelihood-ratio statistic with the standard chi-square approximation and its second-order (w2) refinement.
logW <- log(detT) - p_dim * log(trT / p_dim)
mauchly_W <- exp(logW)
rho <- 1 - (2 * p_dim^2 + p_dim + 2) / (6 * p_dim * df_e_cov)
z <- -df_e_cov * rho * logW
f_df <- p_dim * (p_dim + 1) / 2 - 1
w2 <- (p_dim + 2) * (p_dim - 1) * (p_dim - 2) *
(2 * p_dim^3 + 6 * p_dim^2 + 3 * p_dim + 2) / (288 * (df_e_cov * p_dim * rho)^2)
pr1 <- stats::pchisq(z, f_df, lower.tail = FALSE)
pr2 <- stats::pchisq(z, f_df + 4, lower.tail = FALSE)
mauchly_p <- min(1, max(0, pr1 + w2 * (pr2 - pr1)))Greenhouse-Geisser epsilon from the eigenvalues of T
ev <- suppressWarnings(eigen(Tm, symmetric = TRUE, only.values = TRUE)$values)
ev <- ev[is.finite(ev)]
if (length(ev) == p_dim && sum(ev^2) > 0) {
eps_gg <- min(1, max(eps_lb, sum(ev)^2 / (p_dim * sum(ev^2))))Huynh-Feldt epsilon (N subjects, g between-subjects groups)
hf_num <- n_subjects * p_dim * eps_gg - 2
hf_den <- p_dim * (n_subjects - n_between_groups - p_dim * eps_gg)
eps_hf <- if (is.finite(hf_den) && hf_den > 0)
min(1, max(eps_lb, hf_num / hf_den)) else NA_real_
}
} else {
sph_note <- sprintf("The covariance of the repeated measures is singular(often because two '%s' levels carry identical values), so the sphericity test and its corrections could not be computed. ",
condition_h)
}
} else {
sph_note <- sprintf("Sphericity needs at least as many subjects as conditions to estimate; with %s complete subjects and %s conditions it is not estimable, so no correction could be applied. ",
fmt_n(n_subjects), fmt_n(k))
}
p_gg <- if (!is.na(eps_gg))
stats::pf(f_stat, eps_gg * df1, eps_gg * df2, lower.tail = FALSE) else NA_real_
p_hf <- if (!is.na(eps_hf))
stats::pf(f_stat, eps_hf * df1, eps_hf * df2, lower.tail = FALSE) else NA_real_
sphericity_ok <- !is.na(mauchly_p) && mauchly_p >= 0.05
sphericity_known <- !is.na(mauchly_p)Girden's rule: below 0.75 the Greenhouse-Geisser correction, at or above it the less conservative Huynh-Feldt.
if (!sphericity_known) {
headline_label <- "Uncorrected F test"; headline_p <- p_uncorr
headline_reason <- sph_note
} else if (sphericity_ok) {
headline_label <- "Uncorrected F test"; headline_p <- p_uncorr
headline_reason <- sprintf("Mauchly's test does not reject sphericity (%s), so the uncorrected F test is the one to read. ",
fmt_pp(mauchly_p))
} else if (!is.na(eps_gg) && eps_gg < 0.75) {
headline_label <- "Greenhouse-Geisser corrected"; headline_p <- p_gg
headline_reason <- sprintf("Mauchly's test rejects sphericity (%s) and the estimated epsilon is %s, below 0.75, so the Greenhouse-Geisser correction is the result to lead with. ",
fmt_pp(mauchly_p), r3(eps_gg))
} else if (!is.na(eps_hf)) {
headline_label <- "Huynh-Feldt corrected"; headline_p <- p_hf
headline_reason <- sprintf("Mauchly's test rejects sphericity (%s) but the estimated epsilon is %s, at or above 0.75, so the less conservative Huynh-Feldt correction is the result to lead with. ",
fmt_pp(mauchly_p), r3(eps_gg))
} else {
headline_label <- "Uncorrected F test"; headline_p <- p_uncorr
headline_reason <- sph_note
}
headline_sig <- !is.na(headline_p) && headline_p < 0.05
flips <- !is.na(p_uncorr) && !is.na(headline_p) &&
((p_uncorr < 0.05) != (headline_p < 0.05))Step 11: Post-hoc pairwise comparisons with a multiplicity correction
adj_method <- tolower(as.character(params$p_adjust %||% "holm"))
if (!(adj_method %in% c("holm", "bonferroni", "bh", "by", "hochberg", "hommel", "none"))) {
adj_method <- "holm"
}
adj_arg <- switch(adj_method, bh = "BH", by = "BY", adj_method)
adj_label <- switch(adj_method,
holm = "Holm", bonferroni = "Bonferroni",
bh = "Benjamini-Hochberg", by = "Benjamini-Yekutieli",
hochberg = "Hochberg", hommel = "Hommel",
none = "none(raw p-values)")
pairs <- utils::combn(k, 2)
pw <- lapply(seq_len(ncol(pairs)), function(i) {
ia <- pairs[1, i]; ib <- pairs[2, i]
d <- Y[, ia] - Y[, ib]
sd_d <- stats::sd(d)
tt <- tryCatch(stats::t.test(d), error = function(e) NULL)
data.frame(
comparison = paste0(cond_levels[ia], " vs ", cond_levels[ib]),
difference = round(mean(d), 3),
ci_low = if (!is.null(tt)) round(as.numeric(tt$conf.int)[1], 3) else NA_real_,
ci_high = if (!is.null(tt)) round(as.numeric(tt$conf.int)[2], 3) else NA_real_,
t_statistic = if (!is.null(tt)) round(unname(tt$statistic), 3) else NA_real_,
p_raw_num = if (!is.null(tt)) tt$p.value else NA_real_,
cohens_dz = if (!is.na(sd_d) && sd_d > 0) round(mean(d) / sd_d, 3) else NA_real_,
stringsAsFactors = FALSE
)
})
posthoc_all <- do.call(rbind, pw)
posthoc_all$p_adj_num <- if (adj_method == "none") posthoc_all$p_raw_num else
stats::p.adjust(posthoc_all$p_raw_num, method = adj_arg)
n_sig_pairs <- sum(!is.na(posthoc_all$p_adj_num) & posthoc_all$p_adj_num < 0.05)
n_pairs <- nrow(posthoc_all)
ord <- order(-abs(posthoc_all$difference))
posthoc_ranked <- posthoc_all[ord, , drop = FALSE]
posthoc_df <- head(posthoc_ranked, 15)
posthoc_df <- data.frame(
comparison = posthoc_df$comparison,
difference = posthoc_df$difference,
ci_low = posthoc_df$ci_low, ci_high = posthoc_df$ci_high,
t_statistic = posthoc_df$t_statistic,
p_value = sapply(posthoc_df$p_raw_num, fmt_p),
p_adjusted = sapply(posthoc_df$p_adj_num, fmt_p),
cohens_dz = posthoc_df$cohens_dz,
significant = ifelse(is.na(posthoc_df$p_adj_num), "",
ifelse(posthoc_df$p_adj_num < 0.05, "yes", "no")),
stringsAsFactors = FALSE
)
rownames(posthoc_df) <- NULL
top_pair <- posthoc_ranked[1, ]Step 12: Condition means with confidence intervals (the bar chart)
condition_means_df <- do.call(rbind, lapply(seq_len(k), function(j) {
x <- Y[, j]
m <- mean(x); sdev <- stats::sd(x)
half <- if (!is.na(sdev) && sdev > 0 && n_subjects > 1)
stats::qt(0.975, n_subjects - 1) * sdev / sqrt(n_subjects) else 0
data.frame(condition = cond_levels[j], mean_value = round(m, 3),
ci_low = round(m - half, 3), ci_high = round(m + half, 3),
sd_value = round(if (is.na(sdev)) 0 else sdev, 3),
n = n_subjects, stringsAsFactors = FALSE)
}))
rownames(condition_means_df) <- NULLNA-safe extremes — never which.max over a possibly-all-NA vector
fin <- which(is.finite(condition_means_df$mean_value))
hi_i <- if (length(fin) > 0) fin[which.max(condition_means_df$mean_value[fin])] else 1L
lo_i <- if (length(fin) > 0) fin[which.min(condition_means_df$mean_value[fin])] else 1L
highest_cond <- condition_means_df$condition[hi_i]
lowest_cond <- condition_means_df$condition[lo_i]
spread <- condition_means_df$mean_value[hi_i] - condition_means_df$mean_value[lo_i]Step 13: Per-subject profile grid — up to 60 subjects, seeded sample
subj_means <- rowMeans(Y)
show_n <- min(60, n_subjects)
set.seed(42)
pick <- if (n_subjects > show_n) sort(sample(n_subjects, show_n)) else seq_len(n_subjects)
pick <- pick[order(subj_means[pick])]
profiles_df <- do.call(rbind, lapply(pick, function(i) {
data.frame(condition = cond_levels,
subject_label = rownames(Y)[i],
value = round(as.numeric(Y[i, ]), 3),
stringsAsFactors = FALSE)
}))
rownames(profiles_df) <- NULLHow consistent are the profiles? The share of subjects whose own change from the lowest-mean to the highest-mean condition runs the same way.
same_dir <- mean((Y[, hi_i] - Y[, lo_i]) * sign(spread) > 0)
pct_same_dir <- 100 * (if (is.finite(same_dir)) same_dir else 0)Step 14: The ANOVA table
anova_rows <- list()
add_anova <- function(source, dfv, ssv, fv, pv, peta) {
anova_rows[[length(anova_rows) + 1]] <<- data.frame(
source = source,
df = as.numeric(dfv),
sum_sq = round(as.numeric(ssv), 3),
f_statistic = if (is.na(fv)) NA_real_ else round(as.numeric(fv), 3),
p_value = fmt_p(pv),
partial_eta_sq = if (is.na(peta)) NA_real_ else round(as.numeric(peta), 3),
stringsAsFactors = FALSE
)
}
if (mixed && !is.null(group_eff)) {
g_peta <- if (is.finite(group_eff$ss) && is.finite(group_eff$res_ss) &&
(group_eff$ss + group_eff$res_ss) > 0)
group_eff$ss / (group_eff$ss + group_eff$res_ss) else NA_real_
add_anova(sprintf("%s(between subjects)", group_h), group_eff$df, group_eff$ss,
group_eff$f, group_eff$p, g_peta)
add_anova("Subjects within groups(error)", group_eff$res_df, group_eff$res_ss,
NA_real_, NA_real_, NA_real_)
} else {
add_anova("Subjects(between subjects)",
if (is.finite(subj_res_ss)) n_subjects - 1 else NA_real_,
subj_res_ss, NA_real_, NA_real_, NA_real_)
}
add_anova(sprintf("%s(within subjects)", condition_h), df1, ss_cond,
f_stat, p_uncorr, partial_eta2)
if (mixed && !is.null(inter_eff)) {
add_anova(sprintf("%s by %s(interaction)", group_h, condition_h),
inter_eff$df, inter_eff$ss, inter_eff$f, inter_eff$p, inter_peta2)
}
add_anova("Residual(within subjects)", df2, ss_err, NA_real_, NA_real_, NA_real_)
anova_df <- do.call(rbind, anova_rows)
rownames(anova_df) <- NULLStep 15: The sphericity table
sph_rows <- list()
add_sph <- function(quantity, value, interpretation) {
sph_rows[[length(sph_rows) + 1]] <<- data.frame(
quantity = quantity, value = value, interpretation = interpretation,
stringsAsFactors = FALSE)
}
add_sph("Mauchly's W",
if (is.na(mauchly_W)) "not estimable" else r3(mauchly_W),
sprintf("Likelihood-ratio statistic for sphericity across the %s conditions; 1 means the pairwise differences all have the same variance.",
fmt_n(k)))
add_sph("Mauchly p-value",
if (is.na(mauchly_p)) "not estimable" else fmt_p(mauchly_p),
if (!sphericity_known) "Sphericity could not be tested on this data."
else if (sphericity_ok) "Sphericity is not rejected — the uncorrected F test is appropriate."
else "Sphericity is rejected — read a corrected p-value, not the uncorrected one.")
add_sph("Greenhouse-Geisser epsilon",
if (is.na(eps_gg)) "not estimable" else r3(eps_gg),
sprintf("How far the covariance departs from sphericity, on a scale from the lower bound %s(worst case) to 1 (perfect). Degrees of freedom are multiplied by it.",
r3(eps_lb)))
add_sph("Huynh-Feldt epsilon",
if (is.na(eps_hf)) "not estimable" else r3(eps_hf),
"A less conservative estimate of the same quantity; preferred when the Greenhouse-Geisser epsilon is at or above 0.75.")
add_sph("Uncorrected p-value", fmt_p(p_uncorr),
sprintf("The F test read at %s and %s degrees of freedom.", r2(df1), r2(df2)))
add_sph("Greenhouse-Geisser p-value",
if (is.na(p_gg)) "not estimable" else fmt_p(p_gg),
if (is.na(eps_gg)) "Not available." else
sprintf("The same F read at %s and %s degrees of freedom.",
r2(eps_gg * df1), r2(eps_gg * df2)))
add_sph("Huynh-Feldt p-value",
if (is.na(p_hf)) "not estimable" else fmt_p(p_hf),
if (is.na(eps_hf)) "Not available." else
sprintf("The same F read at %s and %s degrees of freedom.",
r2(eps_hf * df1), r2(eps_hf * df2)))
sphericity_df <- do.call(rbind, sph_rows)
rownames(sphericity_df) <- NULLStep 16: Methods and disclosure
methods_df <- data.frame(
item = c(
"Model",
"Subjects analysed",
"Sphericity test",
"Epsilon corrections",
"Effect sizes",
"Post-hoc comparisons",
"Sums-of-squares cross-check",
"What this does not establish"),
detail = c(
if (mixed)
sprintf("aov(%s ~ %s by %s with Error(%s / %s)) — a mixed design with '%s' between subjects and '%s' within subjects.",
value_h, group_h, condition_h, subject_h, condition_h, group_h, condition_h)
else
sprintf("aov(%s ~ %s with Error(%s / %s)) — the within-subject error stratum is separated from the subject stratum, which is what removes the stable subject-to-subject differences from the test.",
value_h, condition_h, subject_h, condition_h),
sprintf("%s subject(s) with a complete set of all %s conditions; %s subject(s) dropped as incomplete. Replicate measurements in the same subject-condition cell were averaged(%s cell(s) affected).",
fmt_n(n_subjects), fmt_n(k), fmt_n(n_dropped_subjects), fmt_n(n_replicate_cells)),
sprintf("Mauchly's likelihood-ratio test computed directly from the covariance of the repeated measures: W = det(T) divided by the mean-of-eigenvalues to the power %s, where T is the covariance projected onto an orthonormal basis of the %s contrasts orthogonal to the unit vector, with the standard chi-square approximation and its second-order refinement.",
fmt_n(p_dim), fmt_n(p_dim)),
sprintf("Greenhouse-Geisser epsilon is the squared trace of T divided by %s times the trace of T squared; Huynh-Feldt rescales it using the %s subject(s) and %s between-subjects group(s). Both multiply the numerator and denominator degrees of freedom before the F is read.",
fmt_n(p_dim), fmt_n(n_subjects), fmt_n(n_between_groups)),
sprintf("Partial eta-squared is the effect's sum of squares over itself plus its own error stratum (%s here). Generalized eta-squared puts the subject variance back into the denominator (%s), which is why it is much smaller and is the one comparable to a between-subjects study.",
r3(partial_eta2), r3(generalized_eta2)),
sprintf("All %s pairs of conditions compared with a paired t-test on each subject's own difference, corrected for multiplicity with base p.adjust using the %s method.",
fmt_n(n_pairs), adj_label),
if (is.na(ss_check)) "Not applicable to this design."
else sprintf("The condition sum of squares from aov was reproduced from first principles(n times the squared deviations of the condition means); relative disagreement %s.",
formatC(ss_check, format = "f", digits = 8)),
sprintf("A repeated-measures design removes stable differences between subjects, but it does not make the comparison causal. Anything else that changed alongside '%s' — practice, fatigue, maturation, the order the conditions were run in — remains confounded with it.",
condition_h))
, stringsAsFactors = FALSE)Step 17: Metrics + the one-paragraph computed answer
metrics <- list(
`Subjects` = n_subjects,
`Conditions` = k,
`F Statistic` = round(f_stat, 3),
`Headline Test` = headline_label,
`Headline p-value` = fmt_p(headline_p),
`Partial Eta-Squared` = if (is.na(partial_eta2)) NA_real_ else round(partial_eta2, 3),
`Sphericity` = if (!sphericity_known) "not estimable"
else if (sphericity_ok) "not rejected" else "violated",
`Significant Pairs` = as.integer(n_sig_pairs)
)
verdict <- if (headline_sig)
sprintf("%s differs across the %s levels of %s", value_h, fmt_n(k), condition_h)
else
sprintf("no reliable difference in %s across the %s levels of %s was found",
value_h, fmt_n(k), condition_h)
json_output <- list(
answer = paste0(
"Repeated-measures ANOVA of ", value_h, " across ", fmt_n(k), " ", condition_h,
" conditions measured on the same ", fmt_n(n_subjects), " subjects: ",
verdict, " (F(", r2(df1), ", ", r2(df2), ") = ", r2(f_stat), ", ",
headline_label, ", ", fmt_pp(headline_p), "). ",
"Partial eta-squared is ", r3(partial_eta2), " (", eta_word(partial_eta2),
") and generalized eta-squared ", r3(generalized_eta2), ". ",
if (!sphericity_known) sph_note
else if (sphericity_ok)
paste0("Mauchly's test does not reject sphericity (", fmt_pp(mauchly_p),
"), so the uncorrected F test stands. ")
else
paste0("Mauchly's test rejects sphericity (", fmt_pp(mauchly_p),
", Greenhouse-Geisser epsilon ", r3(eps_gg),
"), so the corrected p-value is the one to read. "),
fmt_n(n_sig_pairs), " of ", fmt_n(n_pairs),
" pairwise comparisons remain significant after the ", adj_label, " correction",
if (n_sig_pairs > 0) paste0(", the largest being ", top_pair$comparison,
" (difference ", r2(top_pair$difference), ")") else "",
"."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "condition_means", "subject_profiles",
"anova_table", "sphericity", "posthoc", "methods"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows, rows_removed = rows_removed,
subject_h = subject_h, condition_h = condition_h, value_h = value_h, group_h = group_h,
n_subjects = n_subjects, n_subjects_seen = n_subjects_seen,
n_dropped_subjects = n_dropped_subjects, n_replicate_cells = n_replicate_cells,
n_rows_missing_value = n_rows_missing_value, n_blank_id = n_blank_id,
n_conditions = k, cond_levels = cond_levels,
f_stat = f_stat, df1 = df1, df2 = df2, p_uncorr = p_uncorr,
ss_cond = ss_cond, ss_err = ss_err, subj_res_ss = subj_res_ss, ss_check = ss_check,
partial_eta2 = partial_eta2, generalized_eta2 = generalized_eta2,
mauchly_W = mauchly_W, mauchly_p = mauchly_p,
sphericity_ok = sphericity_ok, sphericity_known = sphericity_known,
sph_note = sph_note, group_note = group_note,
eps_gg = eps_gg, eps_hf = eps_hf, eps_lb = eps_lb, p_gg = p_gg, p_hf = p_hf,
headline_label = headline_label, headline_p = headline_p,
headline_sig = headline_sig, headline_reason = headline_reason, flips = flips,
mixed = mixed, group_levels = group_levels, n_between_groups = n_between_groups,
inter_p = inter_p, inter_peta2 = inter_peta2,
adj_label = adj_label, n_pairs = n_pairs, n_sig_pairs = n_sig_pairs,
top_pair = top_pair, posthoc_df = posthoc_df,
anova_df = anova_df, sphericity_df = sphericity_df, methods_df = methods_df,
condition_means_df = condition_means_df, profiles_df = profiles_df,
highest_cond = highest_cond, lowest_cond = lowest_cond, spread = spread,
pct_same_dir = pct_same_dir, n_profiles_shown = length(pick),
metrics = metrics, json_output = json_output
)
}