Executive Summary
Churn across 1,000 customers
The short answer
Of 1,000 customers, 488 (48.8%) have churned and 512 remain active. The median customer lifetime is 171 days. The strongest churn driver is Plan Type: Pro vs Basic, which reduces churn odds to 0.373 (p = 6.28e-14).
The detail
Churn rate: 488 churned / 1,000 total = 48.8%. Median lifetime (Kaplan-Meier): 171 days. Top driver: Plan Type: Pro vs Basic with odds ratio 0.373 (95% CI: 0.289–0.483, p = 6.28e-14), indicating Pro plan customers have substantially lower churn odds than Basic plan customers. This is the only driver with a confidence interval that does not cross 1.
What this can't tell you
The odds ratio quantifies association; it does not establish whether the plan type itself causes lower churn or whether plan choice reflects unobserved customer commitment or quality differences.
Analysis Overview
Churn rate, survival, and drivers across 1,000 customers.
The short answer
This analysis tracks 1,000 customers from signup through churn or the reference date (2026-06-29), correctly accounting for the fact that 512 still-active customers have not finished their lifetimes yet. The Kaplan-Meier survival estimator counts them properly rather than treating them as failed or dropped, and logistic regression ranks drivers by their association with churn — not their causal effect.
The detail
Of 1,000 observations, 488 churned and 512 remain active as of 2026-06-29. Churn is dated by 'Cancel Date'; still-active customers are censored observations whose true lifetime is known only to exceed the observed duration. The Kaplan-Meier estimator adjusts for this censoring so that survival probabilities reflect the true risk set at each time point. Churn drivers are ranked by odds ratios from logistic regression fit jointly on all predictors; numeric drivers are standardized (per SD), and categorical levels are lumped or blanked as needed. Odds ratios measure association, not causation.
What this can't tell you
This design cannot establish causal effects of any driver on churn. Observed associations may reflect confounding, selection, or reverse causality. The analysis assumes no unobserved confounding and that censoring is independent of the outcome.
Data Quality
Date parsing, churn coercion, and exclusions.
The short answer
All 1,000 rows were retained with no missing dates or data quality issues. The 'Cancel Date' column was read as a churn indicator; all driver columns were usable and required no exclusion for sparsity or constant values.
The detail
Initial rows: 1,000; final rows: 1,000; rows removed: 0. 'Cancel Date' was interpreted as a churn/cancel date field. All mapped driver columns were retained. Numeric drivers are standardized (odds ratios report per standard deviation change); categorical drivers lump rare levels into 'Other' and blank values into 'Missing'. No constant, empty, or identifier-like columns were excluded.
What this can't tell you
Data quality was clean for the fields used; this does not confirm that the data collection process captured all relevant churn drivers or that the 'Cancel Date' field was populated consistently across all churned customers.
Churn by Signup Cohort
Churn rate per month of Signup Date.
The short answer
Churn rates have declined sharply from earlier to later cohorts, falling from 88.3% for customers who signed up in July 2025 to 6.4% for those in June 2026. However, recent cohorts have had less calendar time to churn, so direct comparison across cohorts of different ages overstates improvement.
The detail
Twelve signup cohorts span July 2025 to June 2026. The earliest cohort (2025-07-01) shows 88.3% churn (53 of 60 customers), while the most recent (2026-06-01) shows 6.4%. The decline is not monotonic—the 2025-11-01 cohort (67.1%) is higher than 2025-10-01 (61.6%)—but the overall trajectory from older to newer cohorts is downward. The card's own text notes the critical caveat: recent cohorts have had less time at risk, so their naturally lower churn rates do not necessarily indicate improved retention without adjusting for cohort age.
What this can't tell you
Comparing cohorts of different ages directly conflates two effects: changes in the product or business and the simple passage of time. To assess whether retention has genuinely improved, cohorts must be matched by age (e.g., all cohorts at 30 days post-signup) rather than by calendar date.
Customer Survival Curve
Share of customers still active by lifetime (Kaplan-Meier).
The short answer
The survival curve shows steep attrition in the first month: 69% of customers survive past 90 days and only 17.6% past a year. The curve crosses 50% at 171 days, marking the median customer lifetime. Churn is concentrated early, with the sharpest drops in the first 30 days.
The detail
At 0 days, 99.5% of customers are active; by day 30, survival falls to 86.85%. By day 90, 69% remain active. The curve continues to decline, reaching 17.6% survival at 365 days. The median (50th percentile) occurs at 171 days. The Kaplan-Meier estimator accounts for the 512 still-active customers as censored observations, ensuring that survival probabilities reflect the true risk set rather than treating active customers as failures or drops.
What this can't tell you
The curve does not identify which customer characteristics or behaviors drive the early attrition. It shows the aggregate pattern but not whether churn risk is uniform across all customers or concentrated in specific segments.
Churn Drivers
Odds ratios per driver, with 95% confidence intervals.
The short answer
Plan type dominates the churn picture: Pro customers have 0.373 times the churn odds of Basic customers, a strong protective effect. Five other drivers (region variants, spend, support tickets) show odds ratios near 1 with confidence intervals crossing 1, meaning no statistically significant association with churn.
The detail
Logistic regression on six drivers yields one significant effect: Plan Type: Pro vs Basic (odds ratio 0.373, 95% CI 0.289–0.483, p < 0.05). The remaining five terms—Region: East vs South (0.925), Monthly Spend per SD (0.933), Support Tickets per SD (0.969), Region: North vs South (0.99), and Region: West vs South (1.047)—all have 95% confidence intervals that cross 1.0, indicating no statistically distinguishable effect. Only 1 of 6 terms reaches significance at p < 0.05.
What this can't tell you
These are odds ratios from association, not causation. The strong Pro plan effect could reflect selection (higher-intent customers choose Pro) rather than a causal impact of the plan itself. The dominance of one driver and weakness of others does not rule out unmeasured factors or interactions among the measured ones.
Lifetime Distribution — Churned vs Active
Observed lifetimes split by customer status.
The short answer
Churned customers lasted a median of 74 days before leaving, while still-active customers have already been around a median of 112 days—38 days longer. Active customers already outlast the typical churned lifetime, consistent with churn concentrating in the first few months.
The detail
Box plots of observed lifetimes split by status show: Churned customers have a median tenure of 74 days; Active customers have a median tenure of 112 days. The 38-day difference in medians indicates that the typical active customer has already survived longer than the typical churned customer lasted. This pattern is consistent with early-stage churn concentration—customers who survive the first 3 months are substantially more likely to remain active. The Active group's lifetimes are right-censored (still growing), so their median will increase as these customers continue to accumulate tenure.
What this can't tell you
Right-censoring means the true median for active customers is unknown and will shift upward. The comparison does not adjust for the fact that churned and active cohorts may differ in signup timing, plan mix, or other baseline characteristics. A formal survival analysis (Kaplan-Meier, as shown in the survival_curve card) accounts for censoring; this raw distribution is descriptive only.
Cohort Detail
Per-month cohort sizes and churn.
| Cohort Month | Customers | Churned | Churn Rate PCT |
|---|---|---|---|
| 2025-07-01 | 60 | 53 | 88.3 |
| 2025-08-01 | 106 | 78 | 73.6 |
| 2025-09-01 | 89 | 61 | 68.5 |
| 2025-10-01 | 73 | 45 | 61.6 |
| 2025-11-01 | 76 | 51 | 67.1 |
| 2025-12-01 | 90 | 53 | 58.9 |
| 2026-01-01 | 86 | 46 | 53.5 |
| 2026-02-01 | 78 | 26 | 33.3 |
| 2026-03-01 | 73 | 23 | 31.5 |
| 2026-04-01 | 95 | 29 | 30.5 |
| 2026-05-01 | 96 | 18 | 18.8 |
| 2026-06-01 | 78 | 5 | 6.4 |
The short answer
The largest cohort is 2025-08-01 with 106 customers, of which 78 (73.6%) churned. The oldest cohort, 2025-07-01 (60 customers), shows the highest churn rate at 88.3%. The newest cohort, 2026-06-01 (78 customers), shows the lowest at 6.4%, but has had only weeks at risk.
The detail
Twelve cohorts range from 60 to 106 customers. The 2025-07-01 cohort (60 customers, 53 churned, 88.3% rate) and 2025-08-01 cohort (106 customers, 78 churned, 73.6% rate) are the oldest and largest, respectively, and show the highest churn rates. Churn rates decline as cohorts age into 2026: 2026-01-01 (86 customers, 46 churned, 53.5%), 2026-02-01 (78 customers, 26 churned, 33.3%), and 2026-06-01 (78 customers, 5 churned, 6.4%). The recent cohorts' low rates reflect limited elapsed time, not necessarily improved product-market fit; older cohorts have had 11–12 months to churn.
What this can't tell you
Cohort size varies (60–106 customers), which affects the precision of each rate estimate; smaller cohorts like 2025-07-01 (60 customers) have wider sampling variability. The declining rates from 2025-07 to 2026-06 reflect both customer age and calendar time, making it difficult to isolate which drives the trend without controlling for cohort age at the time of observation.
Methodology
Statistical methodology and diagnostics for Churn Analysis — rate, survival, drivers
Statistical Method
Standard-library analysis: how many customers churn, when they churn, and what predicts it. Map a start date and a churn indicator (a cancel-date column, a 0/1 flag, or yes/no text) and get the overall churn rate, churn by monthly signup cohort, a Kaplan-Meier survival curve of customer lifetime that counts still-active customers correctly, and — if you map candidate driver columns like plan or region — a logistic-regression ranking of churn drivers with odds ratios and confidence intervals.
- One row per customer, with a readable start date
- The churn column is a cancel/churn date (blank = active), a 0/1 flag, or yes/no text
- The last date observed in the file is a reasonable reference date for who is still active
- Without a churn date (flag or yes/no data), churned lifetimes are measured to the reference date and overstate time-to-churn
- Odds ratios measure association, not causation — a driver may proxy for something unmeasured
- Recent signup cohorts have had less time at risk, so their churn rates read low
- The median lifetime is not estimable when fewer than half the observed lifetimes end in churn
Analysis Code
Complete R source code for this analysis
Churn Analysis — Rate, Survival, Drivers
Customer-level churn analysis: overall churn rate, monthly signup-cohort churn, a Kaplan-Meier survival curve of customer lifetime (censoring-aware), and a logistic-regression driver analysis ranking which mapped columns most raise or lower the odds of churning.
Why This Method?
A raw churn rate hides when customers leave and who leaves. Kaplan-Meier survival handles the customers who have not churned yet (censoring) instead of ignoring them, and logistic regression puts every candidate driver on the same odds-ratio scale so they can be ranked honestly.
What This Analysis Covers
- Overall churn rate and median customer lifetime (KM)
- Churn rate by monthly signup cohort
- Survival curve with 95% confidence band
- Churn drivers ranked by odds ratio with confidence intervals
- Lifetime distribution, churned vs active
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {start_date, churn_status, base_1..base_N}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(survival))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
Step 1: Required columns + humanized names
initial_rows <- nrow(df)
if (!"start_date" %in% names(df) || !"churn_status" %in% names(df)) {
stop("column_mapping must map both start_date and churn_status")
}
start_name <- humanize_semantic("start_date", col_map)
churn_name <- humanize_semantic("churn_status", col_map)
base_cols <- grep("^base_[0-9]+$", names(df), value = TRUE)
base_cols <- base_cols[order(as.integer(sub("^base_", "", base_cols)))]
driver_names <- setNames(humanize_semantic(base_cols, col_map), base_cols)Step 2: Parse start dates (multi-format, lubridate)
start_date <- parse_dates_flex(df$start_date)
if (sum(!is.na(start_date)) < 0.5 * initial_rows) {
stop(paste0("Could not read '", start_name, "' as dates — fewer than half ",
"of the values parse in any common date format."))
}Step 3: Coerce churn_status — churn/cancel DATE, 0/1 flag, or yes/no text
raw_churn <- df$churn_status
raw_chr <- trimws(as.character(raw_churn))
is_blank <- is.na(raw_chr) | raw_chr == ""
n_nonblank <- sum(!is_blank)
churned <- rep(NA, initial_rows)
churn_date <- as.Date(rep(NA, initial_rows))
churn_mode <- NULL
cd_try <- parse_dates_flex(raw_churn)
if (n_nonblank > 0 && sum(!is.na(cd_try)) >= 0.8 * n_nonblank) {Date mode: non-blank = churn/cancel date, blank = still active
churn_mode <- "date"
churned <- !is.na(cd_try)
churn_date <- cd_try
} else {
num_try <- suppressWarnings(as.numeric(raw_chr))
num_ok <- !is.na(num_try)
if (n_nonblank > 0 && sum(num_ok) >= 0.95 * n_nonblank &&
all(num_try[num_ok] %in% c(0, 1))) {Flag mode: 1 = churned, 0 = active (blank treated as unknown -> dropped)
churn_mode <- "flag"
churned <- ifelse(is_blank, NA, num_try == 1)
} else {Text mode: yes/churned/cancelled vs no/active
lo <- tolower(raw_chr)
yes_set <- c("yes", "y", "true", "churned", "churn", "cancelled",
"canceled", "cancel", "inactive", "lost", "closed")
no_set <- c("no", "n", "false", "active", "current", "retained",
"open", "subscribed")
mapped <- ifelse(lo %in% yes_set, TRUE, ifelse(lo %in% no_set, FALSE, NA))
if (n_nonblank > 0 && sum(!is.na(mapped[!is_blank])) >= 0.8 * n_nonblank) {
churn_mode <- "text"
churned <- mapped
} else {
stop(paste0("Could not read '", churn_name, "' as a churn indicator — ",
"expected a churn/cancel date column(blank = active), a 0/1 ",
"flag, or yes/no text."))
}
}
}Step 4: Reference date, tenure, row filtering
ref_date <- suppressWarnings(max(c(start_date, churn_date), na.rm = TRUE))
if (!is.finite(as.numeric(ref_date))) {
stop(paste0("No usable dates found in '", start_name, "'."))
}
keep <- !is.na(start_date) & !is.na(churned)Churned rows in date mode need churn_date >= start_date; others need start <= ref
tenure <- ifelse(churned & !is.na(churn_date),
as.numeric(churn_date - start_date),
as.numeric(ref_date - start_date))
keep <- keep & !is.na(tenure) & tenure >= 0
work <- data.frame(start_date = start_date, churned = churned,
tenure = tenure, stringsAsFactors = FALSE)[keep, , drop = FALSE]
for (bc in base_cols) work[[bc]] <- df[[bc]][keep]
final_rows <- nrow(work)
rows_removed <- initial_rows - final_rowsStep 5: Minimum-data guards (clear, humanized messages)
if (final_rows < 30) {
stop(paste0("Only ", final_rows, " usable rows after reading '", start_name,
"' and '", churn_name, "' — churn analysis needs at least 30 ",
"customers."))
}
n_churned <- sum(work$churned)
n_active <- final_rows - n_churned
if (n_churned < 5) {
stop(paste0("Only ", n_churned, " churned customer(s) found in '", churn_name,
"' — at least 5 are needed to analyse churn."))
}
churn_rate <- n_churned / final_rowsStep 6: Signup cohorts (month; lumped to quarter/year if > 24 cohorts)
cohort_unit <- "month"
coh <- lubridate::floor_date(work$start_date, "month")
if (length(unique(coh)) > 24) { cohort_unit <- "quarter"; coh <- lubridate::floor_date(work$start_date, "quarter") }
if (length(unique(coh)) > 24) { cohort_unit <- "year"; coh <- lubridate::floor_date(work$start_date, "year") }
cohort_df <- work %>%
mutate(cohort = coh) %>%
group_by(cohort) %>%
summarise(customers = n(), churned = sum(churned), .groups = "drop") %>%
arrange(cohort) %>%
mutate(churn_rate_pct = round(100 * churned / customers, 1),
cohort_month = as.character(cohort)) %>%
select(cohort_month, customers, churned, churn_rate_pct) %>%
as.data.frame(stringsAsFactors = FALSE)Step 7: Kaplan-Meier survival of customer lifetime (censoring-aware)
km_fit <- survival::survfit(survival::Surv(work$tenure, work$churned) ~ 1,
conf.int = 0.95)
km_full <- data.frame(
tenure_days = km_fit$time,
survival_pct = round(100 * km_fit$surv, 2),
lower_pct = round(100 * ifelse(is.na(km_fit$lower), km_fit$surv, km_fit$lower), 2),
upper_pct = round(100 * ifelse(is.na(km_fit$upper), km_fit$surv, km_fit$upper), 2),
stringsAsFactors = FALSE
)
km_full <- km_full[order(km_full$tenure_days), , drop = FALSE]
if (nrow(km_full) > 500) {
idx <- unique(round(seq(1, nrow(km_full), length.out = 500)))
km_df <- km_full[idx, , drop = FALSE]
} else {
km_df <- km_full
}
rownames(km_df) <- NULL
km_table <- summary(km_fit)$table
median_lifetime <- suppressWarnings(as.numeric(km_table["median"]))
if (length(median_lifetime) == 0) median_lifetime <- NA_real_Step 8: Churn drivers — logistic regression on the mapped base_* columns
drivers_df <- NULL
top_driver <- NULL
dropped_drivers <- character(0)
drivers_skipped_reason <- NULL
usable <- character(0)
model_data <- data.frame(churned = as.integer(work$churned))
term_source <- list() # model column -> list(base, kind)
if (length(base_cols) == 0) {
drivers_skipped_reason <- "no driver columns were mapped"
} else if (n_active == 0 || n_churned == 0) {
drivers_skipped_reason <- paste0(
"driver analysis needs both churned and active customers; this data has ",
if (n_active == 0) "no active(retained) customers" else "no churned customers")
} else {
for (bc in base_cols) {
v <- work[[bc]]
hn <- driver_names[[bc]]
v_chr <- trimws(as.character(v))
nonblank <- !is.na(v_chr) & v_chr != ""
if (sum(nonblank) == 0) {
dropped_drivers <- c(dropped_drivers, paste0(hn, " (empty)")); next
}
conv <- suppressWarnings(as.numeric(v_chr))
if (sum(!is.na(conv[nonblank])) >= 0.95 * sum(nonblank)) {Numeric driver (95% rule): impute median, standardize -> OR per SD
conv[is.na(conv)] <- median(conv, na.rm = TRUE)
sdv <- stats::sd(conv)
if (is.na(sdv) || sdv == 0) {
dropped_drivers <- c(dropped_drivers, paste0(hn, " (constant)")); next
}
model_data[[bc]] <- as.numeric(scale(conv))
term_source[[bc]] <- list(base = bc, kind = "numeric")
usable <- c(usable, bc)
} else {Categorical driver: blanks -> Missing, identifier-like dropped, lump to <=8 levels
lv <- ifelse(nonblank, v_chr, "Missing")
n_lev <- length(unique(lv))
if (n_lev > 0.5 * final_rows || n_lev > 50) {
dropped_drivers <- c(dropped_drivers, paste0(hn, " (identifier-like)")); next
}
if (n_lev < 2) {
dropped_drivers <- c(dropped_drivers, paste0(hn, " (constant)")); next
}
tab <- sort(table(lv), decreasing = TRUE)
if (n_lev > 8) {
keep_lv <- names(tab)[1:8]
lv <- ifelse(lv %in% keep_lv, lv, "Other")
tab <- sort(table(lv), decreasing = TRUE)
}
model_data[[bc]] <- stats::relevel(factor(lv), ref = names(tab)[1])
term_source[[bc]] <- list(base = bc, kind = "categorical",
ref = names(tab)[1])
usable <- c(usable, bc)
}
}
if (length(usable) == 0) {
drivers_skipped_reason <-
"none of the mapped driver columns were usable(constant, empty, or identifier-like)"
} else {
fit <- tryCatch(
stats::glm(churned ~ ., data = model_data, family = stats::binomial()),
error = function(e) NULL, warning = function(w) {
suppressWarnings(stats::glm(churned ~ ., data = model_data,
family = stats::binomial()))
})
if (is.null(fit)) {
drivers_skipped_reason <- "the churn-driver model could not be fitted on this data"
} else {
cf <- summary(fit)$coefficients
cf <- cf[rownames(cf) != "(Intercept)", , drop = FALSE]
rows <- list()
for (tm in rownames(cf)) {
est <- cf[tm, "Estimate"]; se <- cf[tm, "Std. Error"]
pv <- cf[tm, "Pr(>|z|)"]
if (is.na(est) || is.na(se) || abs(est) > 10) next # separation / aliased
src <- NULL
for (bc in usable) if (startsWith(tm, bc)) {
if (is.null(src) || nchar(bc) > nchar(src$base)) src <- term_source[[bc]]
}
if (is.null(src)) next
hn <- driver_names[[src$base]]
label <- if (src$kind == "numeric") {
paste0(hn, " (per SD)")
} else {
paste0(hn, ": ", sub(paste0("^", src$base), "", tm),
" vs ", src$ref)
}
rows[[length(rows) + 1]] <- data.frame(
driver = label,
odds_ratio = round(exp(est), 3),
or_lower = round(exp(est - 1.96 * se), 3),
or_upper = round(exp(est + 1.96 * se), 3),
p_value = signif(pv, 3),
log_or = est,
stringsAsFactors = FALSE
)
}
if (length(rows) > 0) {
dd <- do.call(rbind, rows)
dd$effect <- ifelse(dd$odds_ratio >= 1, "increases churn", "decreases churn")
dd$significance <- ifelse(is.na(dd$p_value), "",
ifelse(dd$p_value < 0.001, "***",
ifelse(dd$p_value < 0.01, "**",
ifelse(dd$p_value < 0.05, "*", ""))))
dd <- dd[order(-abs(dd$log_or)), , drop = FALSE]
rownames(dd) <- NULLTop driver: NA-safe — prefer significant terms, else strongest effect
ok_idx <- which(!is.na(dd$log_or))
sig_idx <- which(!is.na(dd$p_value) & dd$p_value < 0.05)
pick <- if (length(sig_idx) > 0) sig_idx[1] else if (length(ok_idx) > 0) ok_idx[1] else NA
if (!is.na(pick)) {
top_driver <- list(label = dd$driver[pick],
odds_ratio = dd$odds_ratio[pick],
p_value = dd$p_value[pick],
effect = dd$effect[pick])
}
drivers_df <- head(dd[, c("driver", "odds_ratio", "or_lower",
"or_upper", "p_value", "effect",
"significance")], 12)
} else {
drivers_skipped_reason <-
"no driver term produced a stable estimate(separation or aliasing)"
}
}
}
}Step 9: Lifetime distribution by status (box)
set.seed(42)
tidx <- if (final_rows > 1000) sample(final_rows, 1000) else seq_len(final_rows)
tenure_df <- data.frame(
status = ifelse(work$churned[tidx], "Churned", "Active"),
tenure_days = round(work$tenure[tidx], 1),
stringsAsFactors = FALSE
)Step 10: Metrics + json answer (all computed)
med_txt <- if (is.na(median_lifetime)) "not reached in the observed window"
else paste0(round(median_lifetime), " days")
metrics <- list(
`Customers` = final_rows,
`Churned` = as.integer(n_churned),
`Churn Rate %` = round(100 * churn_rate, 1),
`Median Lifetime(days)` = if (is.na(median_lifetime)) NA else round(median_lifetime),
`Top Churn Driver` = if (!is.null(top_driver)) top_driver$label else "n/a"
)
json_output <- list(
answer = paste0(
"Churn analysis of ", format(final_rows, big.mark = ","), " customers(from '",
start_name, "' and '", churn_name, "'): ", round(100 * churn_rate, 1),
"% have churned(", format(n_churned, big.mark = ","), " churned, ",
format(n_active, big.mark = ","), " still active as of ",
as.character(ref_date), "). Median customer lifetime(Kaplan-Meier, ",
"censoring-aware) is ", med_txt, ". ",
if (!is.null(top_driver)) paste0(
"The strongest churn driver is ", top_driver$label, " (odds ratio ",
top_driver$odds_ratio, ", ", top_driver$effect,
if (!is.na(top_driver$p_value)) paste0(", p = ", top_driver$p_value) else "",
").")
else paste0("Driver analysis was skipped: ",
drivers_skipped_reason %||% "no usable drivers", ".")
),
cards = lapply(
c("tldr", "overview", "preprocessing", "churn_by_cohort", "survival_curve",
"churn_drivers", "tenure_distribution", "cohort_table"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed,
start_name = start_name, churn_name = churn_name,
churn_mode = churn_mode, ref_date = ref_date,
n_churned = as.integer(n_churned), n_active = as.integer(n_active),
churn_rate = churn_rate, median_lifetime = median_lifetime,
cohort_df = cohort_df, cohort_unit = cohort_unit,
km_df = km_df, drivers_df = drivers_df, top_driver = top_driver,
driver_names = driver_names, dropped_drivers = dropped_drivers,
drivers_skipped_reason = drivers_skipped_reason,
tenure_df = tenure_df,
metrics = metrics, json_output = json_output
)
}NA-safe landmark reads off the curve
s_at <- function(d) {
idx <- which(km$tenure_days <= d)
if (length(idx) == 0) return(NA_real_)
km$survival_pct[max(idx)]
}
s90 <- s_at(90); s365 <- s_at(365)
med_txt <- if (is.na(shared$median_lifetime))
"The curve never crosses 50%, so the median lifetime is not reached within the observed window."
else paste0("The curve crosses 50% at ", round(shared$median_lifetime),
" days — the median customer lifetime.")
text <- paste0(
"The curve shows the share of customers still active after each number of ",
"days since '", shared$start_name, "', using the Kaplan-Meier estimator so ",
"that still-active(censored) customers are counted correctly rather than ",
"treated as churned or dropped. ",
if (!is.na(s90)) paste0(round(s90, 1), "% of customers survive past 90 days",
if (!is.na(s365)) paste0(" and ", round(s365, 1), "% past a year") else "",
". ") else "",
med_txt
)
km <- km[, c("tenure_days", "survival_pct")]
list(
title = "Customer Survival Curve",
description = "Share of customers still active by lifetime(Kaplan-Meier).",
text = text,
chart_labels = list(
tenure_days = "Customer lifetime(days)",
survival_pct = "Still active(%)"
),
data = list(survival_curve = km)
)
}
# Card: churn_drivers (horizontal_bar)
card_churn_drivers <- function(shared, df, params) {
if (is.null(shared$drivers_df)) {
return(list(
title = "Churn Drivers",
description = "No usable driver columns.",
text = paste0("Driver analysis was skipped: ",
shared$drivers_skipped_reason %||% "no usable drivers",
". Map columns such as plan, price tier, region, or usage ",
"metrics as driver columns to rank what predicts churn."),
data = list()
))
}
dd <- shared$drivers_df
n_sig <- sum(dd$p_value < 0.05, na.rm = TRUE)
risk <- dd[dd$odds_ratio >= 1, , drop = FALSE]
prot <- dd[dd$odds_ratio < 1, , drop = FALSE]
text <- paste0(
"Each bar is a driver's churn odds ratio from a logistic regression fit on ",
"all drivers jointly: above 1 means higher churn odds, below 1 lower; the ",
"whiskers are 95% confidence intervals(an interval crossing 1 means the ",
"effect is not statistically distinguishable from none). ",
shared$top_driver$label, " has the strongest effect(odds ratio ",
shared$top_driver$odds_ratio, ", ", shared$top_driver$effect, "). ",
n_sig, " of ", nrow(dd), " terms are significant at p < 0.05.",
if (nrow(risk) > 0 && nrow(prot) > 0) paste0(
" Risk factors: ", paste(head(risk$driver, 3), collapse = "; "),
". Protective: ", paste(head(prot$driver, 3), collapse = "; "), ".") else "",
" Numeric drivers are per standard deviation; odds ratios measure ",
"association, not causation."
)
list(
title = "Churn Drivers",
description = "Odds ratios per driver, with 95% confidence intervals.",
text = text,
chart_labels = list(
driver = "Driver",
odds_ratio = "Churn odds ratio(1 = no effect)"
),
data = list(churn_drivers = dd[, c("driver", "odds_ratio",
"or_lower", "or_upper")])
)
}
# Card: tenure_distribution (box)
card_tenure_distribution <- function(shared, df, params) {
td <- shared$tenure_df
med_ch <- suppressWarnings(median(td$tenure_days[td$status == "Churned"]))
med_ac <- suppressWarnings(median(td$tenure_days[td$status == "Active"]))
both <- length(unique(td$status)) == 2
text <- paste0(
"Each box shows the spread of customer lifetimes(days since '",
shared$start_name, "'). ",
if (both) paste0(
"Churned customers lasted a median of ", round(med_ch),
" days before leaving, while still-active customers have already been ",
"around a median of ", round(med_ac), " days",
if (!is.na(med_ch) && !is.na(med_ac) && med_ac > med_ch) paste0(
" — active customers already outlast the typical churned lifetime, ",
"consistent with churn concentrating early") else "", ". ")
else paste0("All customers in this data share the status '",
unique(td$status), "'. "),
"Note the Active boxes are right-censored: those lifetimes are still growing."
)
list(
title = "Lifetime Distribution — Churned vs Active",
description = "Observed lifetimes split by customer status.",
text = text,
chart_labels = list(
status = "Customer status",
tenure_days = "Lifetime(days)"
),
data = list(tenure_by_status = td)
)
}
# Card: cohort_table (table)
card_cohort_table <- function(shared, df, params) {
cd <- shared$cohort_df
biggest <- cd[which.max(cd$customers), ]
text <- paste0(
"One row per ", shared$cohort_unit, " signup cohort: how many customers ",
"started, how many have churned, and the cohort's churn rate. The largest ",
"cohort is ", biggest$cohort_month, " with ",
format(biggest$customers, big.mark = ","), " customers(",
biggest$churn_rate_pct, "% churned). Older cohorts have had more time at ",
"risk, so read the rates alongside cohort age."
)
list(
title = "Cohort Detail",
description = paste0("Per-", shared$cohort_unit, " cohort sizes and churn."),
text = text,
data = list(cohort_table = cd)
)
}