Standard Weibull
Executive Summary

Executive Summary

Weibull life verdict across 500 units.

Units
500
Failures
318
Censored %
36.4
Shape (beta)
2.419
Scale (eta)
1012.57
B10 Life
399.33
B50 Life
870.19
Mean Life (MTBF)
897.76
Failure Mode
wear-out (failure rate rising with age)
The shape parameter is 2.42 (95% CI 2.23 to 2.63), entirely above 1, so the fitted failure rate rises as units age — the wear-out pattern, in which older units are progressively more likely to fail. Age-based preventive replacement acts on this pattern; it does not help when the rate is flat. Characteristic life is 1,013 Operating Hours; 10% of units are expected to have failed by 399 (B10) and half by 870 (B50), with a mean life of 898. 36.4% of units were still running when the data ended, and the fit counts them for the time they survived rather than discarding them. The data itself reaches 1,820 Operating Hours, where 1.6% are still fitted to be running; the curve is projected to 2,274 (1.25 times the observed follow-up), and that projected part is the model's assumption rather than anything the data has seen.
Suggested Interpretation

The fitted failure rate rises as units age—a wear-out pattern—because the shape parameter is 2.42 (95% CI 2.23 to 2.63), entirely above 1. This means older bearings are progressively more likely to fail, making age-based preventive replacement a sound strategy. The B10 life is 399.33 Operating Hours (10% of units expected to fail by then), the B50 is 870.19 (half expected to fail), and the mean life (MTBF) is 897.76. The characteristic life is 1,012.57 Operating Hours. The data reaches 1,820 Operating Hours; the curve is projected to 2,274 (1.25 times observed follow-up), and that projection is the model's assumption, not observed fact.

Overview

Analysis Overview

Weibull life fit on 500 units with 318 observed failures.

N Units500
N Failures318
Censored Pct36.4
Shape Beta2.419
Suggested Interpretation

The short answer

These bearings follow a Weibull failure pattern with shape parameter 2.419, indicating an increasing hazard rate—failures grow more common as units age. The analysis fit this pattern across 500 units with 318 observed failures; the remaining 36.4% were still running when observation ended, and including their full operating time (rather than dropping them) is critical to avoid understating bearing life.

The detail

The Weibull fit used maximum likelihood estimation on 500 units, of which 318 failed and 182 remained right-censored. The shape parameter beta = 2.419 sits above 1, confirming an increasing failure rate: older bearings fail at higher rates than younger ones. The censored units (36.4% of the file) are not missing data—each contributes its full observed operating hours to the fit, preventing the downward bias that would result from analyzing only the 318 failures. The parametric Weibull route yields a life equation and B10 life (the hours at which 10% of units are expected to have failed), readable directly from eta and beta in the standard form R(t) = exp(-(t/eta)^beta).

What this can't tell you

The analysis assumes a single Weibull population; if distinct failure modes or bearing types are present, they would violate this assumption and require stratification or a mixture model. B10 life itself is not reported in this overview—it appears in a separate deliverable derived from eta and beta.

Data Preparation

Data Quality

Row exclusions, how the failure flag was read, and the censoring rate.

Initial Rows500
Final Rows500
Rows Removed0
Censored Pct36.4
Suggested Interpretation

All 500 rows were usable; no rows were removed for missing or non-positive durations or an unreadable failure flag. In the Failed column, the value "1" was treated as a failure; all other values count as units still running. The censoring rate is 36.4%—182 of 500 units had not failed when observation stopped at 1,820 Operating Hours. The last observed failure occurred at 1,820 Operating Hours. These 182 censored units contribute the fact that they survived at least their recorded operating hours, which prevents the life estimates from being biased downward. The data quality is clean: no exclusions, no ambiguity in the failure flag reading, and complete follow-up to the stopping point.

Data Table

Weibull Parameters & Failure Mode

Shape and scale with 95% confidence intervals, plus fit quality.

ParameterEstimateCI LowCI HighInterpretation
Shape (beta)2.4192.2252.629Below 1 means the failure rate falls with age, above 1 means it rises, 1 means it is flat. Here: wear-out (failure rate rising with age). 318 observed failures carry the shape estimate; the interval below is the honest range.
Scale (eta) — characteristic life1013967.61060The Operating Hours by which 63.2% of units have failed — the natural scale of this life distribution, not a safe operating limit.
Observed failures carrying the fit318Censored units (182, 36.4% of the file) contribute the time they survived, but only failures pin down the shape.
Probability-plot R-squared0.9948How straight the data is in Weibull coordinates: very close to a straight line. This scores the fit, and a good fit is still a model choice, not a fact.
Largest gap to the observed curve (percentage points)2.74The largest distance between the fitted reliability curve and the assumption-free Kaplan-Meier curve computed from the same data.
Suggested Interpretation

The shape parameter beta is 2.4186 (95% CI 2.2251 to 2.6289), entirely above 1, confirming wear-out: the failure rate rises with age and age-based preventive replacement is operationally sound. The scale (characteristic life) eta is 1,012.572 Operating Hours (95% CI 967.571 to 1,059.665)—the point at which 63.2% of units have failed. The 318 observed failures carry the shape estimate; the interval reflects the honest range given that sample size. Fit quality is strong: the probability-plot R-squared is 0.9948, indicating the data is very close to a straight line in Weibull coordinates. The largest gap to the assumption-free Kaplan-Meier curve is 2.74 percentage points, showing the Weibull assumption imposes minimal distortion over the observed range. A Weibull fit is a model of this data, not a property of the parts.

Visualization

Reliability Curve R(t)

Share of units still running against Operating Hours, fitted and projected, with the observed curve for comparison.

Suggested Interpretation

The fitted reliability curve shows the share of units still running at any age. At 1,820 Operating Hours (the end of observed data), 1.6% are still running; the curve is then extended to 2,274 Operating Hours, where 0.1% remain fitted to be running. The stepped Kaplan-Meier curve computed from the same data without assuming any distribution tracks the smooth Weibull fit closely, meaning the Weibull assumption does no harm over the observed range. Where they diverge, the smooth curve imposes a shape the data does not show. Everything past 1,820 Operating Hours is a separate series for a reason: it is the model extended 1.25 times beyond the longest unit ever observed, and it assumes the same failure process continues unchanged. No unit in the file can confirm that.

Data Table

Life Metrics

B10, B50, characteristic life, MTBF, and reliability at two named times.

MetricEstimateCI LowCI HighInterpretation
B10 life (10% of units failed)399.3364.6437.4The design life engineers usually quote: 10% of units are expected to have failed by 399 Operating Hours.
B50 life (half of units failed)870.2829.8912.5Half of all units are expected to have failed by 870 Operating Hours.
Characteristic life (63.2% failed)1013967.61060The Weibull scale parameter, always the 63.2% point whatever the shape.
Mean life (MTBF)897.8857.8939.6The average life implied by the fit (eta times the gamma function of 1 + 1/beta). It is the mean, not a survival guarantee: 52.6% of units are expected to have failed by then.
Reliability at the end of the data (1,820)1.61The share still running at the last Operating Hours actually covered by this data — the furthest point the file can speak to.
Reliability projected to 2,2740.08The fitted survival share at 2,274 Operating Hours. Beyond 1,820 Operating Hours there is no data — this number is the fitted model extended forward, not an observation.
Suggested Interpretation

B10 life—the design number reliability engineers quote—is 399.334 Operating Hours (95% CI 364.617 to 437.358), meaning 10% of units are expected to have failed by then. The B50 life is 870.187 Operating Hours (95% CI 829.819 to 912.519), at which half are expected to have failed. The characteristic life is 1,012.572 Operating Hours (95% CI 967.571 to 1,059.665). The mean life (MTBF) is 897.759 Operating Hours (95% CI 857.798 to 939.581), but this is the average, not a safe operating point: 52.6% of units are expected to have failed by then. Reliability at the end of observed data (1,820 Operating Hours) is 1.61%, the furthest point this file can speak to. Reliability projected to 2,274 Operating Hours is 0.08%, a model assumption beyond any observed data.

Visualization

Hazard Curve h(t)

The instantaneous failure rate against Operating Hours.

Suggested Interpretation

The failure rate per Operating Hour rises steadily with age, from 0.0000435 at 60.1 Operating Hours to 0.00547 at 1,815. This rising hazard is the operational signature of wear-out and is why the shape parameter matters: a rising hazard makes age-based preventive replacement pay, a falling hazard makes burn-in pay, and a flat hazard means age tells you nothing about risk. Here the reading is clear wear-out (failure rate rising with age). The hazard formula is h(t) = (beta/eta) × (t/eta)^(beta−1), and because beta is 2.42, the exponent is always positive, driving the upward trend. The curve past 1,820 Operating Hours is drawn separately because no unit was observed that long.

Visualization

Weibull Probability Plot

Failure data in Weibull coordinates, with the fitted line.

Suggested Interpretation

In Weibull coordinates—natural log of Operating Hours on one axis, log of minus log of reliability on the other—a genuine Weibull population falls on a straight line with slope equal to the shape parameter. Here, 316 observed failure points (plotted at Kaplan-Meier positions to account for censoring) score an R-squared of 0.995 against the fitted line, reading as very close to straight. No systematic curvature or break into two segments appears, so no evidence of multiple failure modes averaging into a single shape that describes neither. The plot is the gateway diagnostic: if it curved or split, the single Weibull fit would be averaging distinct failure mechanisms and the shape and life estimates would be untrustworthy. Here the data is consistent with a single wear-out mechanism.

Methodology

Methodology

Statistical methodology and diagnostics for Weibull Failure Analysis

Statistical Method

Weibull Failure Analysis

Standard-library analysis: how do components fail over time? Fits a two-parameter Weibull life distribution to your time-to-failure data with right-censoring handled properly — units still running when observation stopped are counted for exactly as long as they lasted instead of being dropped or treated as failures. You get the shape parameter beta with a confidence interval and the failure-mode reading that comes with it (below 1 = infant mortality, around 1 = random failure, above 1 = wear-out), the characteristic life eta, B10 and B50 life, MTBF, the reliability curve R(t) with a clearly-labelled forward projection, the hazard curve, and a Weibull probability plot with the fitted line so you can see whether the distribution actually fits. Works on any life data: map a duration column and a failure flag.

Data
N = 500 observations
Assumptions
  • Durations are positive numbers on one consistent scale (hours, cycles, miles)
  • All units come from a single Weibull population with one failure mode — the probability plot is the check on this
  • Censoring is non-informative: units still running are not systematically different from those that failed
  • The failure indicator is a two-state flag; anything not marked as a failure is treated as still running
Limitations
  • A fitted distribution is a model of this data, not a property of the parts — the probability-plot R-squared and the gap to the Kaplan-Meier curve are reported so the fit can be judged
  • Reliability projected past the longest observed unit is an assumption, not evidence; the projected segment is drawn as a separate series and its extrapolation factor is stated
  • With few observed failures the shape parameter is poorly identified regardless of unit count — the analysis reports the effective failure count and refuses to name a failure mode when the interval spans 1
  • Mixed failure modes bend the probability plot and a single Weibull fit averages them into a shape that describes neither
Software & Citation
MCP Analytics · mcpanalytics.ai
Code Appendix

Analysis Code

Complete R source code for this analysis

Weibull Failure Analysis — How Components Fail Over Time

Fits a two-parameter Weibull life distribution to time-to-failure data with right-censoring (units still running when observation stopped), and reports the two numbers reliability engineering is built on: the shape parameter beta — whose value separates infant mortality from random failure from wear-out — and the characteristic life eta.

Why This Method?

Reliability data is almost never complete: at the end of a test or a warranty window some units have failed and the rest are still running. Dropping the survivors biases every life estimate downward; treating them as failures biases it too. Maximum-likelihood Weibull fitting with right-censoring uses each survivor for exactly as long as it was observed, which is what makes B10 life, MTBF and the reliability curve honest numbers rather than optimistic ones.

What This Analysis Covers

  • Shape (beta) and scale (eta) with 95% confidence intervals
  • The failure-mode reading of beta, hedged by its confidence interval
  • B10 and B50 life, characteristic life, and mean life (MTBF)
  • The reliability curve R(t) with a clearly-labelled forward projection
  • The hazard curve h(t) — whether the failure rate rises or falls
  • A Weibull probability plot with the fitted line and a fit-quality score

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {time, event}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))
suppressPackageStartupMessages(library(survival))

Core Analysis Pipeline

Step 1: Discover mapped columns (humanized names for ALL prose)

initial_rows <- nrow(df)
  for (k in c("time", "event")) {
    if (!k %in% names(df)) {
      stop(sprintf("column_mapping must map the &#x27;%s' column.", k))
    }
  }
  h_time  <- humanize_semantic("time",  col_map)
  h_event <- humanize_semantic("event", col_map)

Step 2: Coerce the duration numeric (95% rule); drop invalid durations

tv <- df$time
  if (!is.numeric(tv)) {
    conv <- suppressWarnings(as.numeric(as.character(tv)))
    n_orig <- sum(!is.na(tv) & as.character(tv) != "")
    if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
      tv <- conv
    } else {
      stop(sprintf(
        paste0("The life/duration column(%s) could not be read as a number. ",
               "Weibull analysis needs a numeric time, cycle, or distance to failure."),
        h_time))
    }
  }
  bad_time <- is.na(tv) | !is.finite(tv) | tv <= 0
  dropped_time_rows <- sum(bad_time)

Step 3: Binarize the failure indicator (censored = still running)

eb <- .binarize_event(df$event, h_event)
  ev <- eb$ev
  failure_level <- eb$level
  bad_event <- is.na(ev)
  dropped_event_rows <- sum(bad_event & !bad_time)

  keep <- !bad_time & !bad_event
  d <- data.frame(time = tv[keep], event = as.integer(ev[keep]),
                  stringsAsFactors = FALSE)
  final_rows <- nrow(d)
  rows_removed <- initial_rows - final_rows

Step 4: Guard the fit — rows, variation, and observed failures

if (final_rows < 20) {
    stop(sprintf(
      paste0("Only %d usable rows across %s and %s — Weibull life analysis needs ",
             "at least 20 units with a positive duration and a readable failure flag."),
      final_rows, h_time, h_event))
  }
  if (length(unique(d$time)) < 2) {
    stop(sprintf(
      paste0("Every usable value of %s is the same number(%s) — a life ",
             "distribution cannot be fitted to a single duration."),
      h_time, .fmt_sig(d$time[1])))
  }
  n_units <- final_rows
  n_failures <- sum(d$event)
  n_censored <- n_units - n_failures
  censored_pct <- 100 * n_censored / n_units
  if (n_failures < 3) {
    stop(sprintf(
      paste0("Only %d failure(s) are recorded in %s — the Weibull shape and scale ",
             "cannot both be estimated from fewer than 3 observed failures. The ",
             "other %d unit(s) are still running(right-censored)."),
      n_failures, h_event, n_censored))
  }
  if (length(unique(d$time[d$event == 1])) < 2) {
    stop(sprintf(
      paste0("Every observed failure in %s happened at the same %s value — the ",
             "shape parameter cannot be estimated without variation between failures."),
      h_event, h_time))
  }
  few_failures <- n_failures < 20

Step 5: MLE Weibull fit with right-censoring, and the parameterization

survreg fits log(T) = mu + sigma * W (W standard extreme value). The reliability convention is R(t) = exp(-(t/eta)^beta), so shape beta = 1 / sigma scale eta = exp(mu)

fit <- tryCatch(
    suppressWarnings(survreg(Surv(time, event) ~ 1, data = d, dist = "weibull")),
    error = function(e) NULL
  )
  if (is.null(fit) || !is.finite(fit$scale) || fit$scale <= 0) {
    stop(sprintf(
      paste0("The Weibull model could not be fitted to %s and %s — the failure ",
             "times carry too little information to identify a life distribution."),
      h_time, h_event))
  }
  mu <- unname(coef(fit)[1])
  sigma <- as.numeric(fit$scale)
  beta <- 1 / sigma
  eta <- exp(mu)
  if (!is.finite(beta) || !is.finite(eta) || beta <= 0 || eta <= 0) {
    stop(sprintf(
      "The fitted Weibull parameters for %s were not finite — the life data cannot support this model.",
      h_time))
  }

95% intervals. The fit's covariance is over (mu, log sigma); log(beta) = -log(sigma), so beta and eta transform directly, and every life quantile gets a delta-method interval on the log scale.

V <- tryCatch(vcov(fit), error = function(e) NULL)
  ci_available <- !is.null(V) && is.matrix(V) && nrow(V) >= 2 &&
    all(is.finite(V[1:2, 1:2])) && V[1, 1] > 0 && V[2, 2] > 0
  if (ci_available) {
    se_mu <- sqrt(V[1, 1]); se_logsigma <- sqrt(V[2, 2])
    beta_lo <- exp(log(beta) - 1.96 * se_logsigma)
    beta_hi <- exp(log(beta) + 1.96 * se_logsigma)
    eta_lo  <- exp(mu - 1.96 * se_mu)
    eta_hi  <- exp(mu + 1.96 * se_mu)
    beta_p  <- 2 * pnorm(-abs(log(beta) / se_logsigma))
  } else {
    beta_lo <- NA_real_; beta_hi <- NA_real_
    eta_lo  <- NA_real_; eta_hi  <- NA_real_
    beta_p  <- NA_real_
  }

Life quantile t_p = eta * (-log(1-p))^sigma, with a delta-method interval.

.life_q <- function(p) {
    kp <- -log(1 - p)
    lt <- mu + sigma * log(kp)
    est <- exp(lt)
    if (!ci_available) return(c(est = est, lo = NA_real_, hi = NA_real_))
    g <- c(1, sigma * log(kp))
    v <- as.numeric(t(g) %*% V[1:2, 1:2] %*% g)
    se <- sqrt(max(v, 0))
    c(est = est, lo = exp(lt - 1.96 * se), hi = exp(lt + 1.96 * se))
  }

Mean life = eta * Gamma(1 + sigma), same delta-method treatment.

.mean_life <- function() {
    lm_ <- mu + lgamma(1 + sigma)
    est <- exp(lm_)
    if (!ci_available) return(c(est = est, lo = NA_real_, hi = NA_real_))
    g <- c(1, digamma(1 + sigma) * sigma)
    v <- as.numeric(t(g) %*% V[1:2, 1:2] %*% g)
    se <- sqrt(max(v, 0))
    c(est = est, lo = exp(lm_ - 1.96 * se), hi = exp(lm_ + 1.96 * se))
  }
  q10 <- .life_q(0.10); q50 <- .life_q(0.50); qm <- .mean_life()
  b10 <- unname(q10["est"]); b10_lo <- unname(q10["lo"]); b10_hi <- unname(q10["hi"])
  b50 <- unname(q50["est"]); b50_lo <- unname(q50["lo"]); b50_hi <- unname(q50["hi"])
  mttf <- unname(qm["est"]); mttf_lo <- unname(qm["lo"]); mttf_hi <- unname(qm["hi"])

Step 6: The failure-mode reading of beta — decided by the CI, not the point

mode_key <- if (!ci_available) {
    "indeterminate"
  } else if (beta_hi < 1) {
    "infant_mortality"
  } else if (beta_lo > 1) {
    "wear_out"
  } else {
    "indeterminate"
  }
  ci_phrase <- if (ci_available) {
    paste0("95% CI ", .fmt_sig(beta_lo), " to ", .fmt_sig(beta_hi))
  } else {
    "no confidence interval could be computed"
  }
  lean_word <- if (beta > 1) "above 1" else if (beta < 1) "below 1" else "at 1"
  mode_label <- switch(mode_key,
    infant_mortality = "infant mortality(failure rate falling with age)",
    wear_out = "wear-out(failure rate rising with age)",
    "not distinguishable from a constant failure rate")
  mode_sentence <- switch(mode_key,
    infant_mortality = paste0(
      "The shape parameter is ", .fmt_sig(beta), " (", ci_phrase,
      "), entirely below 1, so the fitted failure rate falls as units age — the ",
      "infant-mortality pattern, in which the weak units fail early and the ",
      "survivors are more reliable than the population started out. Burn-in or ",
      "screening acts on this pattern; scheduled replacement does not."),
    wear_out = paste0(
      "The shape parameter is ", .fmt_sig(beta), " (", ci_phrase,
      "), entirely above 1, so the fitted failure rate rises as units age — the ",
      "wear-out pattern, in which older units are progressively more likely to ",
      "fail. Age-based preventive replacement acts on this pattern; it does not ",
      "help when the rate is flat."),
    paste0(
      "The shape parameter is ", .fmt_sig(beta), " (", ci_phrase,
      "), and that interval spans 1. On this data the failure rate cannot be ",
      "told apart from a constant one, so no failure mode is claimed: the point ",
      "estimate sits ", lean_word, ", but the evidence does not separate infant ",
      "mortality, random failure, and wear-out."))

Step 7: Observed follow-up, and the horizon the projection reaches

t_max_obs <- max(d$time)
  t_last_failure <- max(d$time[d$event == 1])
  t_r05 <- unname(.life_q(0.95)["est"])          # time by which 95% have failed
  t_horizon <- min(max(t_r05, 1.25 * t_max_obs), 2.5 * t_max_obs)
  if (!is.finite(t_horizon) || t_horizon <= t_max_obs) t_horizon <- 1.25 * t_max_obs
  extrapolation_factor <- t_horizon / t_max_obs
  .reliability <- function(t) exp(-(t / eta)^beta)
  .hazard <- function(t) (beta / eta) * (t / eta)^(beta - 1)
  r_at_obs_end <- .reliability(t_max_obs)
  r_at_horizon <- .reliability(t_horizon)

Step 8: Curves — fitted R(t), the projected segment, and observed KM

t_start <- max(min(d$time) / 2, t_horizon / 500)
  grid <- unique(c(0, seq(t_start, t_horizon, length.out = 160)))
  in_data <- grid <= t_max_obs
  fit_lab  <- "Weibull fit(within observed data)"
  proj_lab <- "Projection beyond observed data(model assumption)"
  km_lab   <- "Observed(Kaplan-Meier)"
  rel_fit <- data.frame(
    time_point = round(grid, 4),
    reliability_pct = round(100 * .reliability(grid), 3),
    series = ifelse(in_data, fit_lab, proj_lab),
    stringsAsFactors = FALSE
  )
  km_fit <- survfit(Surv(time, event) ~ 1, data = d)
  km_s <- summary(km_fit)
  km_t <- km_s$time; km_surv <- km_s$surv
  keep_km <- which(is.finite(km_t) & is.finite(km_surv))
  km_t <- km_t[keep_km]; km_surv <- km_surv[keep_km]
  idx <- if (length(km_t) > 140) unique(round(seq(1, length(km_t), length.out = 140))) else seq_along(km_t)
  rel_km <- data.frame(
    time_point = round(c(0, km_t[idx]), 4),
    reliability_pct = round(100 * c(1, km_surv[idx]), 3),
    series = km_lab,
    stringsAsFactors = FALSE
  )
  reliability_df <- rbind(rel_fit, rel_km)
  rownames(reliability_df) <- NULL

  hazard_grid <- seq(t_start, t_horizon, length.out = 160)
  hazard_df <- data.frame(
    time_point = round(hazard_grid, 4),
    hazard_rate = signif(.hazard(hazard_grid), 6),
    series = ifelse(hazard_grid <= t_max_obs, fit_lab, proj_lab),
    stringsAsFactors = FALSE
  )

Step 9: Weibull probability plot + fit-quality diagnostics

Plotting positions come from the Kaplan-Meier estimate, which is what makes them valid when units are censored. A Weibull population is a straight line in ln(t) vs ln(-ln R) coordinates, with slope beta.

ok_pp <- which(km_surv > 0 & km_surv < 1 & km_t > 0)
  pp_r2 <- NA_real_; km_max_gap_pp <- NA_real_
  probplot_df <- data.frame(log_time = numeric(0), log_log_failure = numeric(0),
                            series = character(0), stringsAsFactors = FALSE)
  if (length(ok_pp) >= 3) {
    x_all <- log(km_t[ok_pp])
    y_all <- log(-log(km_surv[ok_pp]))
    good <- is.finite(x_all) & is.finite(y_all)
    x_all <- x_all[good]; y_all <- y_all[good]
    if (length(x_all) >= 3 && stats::var(x_all) > 0 && stats::var(y_all) > 0) {
      pp_r2 <- suppressWarnings(stats::cor(x_all, y_all))^2
      sidx <- if (length(x_all) > 400) unique(round(seq(1, length(x_all), length.out = 400))) else seq_along(x_all)
      obs_pp <- data.frame(
        log_time = round(x_all[sidx], 4),
        log_log_failure = round(y_all[sidx], 4),
        series = "Observed failures(Kaplan-Meier positions)",
        stringsAsFactors = FALSE
      )
      line_x <- seq(min(x_all), max(x_all), length.out = 60)
      line_pp <- data.frame(
        log_time = round(line_x, 4),
        log_log_failure = round(beta * (line_x - log(eta)), 4),
        series = "Fitted Weibull line",
        stringsAsFactors = FALSE
      )
      probplot_df <- rbind(obs_pp, line_pp)
      rownames(probplot_df) <- NULL
    }

How far the fitted curve sits from the observed KM curve, in points

gaps <- abs(100 * km_surv[ok_pp] - 100 * .reliability(km_t[ok_pp]))
    gaps <- gaps[is.finite(gaps)]
    if (length(gaps) > 0) km_max_gap_pp <- max(gaps)
  }
  fit_quality_word <- if (is.na(pp_r2)) {
    "could not be scored"
  } else if (pp_r2 >= 0.98) {
    "very close to a straight line"
  } else if (pp_r2 >= 0.95) {
    "close to a straight line"
  } else if (pp_r2 >= 0.90) {
    "loosely straight, with visible curvature"
  } else {
    "clearly curved, which argues against a single Weibull population"
  }

Step 10: Tables

ident_note <- if (few_failures) {
    paste0("Only ", n_failures, " failure(s) were observed, so the shape is ",
           "poorly identified — read the interval, not the point estimate.")
  } else {
    paste0(n_failures, " observed failures carry the shape estimate; the ",
           "interval below is the honest range.")
  }
  params_df <- data.frame(
    parameter = c("Shape(beta)",
                  "Scale(eta) — characteristic life",
                  "Observed failures carrying the fit",
                  "Probability-plot R-squared",
                  "Largest gap to the observed curve(percentage points)"),
    estimate = c(round(beta, 4), round(eta, 3), n_failures,
                 if (is.na(pp_r2)) NA_real_ else round(pp_r2, 4),
                 if (is.na(km_max_gap_pp)) NA_real_ else round(km_max_gap_pp, 2)),
    ci_low = c(if (is.na(beta_lo)) NA_real_ else round(beta_lo, 4),
               if (is.na(eta_lo)) NA_real_ else round(eta_lo, 3),
               NA_real_, NA_real_, NA_real_),
    ci_high = c(if (is.na(beta_hi)) NA_real_ else round(beta_hi, 4),
                if (is.na(eta_hi)) NA_real_ else round(eta_hi, 3),
                NA_real_, NA_real_, NA_real_),
    interpretation = c(
      paste0("Below 1 means the failure rate falls with age, above 1 means it rises, ",
             "1 means it is flat. Here: ", mode_label, ". ", ident_note),
      paste0("The ", h_time, " by which 63.2% of units have failed — the natural ",
             "scale of this life distribution, not a safe operating limit."),
      paste0("Censored units(", format(n_censored, big.mark = ","),
             ", ", .fmt_fixed(censored_pct, 1), "% of the file) contribute the ",
             "time they survived, but only failures pin down the shape."),
      paste0("How straight the data is in Weibull coordinates: ", fit_quality_word,
             ". This scores the fit, and a good fit is still a model choice, not a fact."),
      paste0("The largest distance between the fitted reliability curve and the ",
             "assumption-free Kaplan-Meier curve computed from the same data.")
    ),
    stringsAsFactors = FALSE
  )

  proj_note <- paste0(
    "Beyond ", .fmt_sig(t_max_obs), " ", h_time,
    " there is no data — this number is the fitted model extended forward, not an observation.")
  life_df <- data.frame(
    metric = c("B10 life(10% of units failed)",
               "B50 life(half of units failed)",
               "Characteristic life(63.2% failed)",
               "Mean life(MTBF)",
               paste0("Reliability at the end of the data(", .fmt_sig(t_max_obs), ")"),
               paste0("Reliability projected to ", .fmt_sig(t_horizon))),
    estimate = c(round(b10, 3), round(b50, 3), round(eta, 3), round(mttf, 3),
                 round(100 * r_at_obs_end, 2), round(100 * r_at_horizon, 2)),
    ci_low = c(if (is.na(b10_lo)) NA_real_ else round(b10_lo, 3),
               if (is.na(b50_lo)) NA_real_ else round(b50_lo, 3),
               if (is.na(eta_lo)) NA_real_ else round(eta_lo, 3),
               if (is.na(mttf_lo)) NA_real_ else round(mttf_lo, 3),
               NA_real_, NA_real_),
    ci_high = c(if (is.na(b10_hi)) NA_real_ else round(b10_hi, 3),
                if (is.na(b50_hi)) NA_real_ else round(b50_hi, 3),
                if (is.na(eta_hi)) NA_real_ else round(eta_hi, 3),
                if (is.na(mttf_hi)) NA_real_ else round(mttf_hi, 3),
                NA_real_, NA_real_),
    interpretation = c(
      paste0("The design life engineers usually quote: 10% of units are expected to ",
             "have failed by ", .fmt_sig(b10), " ", h_time, "."),
      paste0("Half of all units are expected to have failed by ", .fmt_sig(b50), " ", h_time, "."),
      paste0("The Weibull scale parameter, always the 63.2% point whatever the shape."),
      paste0("The average life implied by the fit(eta times the gamma function of ",
             "1 + 1/beta). It is the mean, not a survival guarantee: ",
             .fmt_fixed(100 * (1 - .reliability(mttf)), 1),
             "% of units are expected to have failed by then."),
      paste0("The share still running at the last ", h_time,
             " actually covered by this data — the furthest point the file can speak to."),
      paste0("The fitted survival share at ", .fmt_sig(t_horizon), " ", h_time, ". ", proj_note)
    ),
    stringsAsFactors = FALSE
  )

  metrics <- list(
    `Units`            = n_units,
    `Failures`         = as.integer(n_failures),
    `Censored %`       = round(censored_pct, 1),
    `Shape(beta)`     = round(beta, 3),
    `Scale(eta)`      = round(eta, 2),
    `B10 Life`         = round(b10, 2),
    `B50 Life`         = round(b50, 2),
    `Mean Life(MTBF)` = round(mttf, 2),
    `Failure Mode`     = mode_label
  )

  cens_phrase <- if (n_censored == 0) {
    paste0("every unit in the file failed, so there is no right-censoring to correct for")
  } else {
    paste0(format(n_censored, big.mark = ","), " of ", format(n_units, big.mark = ","),
           " units(", .fmt_fixed(censored_pct, 1),
           "%) were still running when observation stopped and are counted as censored")
  }
  json_output <- list(
    answer = paste0(
      "Weibull life analysis of ", format(n_units, big.mark = ","), " units with ",
      n_failures, " observed failures — ", cens_phrase, ". Fitted shape beta = ",
      .fmt_sig(beta),
      if (ci_available) paste0(" (95% CI ", .fmt_sig(beta_lo), " to ", .fmt_sig(beta_hi), ")") else "",
      ", scale eta = ", .fmt_sig(eta), " ", h_time,
      ", which reads as ", mode_label, ". B10 life ", .fmt_sig(b10), ", B50 life ",
      .fmt_sig(b50), ", mean life ", .fmt_sig(mttf), " ", h_time, ". ",
      "The data itself only reaches ", .fmt_sig(t_max_obs), " ", h_time,
      "; the reliability curve is projected to ", .fmt_sig(t_horizon),
      " (", .fmt_fixed(extrapolation_factor, 2),
      " times the observed follow-up), and everything past the end of the data is ",
      "the fitted model rather than evidence.",
      if (few_failures) paste0(" With only ", n_failures,
        " failures the shape is poorly identified — read the interval, not the point estimate.") else ""
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "weibull_parameters",
        "reliability_curve", "life_metrics", "hazard_curve", "probability_plot"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows, rows_removed = rows_removed,
    h_time = h_time, h_event = h_event, failure_level = failure_level,
    n_units = n_units, n_failures = n_failures, n_censored = n_censored,
    censored_pct = censored_pct, few_failures = few_failures,
    dropped_time_rows = dropped_time_rows, dropped_event_rows = dropped_event_rows,
    beta = beta, beta_lo = beta_lo, beta_hi = beta_hi, beta_p = beta_p,
    eta = eta, eta_lo = eta_lo, eta_hi = eta_hi, sigma = sigma, mu = mu,
    ci_available = ci_available,
    mode_key = mode_key, mode_label = mode_label, mode_sentence = mode_sentence,
    b10 = b10, b10_lo = b10_lo, b10_hi = b10_hi,
    b50 = b50, b50_lo = b50_lo, b50_hi = b50_hi,
    mttf = mttf, mttf_lo = mttf_lo, mttf_hi = mttf_hi,
    t_max_obs = t_max_obs, t_last_failure = t_last_failure, t_horizon = t_horizon,
    extrapolation_factor = extrapolation_factor,
    r_at_obs_end = r_at_obs_end, r_at_horizon = r_at_horizon,
    pp_r2 = pp_r2, km_max_gap_pp = km_max_gap_pp, fit_quality_word = fit_quality_word,
    reliability_df = reliability_df, hazard_df = hazard_df, probplot_df = probplot_df,
    params_df = params_df, life_df = life_df,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell. Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No Signup Sign Up Free

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing