Standard Gap Analysis
Executive Summary

Executive Summary

The largest importance-performance gap across 4 attributes.

Responses
1200
Attributes Analysed
4
Importance Source
derived
Top Priority
Service Speed
Top Priority Score
1.57
Concentrate Here
1
Importance Boundary
0.385
Performance Boundary
5.157
Reclassified Under Alternative Boundary
1
Across 1,200 responses and 4 attributes, Service Speed has the widest gap between how much it matters and how well it is rated (priority score 1.57 standard deviations, importance 0.58 against performance 4.22). At the grand-mean boundaries (importance 0.38, performance 5.16), 1 attribute sits in Concentrate Here: Service Speed. The weakest case for attention is Product Quality. The picture is boundary-sensitive: under a fixed correlation cut of 0.30 on importance and the scale midpoint (5.0) on performance, Website Ease would change quadrant. Importance here is association, not a proven lever: the map says where the unmet need is, not that closing it will raise the outcome.
Suggested Interpretation

The short answer

Service Speed is the clear priority: it carries the widest gap between importance and performance, with an importance of 0.58 against a performance rating of 4.22, yielding a priority score of 1.57 standard deviations above the set mean. At the stated boundaries, only 1 attribute lands in Concentrate Here—Service Speed—making it the only quadrant-confirmed unmet need.

The detail

Across 1,200 responses and 4 attributes analysed, Service Speed leads the priority ranking with a score of 1.574 standard deviations. Its importance (derived correlation) is 0.58 and its mean performance rating is 4.22. The analysis applies grand-mean boundaries: importance 0.385 and performance 5.157. Under an alternative fixed boundary (importance 0.30, performance 5.0), 1 attribute reclassifies: Website Ease moves from Possible Overkill to Keep Up The Good Work, but Service Speed remains in Concentrate Here in both cases.

What this can't tell you

Importance is association with satisfaction, not a lever that will move it if improved. The map identifies where unmet need sits, not whether closing the gap will raise the outcome.

Overview

Analysis Overview

How 4 attributes are placed on the importance-performance map.

N Observations1200
N Attributes4
N Concentrate1
N Reclassified1
Suggested Interpretation

The short answer

This map shows where unmet customer need sits. It plots 4 attributes on two axes: how much each matters (derived from correlation with overall satisfaction) and how well it is currently rated. The four quadrants divide attributes by importance and performance, with boundaries drawn at the grand mean of each axis—importance 0.38 and performance 5.16. The core insight is that every placement is relative: the map ranks within this attribute set rather than measuring any absolute shortfall, and moving the boundary shifts which attributes demand attention.

The detail

Importance is measured as derived importance—the correlation of each attribute's rating with Overall Satisfaction across 1,200 responses. Performance is the mean rating on the underlying scale. The four quadrants are: Concentrate Here (important, weak), Keep Up The Good Work (important, strong), Low Priority (unimportant, weak), and Possible Overkill (unimportant, strong). The analysis included n_observations of 1,200 and n_attributes of 4. No stated-importance data were used; the map is built entirely from how strongly each attribute tracks satisfaction.

What this can't tell you

The boundary is a methodological choice, not a fact about the data. The grand-mean convention always creates relative winners and losers even when all attributes are genuinely strong or all weak. Derived importance shows association, not causation—an attribute that moves with satisfaction may not move it if changed.

Data Preparation

Data Quality

Rows and attributes used, exclusions, imputation, and the rating scale.

Initial Rows1200
Final Rows1200
Rows Removed0
N Attributes4
N Imputed0
Suggested Interpretation

The short answer

All 1,200 responses and all four attribute columns were usable; no rows were dropped and no ratings were imputed. The 0–10 rating scale was detected, setting the midpoint at 5.0 for the alternative boundary. Attribute names were shortened by removing shared wording, so "Satisfaction: Price" appears as "Price."

The detail

Initial load: 1,200 rows. Final rows: 1,200 (rows removed: 0). No attribute rating was missing, so n_imputed = 0. Four attributes carried to the map: Price, Service Speed, Product Quality, Website Ease. The rating scale runs from 0 to 10, establishing 5.0 as the scale midpoint used in the fixed-boundary alternative.

What this can't tell you

Data quality was clean for this analysis; no imputation decisions or row exclusions created hidden assumptions. The only caveat is that derived importance depends on the quality of the Overall Satisfaction measure—if that item itself was misunderstood or had high missingness in the source data, the correlations would be unreliable, but that check lies outside the scope of this preprocessing report.

Visualization

Importance vs Performance Map

Each attribute placed by derived importance against current performance, split into four quadrants.

Suggested Interpretation

The short answer

Service Speed sits alone in Concentrate Here (upper left): it matters more than the average attribute here but is rated well below average. Website Ease occupies Possible Overkill (upper right), rated well despite below-average importance. Product Quality anchors Keep Up The Good Work (upper right), both important and well-rated. Price lands in Low Priority (lower left), neither important nor well-rated. The pattern is not a tight trend—attributes are scattered across all four quadrants, with Service Speed the clear outlier on the left side.

The detail

All four quadrants are populated. Reference lines sit at grand means: importance 0.3849, performance 5.1573. Service Speed: importance 0.577, performance 4.218. Website Ease: importance 0.372, performance 5.418. Product Quality: importance 0.465, performance 7.35. Price: importance 0.126, performance 3.644. Attributes near a dividing line are not meaningfully different from those just across it; the corners carry the signal. Service Speed is the only point in the upper left, making it the isolated priority.

What this can't tell you

The map shows correlation between attribute ratings and Overall Satisfaction, not causal relationships. An attribute's position depends on the boundary choice—moving the cross-hair changes which attributes are flagged as priorities. The analysis cannot distinguish between halo effects (satisfied respondents rate everything higher) and true attribute importance, though inter-attribute correlations of 0.05 or less suggest each attribute's signal is largely independent.

Visualization

Priority Ranking

Attributes ranked by the size of the importance-performance gap.

Suggested Interpretation

The short answer

Service Speed dominates the priority ranking by a wide margin. It is the only attribute with a positive gap score, meaning it is the only one where importance exceeds performance. The next three attributes all show negative scores, indicating they either deliver more than their importance warrants or both underperform and matter less.

The detail

Service Speed leads at a priority score of 1.574 standard deviations. Website Ease follows at -0.225, Price at -0.426, and Product Quality trails at -0.922. The priority score is calculated as how far above average an attribute sits on importance minus how far above average it sits on performance, both in standard deviations across the 4 attributes. Only 1 attribute has a positive score. The score is relative to this attribute set and will re-rank if attributes are added or dropped.

What this can't tell you

The score does not measure any absolute shortfall—it prioritizes within this list rather than against an external standard. Adding or removing attributes will shift the boundaries and may move attributes across quadrants.

Data Table

Attribute Detail

Importance, performance, gap, and quadrant for each of 4 attributes.

AttributeImportancePerformancePriority ScoreQuadrant
Service Speed0.5774.2181.574Concentrate Here
Website Ease0.3725.418-0.225Possible Overkill
Price0.1263.644-0.426Low Priority
Product Quality0.4657.35-0.922Keep Up The Good Work
Suggested Interpretation

The short answer

Service Speed is the decisive finding: it is the only attribute in Concentrate Here, with an importance of 0.577 and a performance rating of 4.218. Product Quality, the second-strongest attribute on importance (0.465), delivers the highest performance (7.35) and sits in Keep Up The Good Work. The gap between Service Speed's importance and its performance is the largest in the set.

The detail

Service Speed: importance 0.577, performance 4.218, priority score 1.574, quadrant Concentrate Here. Website Ease: importance 0.372, performance 5.418, priority score -0.225, quadrant Possible Overkill. Price: importance 0.126, performance 3.644, priority score -0.426, quadrant Low Priority. Product Quality: importance 0.465, performance 7.35, priority score -0.922, quadrant Keep Up The Good Work. One attribute sits in Concentrate Here and one in Keep Up The Good Work; the other two are in lower-priority quadrants.

What this can't tell you

Two attributes can share a quadrant while sitting at opposite ends of it. The quadrant assignment depends on the boundary chosen; Website Ease moves under the alternative boundary convention.

Data Table

Stated vs Derived Importance

Where the two ways of measuring importance agree, and where they contradict each other.

AttributeStated ImportanceDerived ImportanceDerived EvidenceStated RankDerived RankRank Shift
Service Speed0.577p < 0.0011
Website Ease0.372p < 0.0013
Price0.126p < 0.0014
Product Quality0.465p < 0.0012
Suggested Interpretation

The short answer

Only derived importance is available—no stated-importance columns were mapped, so the analysis cannot check whether customers themselves would call these attributes important. Derived importance is inferred from each attribute's correlation with Overall Satisfaction. Service Speed ranks first (0.577 correlation, p < 0.001), Product Quality second (0.465, p < 0.001), Website Ease third (0.372, p < 0.001), and Price fourth (0.126, p < 0.001).

The detail

Stated importance: not available. Derived importance and derived evidence (all p < 0.001): Service Speed 0.577, Website Ease 0.372, Price 0.126, Product Quality 0.465. Derived ranks: Service Speed 1, Product Quality 2, Website Ease 3, Price 4. No rank shift can be computed because stated importance does not exist. Halo effects (respondents who are satisfied overall rating every attribute higher) could inflate all correlations at once, but inter-attribute correlations are at most 0.05 here, so each attribute's score is close to its own separate contribution.

What this can't tell you

The analysis cannot compare stated and derived importance to see where customer perception diverges from behavior. Mapping an attribute-level importance rating per respondent would enable that cross-check and is where this analysis is most useful. Derived importance is association, not causation: it shows which attributes move with Overall Satisfaction, not which ones would move it if changed.

Data Table

Boundary Sensitivity

How the quadrant picture changes under the other boundary convention.

AttributeQuadrant Grand MeanQuadrant AlternativeChanged
Service SpeedConcentrate HereConcentrate Hereno
Website EasePossible OverkillKeep Up The Good Workyes
PriceLow PriorityLow Priorityno
Product QualityKeep Up The Good WorkKeep Up The Good Workno
Suggested Interpretation

The short answer

Service Speed stays in Concentrate Here under both boundary conventions, confirming it as a robust priority. Website Ease is the only attribute that moves: it shifts from Possible Overkill (grand-mean boundary) to Keep Up The Good Work (fixed alternative boundary). This sensitivity means any recommendation that rests solely on Website Ease's quadrant placement is a recommendation about the boundary, not the data.

The detail

At the grand-mean boundaries (importance 0.38, performance 5.16), Service Speed remains in Concentrate Here. Website Ease moves from Possible Overkill to Keep Up The Good Work under the fixed alternative (importance 0.30, performance 5.0). Price and Product Quality do not move. The grand-mean boundary always creates relative winners and losers within the set; the fixed alternative is independent of this attribute set but ignores how this particular distribution is shaped.

What this can't tell you

Neither boundary convention is correct. The choice between them depends on your use case: use grand means to rank within your own list; use the fixed alternative when you need the map to remain consistent across surveys or over time.

Data Table

Prioritized Action List

The recommended next step for each quadrant, ordered by urgency.

QuadrantN AttributesAttributesAction
Concentrate Here1Service SpeedFix these first. They matter more than average and are rated below average, so this is where the largest unmet need sits.
Keep Up The Good Work1Product QualityProtect these. They matter and are already rated well — they are the strengths worth defending rather than improving further.
Low Priority1PriceLeave these alone for now. They are rated below average but also matter less than average, so weak scores here cost comparatively little.
Possible Overkill1Website EaseConsider easing off. These are rated well but matter less than average, so effort spent here may be buying little.
Suggested Interpretation

The short answer

Start with Service Speed in Concentrate Here—it is the only attribute where importance clearly exceeds performance, making it the largest unmet need. Protect Product Quality in Keep Up The Good Work rather than investing further. Leave Price alone in Low Priority and consider easing off Website Ease in Possible Overkill. Before committing budget, verify that the boundary-sensitivity finding (1 attribute moves under the alternative convention) does not change your strategic choice.

The detail

Concentrate Here (Service Speed, n_attributes 1): Fix first—largest unmet need. Keep Up The Good Work (Product Quality, n_attributes 1): Protect rather than improve. Low Priority (Price, n_attributes 1): Leave alone while unimportant. Possible Overkill (Website Ease, n_attributes 1): Consider releasing effort. All recommendations rest on association between attribute ratings and Overall Satisfaction; treat each as a hypothesis worth testing.

What this can't tell you

Derived importance is correlation, not proof of causation. Improving Service Speed may or may not raise Overall Satisfaction; the map shows where the unmet need is, not that closing it will move the outcome.

Methodology

Methodology

Statistical methodology and diagnostics for Importance-Performance Gap Analysis

Statistical Method

Importance-Performance Gap Analysis

Standard-library analysis: what matters most that we do worst? Classic Importance-Performance Analysis on survey data. Map your attribute satisfaction columns and either an importance rating per attribute or an overall satisfaction score, and get the four-quadrant map (Concentrate Here, Keep Up The Good Work, Low Priority, Possible Overkill), every attribute ranked by the size of its importance-performance gap, and a prioritized action list. When both stated importance and an overall score are available, both are computed and the disagreement between them is reported rather than one being picked silently — and because the quadrant boundary is a methodological choice, the analysis states the values it used and names every attribute that would move under the other convention.

Data
N = 1200 observations
Assumptions
  • Each row is one respondent, with one rating per attribute
  • Attribute ratings are numeric or cleanly convertible, and on a common scale
  • Attributes are rated on the same scale as stated importance, if the raw gap is to be read in scale points
  • For derived importance, the overall score reflects the attributes measured rather than something outside the survey
Limitations
  • Quadrant placement is relative to the attributes you mapped — adding or removing an attribute moves the boundaries and can move other attributes across them
  • Derived importance is a correlation with the overall score, not proof that improving an attribute will raise it
  • Derived importance is inflated by halo, where a respondent who is satisfied overall rates every attribute higher; when the attributes are correlated with each other, a bivariate score shares credit between them
  • Stated importance suffers ceiling effects — respondents often rate nearly everything as important, which flattens the horizontal axis
Software & Citation
MCP Analytics · mcpanalytics.ai
Code Appendix

Analysis Code

Complete R source code for this analysis

Importance-Performance Gap Analysis

What matters most that we do worst? Classic Importance-Performance Analysis (IPA): every attribute is placed on a four-quadrant map by how important it is against how well it is currently rated, ranked by the gap between the two, and turned into a prioritized action list.

Why This Method?

Ranking attributes by performance alone tells you where you are weak but not whether anyone cares. Ranking by importance alone tells you what matters but not where you are failing. Crossing the two is the whole point: the attributes that are important AND under-performing are the ones worth money, and the map makes that visible in one picture.

Two ways to measure importance — and they disagree

STATED importance is what respondents said mattered. DERIVED importance is inferred from how strongly each attribute tracks an overall satisfaction score. They routinely disagree, and the disagreement is informative rather than an error: people under-report what actually drives them and over-report what they think they should care about. This module detects which inputs it was given, computes BOTH when both are available, and shows where they contradict each other instead of quietly picking one.

What This Analysis Covers

  • Attribute-level importance and performance on a four-quadrant map
  • The priority ranking by importance-performance gap
  • Stated versus derived importance, and where they disagree
  • How the quadrant picture changes under the other boundary convention
  • A per-quadrant action list

Sibling tool

standard_key_drivers answers "what drives satisfaction?" — it ranks drivers by derived importance and reports model fit. This tool answers "what should we fix first?" — it takes importance as given (stated or derived), crosses it with current performance, and produces a prioritized gap list. Same vocabulary, different question.

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {performance_1..N, importance_1..N, overall}. All narrative is derived from the user's own column names and computed values. Derived importance is CORRELATIONAL — it reports association, never proven causation.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Longest shared leading run of whole words.

n_lead <- 0
  repeat {
    k <- n_lead + 1
    if (any(sapply(parts, length) <= k)) break
    w <- sapply(parts, function(p) p[k])
    if (length(unique(w)) != 1) break
    n_lead <- k
  }

Longest shared trailing run of whole words.

n_tail <- 0
  repeat {
    k <- n_tail + 1
    if (any(sapply(parts, length) <= n_lead + k)) break
    w <- sapply(parts, function(p) p[length(p) - k + 1])
    if (length(unique(w)) != 1) break
    n_tail <- k
  }
  if (n_lead == 0 && n_tail == 0) return(labels)

  out <- sapply(parts, function(p) {
    keep <- p[(n_lead + 1):(length(p) - n_tail)]
    trimws(paste(keep, collapse = " "))
  }, USE.NAMES = FALSE)

Refuse the strip if it empties or collides any label.

if (any(nchar(out) == 0) || anyDuplicated(out) > 0) return(labels)
  out
}

Step 1: Row accounting + semantic column discovery

initial_rows <- nrow(df)
  perf_cols <- grep("^performance_[0-9]+$", names(df), value = TRUE)
  perf_cols <- perf_cols[order(as.integer(sub("^performance_", "", perf_cols)))]
  imp_cols <- grep("^importance_[0-9]+$", names(df), value = TRUE)
  imp_cols <- imp_cols[order(as.integer(sub("^importance_", "", imp_cols)))]
  has_overall_col <- "overall" %in% names(df)

  if (length(perf_cols) == 0) {
    stop(paste0("column_mapping must map at least ", MIN_ATTRS,
                " performance columns(performance_1, performance_2, ...) — ",
                "one per attribute you rate."))
  }

  perf_names_raw <- humanize_semantic(perf_cols, col_map)
  imp_names_raw  <- if (length(imp_cols) > 0) humanize_semantic(imp_cols, col_map) else character(0)
  overall_name   <- if (has_overall_col) humanize_semantic("overall", col_map) else NA_character_

Attribute labels come from the PERFORMANCE columns, with any wording shared by all of them removed ("Satisfaction: Price" -> "Price").

attr_labels_all <- strip_common_affix(perf_names_raw)
  affix_stripped <- !identical(attr_labels_all, perf_names_raw)
  names(attr_labels_all) <- perf_cols

Stated importance is paired to performance BY POSITION: importance_1 describes the same attribute as performance_1. A partial or mismatched set cannot be paired safely, so it is refused rather than guessed.

paired_stated <- length(imp_cols) > 0 && length(imp_cols) == length(perf_cols) &&
    identical(sub("^importance_", "", imp_cols), sub("^performance_", "", perf_cols))
  if (length(imp_cols) > 0 && !paired_stated) {
    stop(paste0(
      "Stated importance must be mapped one-for-one with performance: ",
      n_things(length(imp_cols), "importance column"), " were mapped(",
      paste(imp_names_raw, collapse = ", "), ") against ",
      n_things(length(perf_cols), "performance column"), " (",
      paste(perf_names_raw, collapse = ", "),
      "). Map the same attributes in the same order, or map none and supply ",
      "an overall satisfaction column instead."))
  }

Step 2: Coerce every mapped rating to numeric (95% rule)

A column that will not convert cleanly, or that has no usable values, is dropped and reported rather than silently coerced to NA.

dropped_cols <- character(0)
  coerce <- function(dd, cc) {
    v <- dd[[cc]]
    if (is.numeric(v)) return(v)
    conv <- suppressWarnings(as.numeric(as.character(v)))
    n_orig <- sum(!is.na(v) & as.character(v) != "")
    if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) conv else NULL
  }
  bad_perf <- character(0)
  for (pc in perf_cols) {
    conv <- coerce(df, pc)
    if (is.null(conv)) { bad_perf <- c(bad_perf, pc); next }
    df[[pc]] <- conv
  }
  bad_imp <- character(0)
  if (paired_stated) {
    for (ic in imp_cols) {
      conv <- coerce(df, ic)
      if (is.null(conv)) { bad_imp <- c(bad_imp, ic); next }
      df[[ic]] <- conv
    }
  }
  overall_usable <- FALSE
  if (has_overall_col) {
    conv <- coerce(df, "overall")
    if (!is.null(conv)) { df$overall <- conv; overall_usable <- TRUE }
  }

Step 3: Drop constant / all-missing attributes — and their partner

An attribute is only usable if BOTH sides that will be plotted survive.

for (i in seq_along(perf_cols)) {
    pc <- perf_cols[i]
    if (pc %in% bad_perf) next
    v <- df[[pc]]
    if (all(is.na(v))) { bad_perf <- c(bad_perf, pc); next }
    if (isTRUE(stats::var(v, na.rm = TRUE) == 0) ||
        is.na(stats::var(v, na.rm = TRUE))) {
      bad_perf <- c(bad_perf, pc)
    }
  }
  drop_idx <- which(perf_cols %in% bad_perf)
  if (paired_stated) drop_idx <- union(drop_idx, which(imp_cols %in% bad_imp))
  keep_idx <- setdiff(seq_along(perf_cols), drop_idx)
  if (length(drop_idx) > 0) {
    dropped_cols <- unname(attr_labels_all[perf_cols[drop_idx]])
  }
  perf_use <- perf_cols[keep_idx]
  imp_use  <- if (paired_stated) imp_cols[keep_idx] else character(0)
  attr_labels <- unname(attr_labels_all[perf_use])

  if (length(perf_use) < MIN_ATTRS) {
    kept_note <- if (length(perf_use) > 0)
      paste0(" The attributes read were: ", paste(attr_labels, collapse = ", "), ".")
    else
      paste0(" The columns mapped were: ", paste(perf_names_raw, collapse = ", "), ".")
    stop(sprintf(paste0(
      "Importance-Performance Analysis needs at least %d usable attributes; ",
      "only %d of the %d mapped performance columns survived cleaning%s.%s ",
      "A four-quadrant map of fewer than %d attributes is not meaningful."),
      MIN_ATTRS, length(perf_use), length(perf_cols),
      if (length(dropped_cols) > 0)
        paste0(" (excluded as constant, empty, or non-numeric: ",
               paste(dropped_cols, collapse = ", "), ")") else "",
      kept_note, MIN_ATTRS))
  }

Step 4: Which importance sources do we actually have?

have_stated  <- paired_stated && length(imp_use) == length(perf_use)
  have_derived <- overall_usable && sum(!is.na(df$overall)) >= MIN_ROWS
  if (!have_stated && !have_derived) {
    stop(paste0(
      "Importance-Performance Analysis needs importance from somewhere. ",
      "Either map an importance column for each attribute(",
      paste(attr_labels, collapse = ", "),
      "), or map an overall satisfaction column so importance can be ",
      "derived from how each attribute tracks it."))
  }

Step 5: Rows — derived importance needs a usable overall value

if (have_derived) df <- df[!is.na(df$overall), , drop = FALSE]
  keep_row <- rowSums(!is.na(df[, perf_use, drop = FALSE])) > 0
  df <- df[keep_row, , drop = FALSE]
  final_rows <- nrow(df)
  rows_removed <- initial_rows - final_rows
  if (final_rows < MIN_ROWS) {
    stop(sprintf(paste0(
      "Only %d usable responses remain after cleaning%s — at least %d are ",
      "required before an importance-performance map means anything. ",
      "Attributes read: %s."),
      final_rows,
      if (have_derived) sprintf(" (rows with no &#x27;%s' value are dropped)", overall_name) else "",
      MIN_ROWS, paste(attr_labels, collapse = ", ")))
  }

Step 6: Impute remaining missing ratings with each column's median

n_imputed <- 0L
  for (cc in c(perf_use, imp_use)) {
    v <- df[[cc]]
    miss <- is.na(v)
    if (any(miss)) {
      med <- stats::median(v, na.rm = TRUE)
      if (!is.na(med)) { v[miss] <- med; n_imputed <- n_imputed + sum(miss) }
      df[[cc]] <- v
    }
  }

Step 7: Rating scale — needed for the midpoint boundary convention

rating_vals <- unlist(df[, c(perf_use, imp_use), drop = FALSE], use.names = FALSE)
  rating_vals <- rating_vals[is.finite(rating_vals)]
  vmin <- min(rating_vals); vmax <- max(rating_vals)
  scale_detected <- vmin >= 0 && vmax <= 10
  if (scale_detected) {
    scale_min <- if (vmin < 1) 0 else 1
    scale_max <- if (vmax <= 5) 5 else if (vmax <= 7) 7 else 10
  } else {
    scale_min <- vmin; scale_max <- vmax
  }
  scale_mid <- (scale_min + scale_max) / 2

Step 8: Performance — the mean rating per attribute

perf_mean <- sapply(perf_use, function(pc) mean(df[[pc]], na.rm = TRUE))
  perf_mean <- as.numeric(perf_mean)

Step 9: Stated importance — the mean importance rating per attribute

stated_imp <- if (have_stated) {
    as.numeric(sapply(imp_use, function(ic) mean(df[[ic]], na.rm = TRUE)))
  } else rep(NA_real_, length(perf_use))

Step 10: Derived importance — how each attribute tracks the overall

score. Bivariate Pearson correlation is the standard derived-importance statistic; a joint standardized regression is also fitted purely to disclose how much of that association is shared rather than unique.

derived_imp <- rep(NA_real_, length(perf_use))
  derived_p   <- rep(NA_real_, length(perf_use))
  max_attr_cor <- NA_real_
  top_beta <- NA_real_
  if (have_derived) {
    for (i in seq_along(perf_use)) {
      ct <- tryCatch(stats::cor.test(df[[perf_use[i]]], df$overall),
                     error = function(e) NULL)
      if (!is.null(ct) && is.finite(ct$estimate)) {
        derived_imp[i] <- as.numeric(ct$estimate)
        derived_p[i]   <- as.numeric(ct$p.value)
      } else {
        derived_imp[i] <- 0
      }
    }
    if (length(perf_use) >= 2) {
      cm <- suppressWarnings(stats::cor(df[, perf_use, drop = FALSE],
                                        use = "pairwise.complete.obs"))
      cm[!is.finite(cm)] <- 0
      diag(cm) <- 0
      max_attr_cor <- max(abs(cm))
    }
    zfit <- tryCatch({
      zd <- as.data.frame(scale(df[, c("overall", perf_use), drop = FALSE]))
      stats::lm(overall ~ ., data = zd)
    }, error = function(e) NULL)
    if (!is.null(zfit)) {
      bb <- stats::coef(zfit)[perf_use]
      ok <- which(!is.na(derived_imp))
      if (length(ok) > 0) {
        lead <- ok[order(-abs(derived_imp[ok]))][1]
        top_beta <- unname(bb[lead])
      }
    }
  }

Step 11: Which importance source drives the primary map?

Stated is preferred when present because it shares the performance scale, which makes the raw gap directly readable. The choice is stated in the prose and the other source is reported alongside it — never silently dropped.

req_source <- tolower(as.character(params$importance_source %||% "auto"))
  imp_source <- if (req_source == "derived" && have_derived) "derived"
                else if (req_source == "stated" && have_stated) "stated"
                else if (have_stated) "stated" else "derived"
  importance <- if (imp_source == "stated") stated_imp else derived_imp
  imp_source_h <- if (imp_source == "stated")
    "stated importance(the average importance rating respondents gave)"
  else
    paste0("derived importance(each attribute&#x27;s correlation with ", overall_name, ")")

Short form, for sentences that already carry their own parentheses.

imp_source_short <- if (imp_source == "stated") "stated importance" else "derived importance"

Step 12: Boundaries — the methodological choice that moves the map

Primary: the data-driven grand mean of each axis (the cross-hair sits at the average attribute). Alternative: the scale midpoint, which is fixed and independent of this attribute set. For derived importance the scale midpoint does not exist, so the alternative is the conventional moderate-association cut of 0.30.

imp_boundary  <- mean(importance, na.rm = TRUE)
  perf_boundary <- mean(perf_mean, na.rm = TRUE)
  perf_boundary_alt <- scale_mid
  imp_boundary_alt  <- if (imp_source == "stated") scale_mid else DERIVED_ALT_CUT
  alt_label <- if (imp_source == "stated") {
    paste0("the scale midpoint(", fmt_n(scale_mid, 1), " on a ",
           fmt_n(scale_min, 0), "-to-", fmt_n(scale_max, 0), " scale)")
  } else {
    paste0("a fixed correlation cut of ", fmt_n(DERIVED_ALT_CUT, 2),
           " on importance and the scale midpoint(", fmt_n(scale_mid, 1),
           ") on performance")
  }

  quad_of <- function(imp, perf, bi, bp) {
    ifelse(imp >= bi & perf <  bp, "Concentrate Here",
    ifelse(imp >= bi & perf >= bp, "Keep Up The Good Work",
    ifelse(imp <  bi & perf <  bp, "Low Priority",
                                   "Possible Overkill")))
  }
  quadrant     <- quad_of(importance, perf_mean, imp_boundary, perf_boundary)
  quadrant_alt <- quad_of(importance, perf_mean, imp_boundary_alt, perf_boundary_alt)
  reclassified <- quadrant != quadrant_alt
  n_reclassified <- sum(reclassified)
  reclassified_names <- attr_labels[reclassified]

Step 13: The gap. Two measures, both computed.

priority_score standardizes each axis ACROSS THE ATTRIBUTE SET, so it works whether importance is a rating or a correlation. raw_gap is the plain importance-minus-performance difference in scale points, and is only defined when both axes are on the same rating scale.

z_of <- function(v) {
    s <- stats::sd(v, na.rm = TRUE)
    if (!is.finite(s) || s == 0) rep(0, length(v))
    else (v - mean(v, na.rm = TRUE)) / s
  }
  priority_score <- z_of(importance) - z_of(perf_mean)
  imp_range <- diff(range(importance, na.rm = TRUE))
  scale_span <- if (is.finite(scale_max - scale_min) && (scale_max - scale_min) > 0)
    scale_max - scale_min else NA_real_
  imp_tied <- if (imp_source == "stated" && is.finite(scale_span))
    imp_range < 0.05 * scale_span else FALSE
  raw_gap_available <- imp_source == "stated"
  raw_gap <- if (raw_gap_available) stated_imp - perf_mean else rep(NA_real_, length(perf_use))

Step 14: Master attribute frame, ranked by the gap

attrs_df <- data.frame(
    attribute      = attr_labels,
    importance     = round(importance, 3),
    performance    = round(perf_mean, 3),
    priority_score = round(priority_score, 3),
    raw_gap        = round(raw_gap, 3),
    quadrant       = quadrant,
    quadrant_alt   = quadrant_alt,
    stated_importance  = round(stated_imp, 3),
    derived_importance = round(derived_imp, 3),
    derived_p          = derived_p,
    stringsAsFactors = FALSE
  )
  attrs_df <- attrs_df[order(-attrs_df$priority_score, attrs_df$attribute), , drop = FALSE]
  rownames(attrs_df) <- NULL

  top_priority_name  <- attrs_df$attribute[1]
  top_priority_score <- attrs_df$priority_score[1]
  bottom_name        <- attrs_df$attribute[nrow(attrs_df)]
  concentrate_names  <- attrs_df$attribute[attrs_df$quadrant == "Concentrate Here"]
  n_concentrate      <- length(concentrate_names)

Step 15: Stated versus derived — ranks, shifts, and quadrant flips

have_both <- have_stated && have_derived
  rank_rho <- NA_real_
  n_source_flips <- 0L
  source_flip_names <- character(0)
  max_shift_name <- NA_character_
  max_shift <- NA_real_
  stated_rank  <- rep(NA_real_, nrow(attrs_df))
  derived_rank <- rep(NA_real_, nrow(attrs_df))
  if (have_stated) stated_rank  <- rank(-attrs_df$stated_importance, ties.method = "min")
  if (have_derived) derived_rank <- rank(-attrs_df$derived_importance, ties.method = "min")
  rank_shift <- if (have_both) abs(stated_rank - derived_rank) else rep(NA_real_, nrow(attrs_df))
  if (have_both) {
    rank_rho <- suppressWarnings(stats::cor(attrs_df$stated_importance,
                                            attrs_df$derived_importance,
                                            method = "spearman"))
    if (!is.finite(rank_rho)) rank_rho <- NA_real_
    ord <- order(-rank_shift, attrs_df$attribute)
    max_shift_name <- attrs_df$attribute[ord[1]]
    max_shift <- rank_shift[ord[1]]
    q_stated  <- quad_of(attrs_df$stated_importance, attrs_df$performance,
                         mean(attrs_df$stated_importance), perf_boundary)
    q_derived <- quad_of(attrs_df$derived_importance, attrs_df$performance,
                         mean(attrs_df$derived_importance), perf_boundary)
    flips <- q_stated != q_derived
    n_source_flips <- sum(flips)
    source_flip_names <- attrs_df$attribute[flips]
  }

Derived importance is an estimate, so it carries its own evidence: an attribute whose correlation is not distinguishable from zero must not be read as "unimportant with confidence".

derived_evidence <- if (have_derived) {
    sapply(seq_len(nrow(attrs_df)), function(i) {
      pv <- attrs_df$derived_p[i]
      if (is.na(pv)) return("not available")
      paste0(fmt_p(pv), if (pv < 0.05) "" else " (not distinguishable from zero)")
    })
  } else rep(NA_character_, nrow(attrs_df))

  importance_sources_df <- data.frame(
    attribute          = attrs_df$attribute,
    stated_importance  = attrs_df$stated_importance,
    derived_importance = attrs_df$derived_importance,
    derived_evidence   = derived_evidence,
    stated_rank        = stated_rank,
    derived_rank       = derived_rank,
    rank_shift         = rank_shift,
    stringsAsFactors   = FALSE
  )
  boundary_sensitivity_df <- data.frame(
    attribute             = attrs_df$attribute,
    quadrant_grand_mean   = attrs_df$quadrant,
    quadrant_alternative  = attrs_df$quadrant_alt,
    changed               = ifelse(attrs_df$quadrant != attrs_df$quadrant_alt,
                                   "yes", "no"),
    stringsAsFactors = FALSE
  )
  action_plan_df <- do.call(rbind, lapply(QUADS, function(q) {
    members <- attrs_df$attribute[attrs_df$quadrant == q]
    data.frame(
      quadrant     = q,
      n_attributes = length(members),
      attributes   = if (length(members) > 0) paste(members, collapse = ", ") else "None",
      action       = unname(QUAD_ACTIONS[q]),
      stringsAsFactors = FALSE
    )
  }))
  rownames(action_plan_df) <- NULL

Step 17: KPI metrics

metrics <- list(
    `Responses`            = final_rows,
    `Attributes Analysed`  = length(perf_use),
    `Importance Source`    = if (imp_source == "stated") "stated" else "derived",
    `Top Priority`         = top_priority_name,
    `Top Priority Score`   = round(top_priority_score, 2),
    `Concentrate Here`     = as.integer(n_concentrate),
    `Importance Boundary`  = round(imp_boundary, 3),
    `Performance Boundary` = round(perf_boundary, 3),
    `Reclassified Under Alternative Boundary` = as.integer(n_reclassified)
  )

Step 18: json_output machine channel

concentrate_clause <- if (n_concentrate > 0) {
    paste0(n_things(n_concentrate, "attribute"), " ",
           vform(n_concentrate, "falls", "fall"), " in Concentrate Here(",
           paste(concentrate_names, collapse = ", "), ")")
  } else {
    "No attribute falls in Concentrate Here"
  }
  disagree_clause <- if (have_both) {
    paste0(" Stated and derived importance rank the attributes differently ",
           "(Spearman rank correlation ", fmt_n(rank_rho, 2), "); ",
           max_shift_name, " moves the most, by ",
           n_things(as.integer(max_shift), "rank position"),
           ", and ", n_things(as.integer(n_source_flips), "attribute"), " ",
           vform(n_source_flips, "changes", "change"),
           " quadrant depending on which importance source is used.")
  } else if (imp_source == "derived") {
    paste0(" Importance was derived from ", overall_name,
           " because no stated-importance columns were mapped.")
  } else {
    " Importance is as stated by respondents; no overall satisfaction column was mapped to cross-check it."
  }
  json_output <- list(
    answer = paste0(
      "Importance-Performance Analysis of ",
      n_things(length(perf_use), "attribute"), " across ",
      n_things(final_rows, "response"), ", using ", imp_source_h, ". ",
      top_priority_name, " has the largest importance-performance gap ",
      "(priority score ", fmt_n(top_priority_score, 2),
      " standard deviations), and ", bottom_name, " the smallest. ",
      concentrate_clause, " at the grand-mean boundaries(importance ",
      fmt_n(imp_boundary, 2), ", performance ", fmt_n(perf_boundary, 2), "); ",
      "under ", alt_label, ", ",
      n_things(as.integer(n_reclassified), "attribute"), " ",
      vform(n_reclassified, "changes", "change"), " quadrant.",
      disagree_clause,
      " These are associations between attribute ratings and priority, not ",
      "proof that improving an attribute will move the outcome."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "quadrant_map", "gap_ranking",
        "attribute_table", "importance_comparison", "boundary_sensitivity",
        "action_plan"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows, rows_removed = rows_removed,
    n_imputed = n_imputed,
    attr_labels = attr_labels, perf_cols = perf_use, dropped_cols = dropped_cols,
    affix_stripped = affix_stripped, perf_names_raw = perf_names_raw,
    have_stated = have_stated, have_derived = have_derived, have_both = have_both,
    imp_source = imp_source, imp_source_h = imp_source_h,
    imp_source_short = imp_source_short,
    overall_name = overall_name,
    scale_detected = scale_detected, scale_min = scale_min,
    scale_max = scale_max, scale_mid = scale_mid,
    imp_boundary = imp_boundary, perf_boundary = perf_boundary,
    imp_boundary_alt = imp_boundary_alt, perf_boundary_alt = perf_boundary_alt,
    alt_label = alt_label,
    attrs_df = attrs_df,
    quadrant_points_df = quadrant_points_df, gap_ranking_df = gap_ranking_df,
    attribute_detail_df = attribute_detail_df,
    importance_sources_df = importance_sources_df,
    boundary_sensitivity_df = boundary_sensitivity_df,
    action_plan_df = action_plan_df,
    n_reclassified = n_reclassified, reclassified_names = reclassified_names,
    n_source_flips = n_source_flips, source_flip_names = source_flip_names,
    rank_rho = rank_rho, max_shift_name = max_shift_name, max_shift = max_shift,
    top_priority_name = top_priority_name, top_priority_score = top_priority_score,
    bottom_name = bottom_name, concentrate_names = concentrate_names,
    n_concentrate = n_concentrate, imp_tied = imp_tied,
    raw_gap_available = raw_gap_available,
    max_attr_cor = max_attr_cor, top_beta = top_beta,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell. Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No Signup Sign Up Free

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing