Executive Summary
The largest importance-performance gap across 4 attributes.
The short answer
Service Speed is the clear priority: it carries the widest gap between importance and performance, with an importance of 0.58 against a performance rating of 4.22, yielding a priority score of 1.57 standard deviations above the set mean. At the stated boundaries, only 1 attribute lands in Concentrate Here—Service Speed—making it the only quadrant-confirmed unmet need.
The detail
Across 1,200 responses and 4 attributes analysed, Service Speed leads the priority ranking with a score of 1.574 standard deviations. Its importance (derived correlation) is 0.58 and its mean performance rating is 4.22. The analysis applies grand-mean boundaries: importance 0.385 and performance 5.157. Under an alternative fixed boundary (importance 0.30, performance 5.0), 1 attribute reclassifies: Website Ease moves from Possible Overkill to Keep Up The Good Work, but Service Speed remains in Concentrate Here in both cases.
What this can't tell you
Importance is association with satisfaction, not a lever that will move it if improved. The map identifies where unmet need sits, not whether closing the gap will raise the outcome.
Analysis Overview
How 4 attributes are placed on the importance-performance map.
The short answer
This map shows where unmet customer need sits. It plots 4 attributes on two axes: how much each matters (derived from correlation with overall satisfaction) and how well it is currently rated. The four quadrants divide attributes by importance and performance, with boundaries drawn at the grand mean of each axis—importance 0.38 and performance 5.16. The core insight is that every placement is relative: the map ranks within this attribute set rather than measuring any absolute shortfall, and moving the boundary shifts which attributes demand attention.
The detail
Importance is measured as derived importance—the correlation of each attribute's rating with Overall Satisfaction across 1,200 responses. Performance is the mean rating on the underlying scale. The four quadrants are: Concentrate Here (important, weak), Keep Up The Good Work (important, strong), Low Priority (unimportant, weak), and Possible Overkill (unimportant, strong). The analysis included n_observations of 1,200 and n_attributes of 4. No stated-importance data were used; the map is built entirely from how strongly each attribute tracks satisfaction.
What this can't tell you
The boundary is a methodological choice, not a fact about the data. The grand-mean convention always creates relative winners and losers even when all attributes are genuinely strong or all weak. Derived importance shows association, not causation—an attribute that moves with satisfaction may not move it if changed.
Data Quality
Rows and attributes used, exclusions, imputation, and the rating scale.
The short answer
All 1,200 responses and all four attribute columns were usable; no rows were dropped and no ratings were imputed. The 0–10 rating scale was detected, setting the midpoint at 5.0 for the alternative boundary. Attribute names were shortened by removing shared wording, so "Satisfaction: Price" appears as "Price."
The detail
Initial load: 1,200 rows. Final rows: 1,200 (rows removed: 0). No attribute rating was missing, so n_imputed = 0. Four attributes carried to the map: Price, Service Speed, Product Quality, Website Ease. The rating scale runs from 0 to 10, establishing 5.0 as the scale midpoint used in the fixed-boundary alternative.
What this can't tell you
Data quality was clean for this analysis; no imputation decisions or row exclusions created hidden assumptions. The only caveat is that derived importance depends on the quality of the Overall Satisfaction measure—if that item itself was misunderstood or had high missingness in the source data, the correlations would be unreliable, but that check lies outside the scope of this preprocessing report.
Importance vs Performance Map
Each attribute placed by derived importance against current performance, split into four quadrants.
The short answer
Service Speed sits alone in Concentrate Here (upper left): it matters more than the average attribute here but is rated well below average. Website Ease occupies Possible Overkill (upper right), rated well despite below-average importance. Product Quality anchors Keep Up The Good Work (upper right), both important and well-rated. Price lands in Low Priority (lower left), neither important nor well-rated. The pattern is not a tight trend—attributes are scattered across all four quadrants, with Service Speed the clear outlier on the left side.
The detail
All four quadrants are populated. Reference lines sit at grand means: importance 0.3849, performance 5.1573. Service Speed: importance 0.577, performance 4.218. Website Ease: importance 0.372, performance 5.418. Product Quality: importance 0.465, performance 7.35. Price: importance 0.126, performance 3.644. Attributes near a dividing line are not meaningfully different from those just across it; the corners carry the signal. Service Speed is the only point in the upper left, making it the isolated priority.
What this can't tell you
The map shows correlation between attribute ratings and Overall Satisfaction, not causal relationships. An attribute's position depends on the boundary choice—moving the cross-hair changes which attributes are flagged as priorities. The analysis cannot distinguish between halo effects (satisfied respondents rate everything higher) and true attribute importance, though inter-attribute correlations of 0.05 or less suggest each attribute's signal is largely independent.
Priority Ranking
Attributes ranked by the size of the importance-performance gap.
The short answer
Service Speed dominates the priority ranking by a wide margin. It is the only attribute with a positive gap score, meaning it is the only one where importance exceeds performance. The next three attributes all show negative scores, indicating they either deliver more than their importance warrants or both underperform and matter less.
The detail
Service Speed leads at a priority score of 1.574 standard deviations. Website Ease follows at -0.225, Price at -0.426, and Product Quality trails at -0.922. The priority score is calculated as how far above average an attribute sits on importance minus how far above average it sits on performance, both in standard deviations across the 4 attributes. Only 1 attribute has a positive score. The score is relative to this attribute set and will re-rank if attributes are added or dropped.
What this can't tell you
The score does not measure any absolute shortfall—it prioritizes within this list rather than against an external standard. Adding or removing attributes will shift the boundaries and may move attributes across quadrants.
Attribute Detail
Importance, performance, gap, and quadrant for each of 4 attributes.
| Attribute | Importance | Performance | Priority Score | Quadrant |
|---|---|---|---|---|
| Service Speed | 0.577 | 4.218 | 1.574 | Concentrate Here |
| Website Ease | 0.372 | 5.418 | -0.225 | Possible Overkill |
| Price | 0.126 | 3.644 | -0.426 | Low Priority |
| Product Quality | 0.465 | 7.35 | -0.922 | Keep Up The Good Work |
The short answer
Service Speed is the decisive finding: it is the only attribute in Concentrate Here, with an importance of 0.577 and a performance rating of 4.218. Product Quality, the second-strongest attribute on importance (0.465), delivers the highest performance (7.35) and sits in Keep Up The Good Work. The gap between Service Speed's importance and its performance is the largest in the set.
The detail
Service Speed: importance 0.577, performance 4.218, priority score 1.574, quadrant Concentrate Here. Website Ease: importance 0.372, performance 5.418, priority score -0.225, quadrant Possible Overkill. Price: importance 0.126, performance 3.644, priority score -0.426, quadrant Low Priority. Product Quality: importance 0.465, performance 7.35, priority score -0.922, quadrant Keep Up The Good Work. One attribute sits in Concentrate Here and one in Keep Up The Good Work; the other two are in lower-priority quadrants.
What this can't tell you
Two attributes can share a quadrant while sitting at opposite ends of it. The quadrant assignment depends on the boundary chosen; Website Ease moves under the alternative boundary convention.
Stated vs Derived Importance
Where the two ways of measuring importance agree, and where they contradict each other.
| Attribute | Stated Importance | Derived Importance | Derived Evidence | Stated Rank | Derived Rank | Rank Shift |
|---|---|---|---|---|---|---|
| Service Speed | — | 0.577 | p < 0.001 | — | 1 | — |
| Website Ease | — | 0.372 | p < 0.001 | — | 3 | — |
| Price | — | 0.126 | p < 0.001 | — | 4 | — |
| Product Quality | — | 0.465 | p < 0.001 | — | 2 | — |
The short answer
Only derived importance is available—no stated-importance columns were mapped, so the analysis cannot check whether customers themselves would call these attributes important. Derived importance is inferred from each attribute's correlation with Overall Satisfaction. Service Speed ranks first (0.577 correlation, p < 0.001), Product Quality second (0.465, p < 0.001), Website Ease third (0.372, p < 0.001), and Price fourth (0.126, p < 0.001).
The detail
Stated importance: not available. Derived importance and derived evidence (all p < 0.001): Service Speed 0.577, Website Ease 0.372, Price 0.126, Product Quality 0.465. Derived ranks: Service Speed 1, Product Quality 2, Website Ease 3, Price 4. No rank shift can be computed because stated importance does not exist. Halo effects (respondents who are satisfied overall rating every attribute higher) could inflate all correlations at once, but inter-attribute correlations are at most 0.05 here, so each attribute's score is close to its own separate contribution.
What this can't tell you
The analysis cannot compare stated and derived importance to see where customer perception diverges from behavior. Mapping an attribute-level importance rating per respondent would enable that cross-check and is where this analysis is most useful. Derived importance is association, not causation: it shows which attributes move with Overall Satisfaction, not which ones would move it if changed.
Boundary Sensitivity
How the quadrant picture changes under the other boundary convention.
| Attribute | Quadrant Grand Mean | Quadrant Alternative | Changed |
|---|---|---|---|
| Service Speed | Concentrate Here | Concentrate Here | no |
| Website Ease | Possible Overkill | Keep Up The Good Work | yes |
| Price | Low Priority | Low Priority | no |
| Product Quality | Keep Up The Good Work | Keep Up The Good Work | no |
The short answer
Service Speed stays in Concentrate Here under both boundary conventions, confirming it as a robust priority. Website Ease is the only attribute that moves: it shifts from Possible Overkill (grand-mean boundary) to Keep Up The Good Work (fixed alternative boundary). This sensitivity means any recommendation that rests solely on Website Ease's quadrant placement is a recommendation about the boundary, not the data.
The detail
At the grand-mean boundaries (importance 0.38, performance 5.16), Service Speed remains in Concentrate Here. Website Ease moves from Possible Overkill to Keep Up The Good Work under the fixed alternative (importance 0.30, performance 5.0). Price and Product Quality do not move. The grand-mean boundary always creates relative winners and losers within the set; the fixed alternative is independent of this attribute set but ignores how this particular distribution is shaped.
What this can't tell you
Neither boundary convention is correct. The choice between them depends on your use case: use grand means to rank within your own list; use the fixed alternative when you need the map to remain consistent across surveys or over time.
Prioritized Action List
The recommended next step for each quadrant, ordered by urgency.
| Quadrant | N Attributes | Attributes | Action |
|---|---|---|---|
| Concentrate Here | 1 | Service Speed | Fix these first. They matter more than average and are rated below average, so this is where the largest unmet need sits. |
| Keep Up The Good Work | 1 | Product Quality | Protect these. They matter and are already rated well — they are the strengths worth defending rather than improving further. |
| Low Priority | 1 | Price | Leave these alone for now. They are rated below average but also matter less than average, so weak scores here cost comparatively little. |
| Possible Overkill | 1 | Website Ease | Consider easing off. These are rated well but matter less than average, so effort spent here may be buying little. |
The short answer
Start with Service Speed in Concentrate Here—it is the only attribute where importance clearly exceeds performance, making it the largest unmet need. Protect Product Quality in Keep Up The Good Work rather than investing further. Leave Price alone in Low Priority and consider easing off Website Ease in Possible Overkill. Before committing budget, verify that the boundary-sensitivity finding (1 attribute moves under the alternative convention) does not change your strategic choice.
The detail
Concentrate Here (Service Speed, n_attributes 1): Fix first—largest unmet need. Keep Up The Good Work (Product Quality, n_attributes 1): Protect rather than improve. Low Priority (Price, n_attributes 1): Leave alone while unimportant. Possible Overkill (Website Ease, n_attributes 1): Consider releasing effort. All recommendations rest on association between attribute ratings and Overall Satisfaction; treat each as a hypothesis worth testing.
What this can't tell you
Derived importance is correlation, not proof of causation. Improving Service Speed may or may not raise Overall Satisfaction; the map shows where the unmet need is, not that closing it will move the outcome.
Methodology
Statistical methodology and diagnostics for Importance-Performance Gap Analysis
Statistical Method
Standard-library analysis: what matters most that we do worst? Classic Importance-Performance Analysis on survey data. Map your attribute satisfaction columns and either an importance rating per attribute or an overall satisfaction score, and get the four-quadrant map (Concentrate Here, Keep Up The Good Work, Low Priority, Possible Overkill), every attribute ranked by the size of its importance-performance gap, and a prioritized action list. When both stated importance and an overall score are available, both are computed and the disagreement between them is reported rather than one being picked silently — and because the quadrant boundary is a methodological choice, the analysis states the values it used and names every attribute that would move under the other convention.
- Each row is one respondent, with one rating per attribute
- Attribute ratings are numeric or cleanly convertible, and on a common scale
- Attributes are rated on the same scale as stated importance, if the raw gap is to be read in scale points
- For derived importance, the overall score reflects the attributes measured rather than something outside the survey
- Quadrant placement is relative to the attributes you mapped — adding or removing an attribute moves the boundaries and can move other attributes across them
- Derived importance is a correlation with the overall score, not proof that improving an attribute will raise it
- Derived importance is inflated by halo, where a respondent who is satisfied overall rates every attribute higher; when the attributes are correlated with each other, a bivariate score shares credit between them
- Stated importance suffers ceiling effects — respondents often rate nearly everything as important, which flattens the horizontal axis
Analysis Code
Complete R source code for this analysis
Importance-Performance Gap Analysis
What matters most that we do worst? Classic Importance-Performance Analysis (IPA): every attribute is placed on a four-quadrant map by how important it is against how well it is currently rated, ranked by the gap between the two, and turned into a prioritized action list.
Why This Method?
Ranking attributes by performance alone tells you where you are weak but not whether anyone cares. Ranking by importance alone tells you what matters but not where you are failing. Crossing the two is the whole point: the attributes that are important AND under-performing are the ones worth money, and the map makes that visible in one picture.
Two ways to measure importance — and they disagree
STATED importance is what respondents said mattered. DERIVED importance is inferred from how strongly each attribute tracks an overall satisfaction score. They routinely disagree, and the disagreement is informative rather than an error: people under-report what actually drives them and over-report what they think they should care about. This module detects which inputs it was given, computes BOTH when both are available, and shows where they contradict each other instead of quietly picking one.
What This Analysis Covers
- Attribute-level importance and performance on a four-quadrant map
- The priority ranking by importance-performance gap
- Stated versus derived importance, and where they disagree
- How the quadrant picture changes under the other boundary convention
- A per-quadrant action list
Sibling tool
standard_key_drivers answers "what drives satisfaction?" — it ranks drivers by derived importance and reports model fit. This tool answers "what should we fix first?" — it takes importance as given (stated or derived), crosses it with current performance, and produces a prioritized gap list. Same vocabulary, different question.
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {performance_1..N, importance_1..N, overall}. All narrative is derived from the user's own column names and computed values. Derived importance is CORRELATIONAL — it reports association, never proven causation.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Longest shared leading run of whole words.
n_lead <- 0
repeat {
k <- n_lead + 1
if (any(sapply(parts, length) <= k)) break
w <- sapply(parts, function(p) p[k])
if (length(unique(w)) != 1) break
n_lead <- k
}Longest shared trailing run of whole words.
n_tail <- 0
repeat {
k <- n_tail + 1
if (any(sapply(parts, length) <= n_lead + k)) break
w <- sapply(parts, function(p) p[length(p) - k + 1])
if (length(unique(w)) != 1) break
n_tail <- k
}
if (n_lead == 0 && n_tail == 0) return(labels)
out <- sapply(parts, function(p) {
keep <- p[(n_lead + 1):(length(p) - n_tail)]
trimws(paste(keep, collapse = " "))
}, USE.NAMES = FALSE)Refuse the strip if it empties or collides any label.
if (any(nchar(out) == 0) || anyDuplicated(out) > 0) return(labels)
out
}Step 1: Row accounting + semantic column discovery
initial_rows <- nrow(df)
perf_cols <- grep("^performance_[0-9]+$", names(df), value = TRUE)
perf_cols <- perf_cols[order(as.integer(sub("^performance_", "", perf_cols)))]
imp_cols <- grep("^importance_[0-9]+$", names(df), value = TRUE)
imp_cols <- imp_cols[order(as.integer(sub("^importance_", "", imp_cols)))]
has_overall_col <- "overall" %in% names(df)
if (length(perf_cols) == 0) {
stop(paste0("column_mapping must map at least ", MIN_ATTRS,
" performance columns(performance_1, performance_2, ...) — ",
"one per attribute you rate."))
}
perf_names_raw <- humanize_semantic(perf_cols, col_map)
imp_names_raw <- if (length(imp_cols) > 0) humanize_semantic(imp_cols, col_map) else character(0)
overall_name <- if (has_overall_col) humanize_semantic("overall", col_map) else NA_character_Attribute labels come from the PERFORMANCE columns, with any wording shared by all of them removed ("Satisfaction: Price" -> "Price").
attr_labels_all <- strip_common_affix(perf_names_raw)
affix_stripped <- !identical(attr_labels_all, perf_names_raw)
names(attr_labels_all) <- perf_colsStated importance is paired to performance BY POSITION: importance_1 describes the same attribute as performance_1. A partial or mismatched set cannot be paired safely, so it is refused rather than guessed.
paired_stated <- length(imp_cols) > 0 && length(imp_cols) == length(perf_cols) &&
identical(sub("^importance_", "", imp_cols), sub("^performance_", "", perf_cols))
if (length(imp_cols) > 0 && !paired_stated) {
stop(paste0(
"Stated importance must be mapped one-for-one with performance: ",
n_things(length(imp_cols), "importance column"), " were mapped(",
paste(imp_names_raw, collapse = ", "), ") against ",
n_things(length(perf_cols), "performance column"), " (",
paste(perf_names_raw, collapse = ", "),
"). Map the same attributes in the same order, or map none and supply ",
"an overall satisfaction column instead."))
}Step 2: Coerce every mapped rating to numeric (95% rule)
A column that will not convert cleanly, or that has no usable values, is dropped and reported rather than silently coerced to NA.
dropped_cols <- character(0)
coerce <- function(dd, cc) {
v <- dd[[cc]]
if (is.numeric(v)) return(v)
conv <- suppressWarnings(as.numeric(as.character(v)))
n_orig <- sum(!is.na(v) & as.character(v) != "")
if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) conv else NULL
}
bad_perf <- character(0)
for (pc in perf_cols) {
conv <- coerce(df, pc)
if (is.null(conv)) { bad_perf <- c(bad_perf, pc); next }
df[[pc]] <- conv
}
bad_imp <- character(0)
if (paired_stated) {
for (ic in imp_cols) {
conv <- coerce(df, ic)
if (is.null(conv)) { bad_imp <- c(bad_imp, ic); next }
df[[ic]] <- conv
}
}
overall_usable <- FALSE
if (has_overall_col) {
conv <- coerce(df, "overall")
if (!is.null(conv)) { df$overall <- conv; overall_usable <- TRUE }
}Step 3: Drop constant / all-missing attributes — and their partner
An attribute is only usable if BOTH sides that will be plotted survive.
for (i in seq_along(perf_cols)) {
pc <- perf_cols[i]
if (pc %in% bad_perf) next
v <- df[[pc]]
if (all(is.na(v))) { bad_perf <- c(bad_perf, pc); next }
if (isTRUE(stats::var(v, na.rm = TRUE) == 0) ||
is.na(stats::var(v, na.rm = TRUE))) {
bad_perf <- c(bad_perf, pc)
}
}
drop_idx <- which(perf_cols %in% bad_perf)
if (paired_stated) drop_idx <- union(drop_idx, which(imp_cols %in% bad_imp))
keep_idx <- setdiff(seq_along(perf_cols), drop_idx)
if (length(drop_idx) > 0) {
dropped_cols <- unname(attr_labels_all[perf_cols[drop_idx]])
}
perf_use <- perf_cols[keep_idx]
imp_use <- if (paired_stated) imp_cols[keep_idx] else character(0)
attr_labels <- unname(attr_labels_all[perf_use])
if (length(perf_use) < MIN_ATTRS) {
kept_note <- if (length(perf_use) > 0)
paste0(" The attributes read were: ", paste(attr_labels, collapse = ", "), ".")
else
paste0(" The columns mapped were: ", paste(perf_names_raw, collapse = ", "), ".")
stop(sprintf(paste0(
"Importance-Performance Analysis needs at least %d usable attributes; ",
"only %d of the %d mapped performance columns survived cleaning%s.%s ",
"A four-quadrant map of fewer than %d attributes is not meaningful."),
MIN_ATTRS, length(perf_use), length(perf_cols),
if (length(dropped_cols) > 0)
paste0(" (excluded as constant, empty, or non-numeric: ",
paste(dropped_cols, collapse = ", "), ")") else "",
kept_note, MIN_ATTRS))
}Step 4: Which importance sources do we actually have?
have_stated <- paired_stated && length(imp_use) == length(perf_use)
have_derived <- overall_usable && sum(!is.na(df$overall)) >= MIN_ROWS
if (!have_stated && !have_derived) {
stop(paste0(
"Importance-Performance Analysis needs importance from somewhere. ",
"Either map an importance column for each attribute(",
paste(attr_labels, collapse = ", "),
"), or map an overall satisfaction column so importance can be ",
"derived from how each attribute tracks it."))
}Step 5: Rows — derived importance needs a usable overall value
if (have_derived) df <- df[!is.na(df$overall), , drop = FALSE]
keep_row <- rowSums(!is.na(df[, perf_use, drop = FALSE])) > 0
df <- df[keep_row, , drop = FALSE]
final_rows <- nrow(df)
rows_removed <- initial_rows - final_rows
if (final_rows < MIN_ROWS) {
stop(sprintf(paste0(
"Only %d usable responses remain after cleaning%s — at least %d are ",
"required before an importance-performance map means anything. ",
"Attributes read: %s."),
final_rows,
if (have_derived) sprintf(" (rows with no '%s' value are dropped)", overall_name) else "",
MIN_ROWS, paste(attr_labels, collapse = ", ")))
}Step 6: Impute remaining missing ratings with each column's median
n_imputed <- 0L
for (cc in c(perf_use, imp_use)) {
v <- df[[cc]]
miss <- is.na(v)
if (any(miss)) {
med <- stats::median(v, na.rm = TRUE)
if (!is.na(med)) { v[miss] <- med; n_imputed <- n_imputed + sum(miss) }
df[[cc]] <- v
}
}Step 7: Rating scale — needed for the midpoint boundary convention
rating_vals <- unlist(df[, c(perf_use, imp_use), drop = FALSE], use.names = FALSE)
rating_vals <- rating_vals[is.finite(rating_vals)]
vmin <- min(rating_vals); vmax <- max(rating_vals)
scale_detected <- vmin >= 0 && vmax <= 10
if (scale_detected) {
scale_min <- if (vmin < 1) 0 else 1
scale_max <- if (vmax <= 5) 5 else if (vmax <= 7) 7 else 10
} else {
scale_min <- vmin; scale_max <- vmax
}
scale_mid <- (scale_min + scale_max) / 2Step 8: Performance — the mean rating per attribute
perf_mean <- sapply(perf_use, function(pc) mean(df[[pc]], na.rm = TRUE))
perf_mean <- as.numeric(perf_mean)Step 9: Stated importance — the mean importance rating per attribute
stated_imp <- if (have_stated) {
as.numeric(sapply(imp_use, function(ic) mean(df[[ic]], na.rm = TRUE)))
} else rep(NA_real_, length(perf_use))Step 10: Derived importance — how each attribute tracks the overall
score. Bivariate Pearson correlation is the standard derived-importance statistic; a joint standardized regression is also fitted purely to disclose how much of that association is shared rather than unique.
derived_imp <- rep(NA_real_, length(perf_use))
derived_p <- rep(NA_real_, length(perf_use))
max_attr_cor <- NA_real_
top_beta <- NA_real_
if (have_derived) {
for (i in seq_along(perf_use)) {
ct <- tryCatch(stats::cor.test(df[[perf_use[i]]], df$overall),
error = function(e) NULL)
if (!is.null(ct) && is.finite(ct$estimate)) {
derived_imp[i] <- as.numeric(ct$estimate)
derived_p[i] <- as.numeric(ct$p.value)
} else {
derived_imp[i] <- 0
}
}
if (length(perf_use) >= 2) {
cm <- suppressWarnings(stats::cor(df[, perf_use, drop = FALSE],
use = "pairwise.complete.obs"))
cm[!is.finite(cm)] <- 0
diag(cm) <- 0
max_attr_cor <- max(abs(cm))
}
zfit <- tryCatch({
zd <- as.data.frame(scale(df[, c("overall", perf_use), drop = FALSE]))
stats::lm(overall ~ ., data = zd)
}, error = function(e) NULL)
if (!is.null(zfit)) {
bb <- stats::coef(zfit)[perf_use]
ok <- which(!is.na(derived_imp))
if (length(ok) > 0) {
lead <- ok[order(-abs(derived_imp[ok]))][1]
top_beta <- unname(bb[lead])
}
}
}Step 11: Which importance source drives the primary map?
Stated is preferred when present because it shares the performance scale, which makes the raw gap directly readable. The choice is stated in the prose and the other source is reported alongside it — never silently dropped.
req_source <- tolower(as.character(params$importance_source %||% "auto"))
imp_source <- if (req_source == "derived" && have_derived) "derived"
else if (req_source == "stated" && have_stated) "stated"
else if (have_stated) "stated" else "derived"
importance <- if (imp_source == "stated") stated_imp else derived_imp
imp_source_h <- if (imp_source == "stated")
"stated importance(the average importance rating respondents gave)"
else
paste0("derived importance(each attribute's correlation with ", overall_name, ")")Short form, for sentences that already carry their own parentheses.
imp_source_short <- if (imp_source == "stated") "stated importance" else "derived importance"Step 12: Boundaries — the methodological choice that moves the map
Primary: the data-driven grand mean of each axis (the cross-hair sits at the average attribute). Alternative: the scale midpoint, which is fixed and independent of this attribute set. For derived importance the scale midpoint does not exist, so the alternative is the conventional moderate-association cut of 0.30.
imp_boundary <- mean(importance, na.rm = TRUE)
perf_boundary <- mean(perf_mean, na.rm = TRUE)
perf_boundary_alt <- scale_mid
imp_boundary_alt <- if (imp_source == "stated") scale_mid else DERIVED_ALT_CUT
alt_label <- if (imp_source == "stated") {
paste0("the scale midpoint(", fmt_n(scale_mid, 1), " on a ",
fmt_n(scale_min, 0), "-to-", fmt_n(scale_max, 0), " scale)")
} else {
paste0("a fixed correlation cut of ", fmt_n(DERIVED_ALT_CUT, 2),
" on importance and the scale midpoint(", fmt_n(scale_mid, 1),
") on performance")
}
quad_of <- function(imp, perf, bi, bp) {
ifelse(imp >= bi & perf < bp, "Concentrate Here",
ifelse(imp >= bi & perf >= bp, "Keep Up The Good Work",
ifelse(imp < bi & perf < bp, "Low Priority",
"Possible Overkill")))
}
quadrant <- quad_of(importance, perf_mean, imp_boundary, perf_boundary)
quadrant_alt <- quad_of(importance, perf_mean, imp_boundary_alt, perf_boundary_alt)
reclassified <- quadrant != quadrant_alt
n_reclassified <- sum(reclassified)
reclassified_names <- attr_labels[reclassified]Step 13: The gap. Two measures, both computed.
priority_score standardizes each axis ACROSS THE ATTRIBUTE SET, so it works whether importance is a rating or a correlation. raw_gap is the plain importance-minus-performance difference in scale points, and is only defined when both axes are on the same rating scale.
z_of <- function(v) {
s <- stats::sd(v, na.rm = TRUE)
if (!is.finite(s) || s == 0) rep(0, length(v))
else (v - mean(v, na.rm = TRUE)) / s
}
priority_score <- z_of(importance) - z_of(perf_mean)
imp_range <- diff(range(importance, na.rm = TRUE))
scale_span <- if (is.finite(scale_max - scale_min) && (scale_max - scale_min) > 0)
scale_max - scale_min else NA_real_
imp_tied <- if (imp_source == "stated" && is.finite(scale_span))
imp_range < 0.05 * scale_span else FALSE
raw_gap_available <- imp_source == "stated"
raw_gap <- if (raw_gap_available) stated_imp - perf_mean else rep(NA_real_, length(perf_use))Step 14: Master attribute frame, ranked by the gap
attrs_df <- data.frame(
attribute = attr_labels,
importance = round(importance, 3),
performance = round(perf_mean, 3),
priority_score = round(priority_score, 3),
raw_gap = round(raw_gap, 3),
quadrant = quadrant,
quadrant_alt = quadrant_alt,
stated_importance = round(stated_imp, 3),
derived_importance = round(derived_imp, 3),
derived_p = derived_p,
stringsAsFactors = FALSE
)
attrs_df <- attrs_df[order(-attrs_df$priority_score, attrs_df$attribute), , drop = FALSE]
rownames(attrs_df) <- NULL
top_priority_name <- attrs_df$attribute[1]
top_priority_score <- attrs_df$priority_score[1]
bottom_name <- attrs_df$attribute[nrow(attrs_df)]
concentrate_names <- attrs_df$attribute[attrs_df$quadrant == "Concentrate Here"]
n_concentrate <- length(concentrate_names)Step 15: Stated versus derived — ranks, shifts, and quadrant flips
have_both <- have_stated && have_derived
rank_rho <- NA_real_
n_source_flips <- 0L
source_flip_names <- character(0)
max_shift_name <- NA_character_
max_shift <- NA_real_
stated_rank <- rep(NA_real_, nrow(attrs_df))
derived_rank <- rep(NA_real_, nrow(attrs_df))
if (have_stated) stated_rank <- rank(-attrs_df$stated_importance, ties.method = "min")
if (have_derived) derived_rank <- rank(-attrs_df$derived_importance, ties.method = "min")
rank_shift <- if (have_both) abs(stated_rank - derived_rank) else rep(NA_real_, nrow(attrs_df))
if (have_both) {
rank_rho <- suppressWarnings(stats::cor(attrs_df$stated_importance,
attrs_df$derived_importance,
method = "spearman"))
if (!is.finite(rank_rho)) rank_rho <- NA_real_
ord <- order(-rank_shift, attrs_df$attribute)
max_shift_name <- attrs_df$attribute[ord[1]]
max_shift <- rank_shift[ord[1]]
q_stated <- quad_of(attrs_df$stated_importance, attrs_df$performance,
mean(attrs_df$stated_importance), perf_boundary)
q_derived <- quad_of(attrs_df$derived_importance, attrs_df$performance,
mean(attrs_df$derived_importance), perf_boundary)
flips <- q_stated != q_derived
n_source_flips <- sum(flips)
source_flip_names <- attrs_df$attribute[flips]
}Derived importance is an estimate, so it carries its own evidence: an attribute whose correlation is not distinguishable from zero must not be read as "unimportant with confidence".
derived_evidence <- if (have_derived) {
sapply(seq_len(nrow(attrs_df)), function(i) {
pv <- attrs_df$derived_p[i]
if (is.na(pv)) return("not available")
paste0(fmt_p(pv), if (pv < 0.05) "" else " (not distinguishable from zero)")
})
} else rep(NA_character_, nrow(attrs_df))
importance_sources_df <- data.frame(
attribute = attrs_df$attribute,
stated_importance = attrs_df$stated_importance,
derived_importance = attrs_df$derived_importance,
derived_evidence = derived_evidence,
stated_rank = stated_rank,
derived_rank = derived_rank,
rank_shift = rank_shift,
stringsAsFactors = FALSE
)
boundary_sensitivity_df <- data.frame(
attribute = attrs_df$attribute,
quadrant_grand_mean = attrs_df$quadrant,
quadrant_alternative = attrs_df$quadrant_alt,
changed = ifelse(attrs_df$quadrant != attrs_df$quadrant_alt,
"yes", "no"),
stringsAsFactors = FALSE
)
action_plan_df <- do.call(rbind, lapply(QUADS, function(q) {
members <- attrs_df$attribute[attrs_df$quadrant == q]
data.frame(
quadrant = q,
n_attributes = length(members),
attributes = if (length(members) > 0) paste(members, collapse = ", ") else "None",
action = unname(QUAD_ACTIONS[q]),
stringsAsFactors = FALSE
)
}))
rownames(action_plan_df) <- NULLStep 17: KPI metrics
metrics <- list(
`Responses` = final_rows,
`Attributes Analysed` = length(perf_use),
`Importance Source` = if (imp_source == "stated") "stated" else "derived",
`Top Priority` = top_priority_name,
`Top Priority Score` = round(top_priority_score, 2),
`Concentrate Here` = as.integer(n_concentrate),
`Importance Boundary` = round(imp_boundary, 3),
`Performance Boundary` = round(perf_boundary, 3),
`Reclassified Under Alternative Boundary` = as.integer(n_reclassified)
)Step 18: json_output machine channel
concentrate_clause <- if (n_concentrate > 0) {
paste0(n_things(n_concentrate, "attribute"), " ",
vform(n_concentrate, "falls", "fall"), " in Concentrate Here(",
paste(concentrate_names, collapse = ", "), ")")
} else {
"No attribute falls in Concentrate Here"
}
disagree_clause <- if (have_both) {
paste0(" Stated and derived importance rank the attributes differently ",
"(Spearman rank correlation ", fmt_n(rank_rho, 2), "); ",
max_shift_name, " moves the most, by ",
n_things(as.integer(max_shift), "rank position"),
", and ", n_things(as.integer(n_source_flips), "attribute"), " ",
vform(n_source_flips, "changes", "change"),
" quadrant depending on which importance source is used.")
} else if (imp_source == "derived") {
paste0(" Importance was derived from ", overall_name,
" because no stated-importance columns were mapped.")
} else {
" Importance is as stated by respondents; no overall satisfaction column was mapped to cross-check it."
}
json_output <- list(
answer = paste0(
"Importance-Performance Analysis of ",
n_things(length(perf_use), "attribute"), " across ",
n_things(final_rows, "response"), ", using ", imp_source_h, ". ",
top_priority_name, " has the largest importance-performance gap ",
"(priority score ", fmt_n(top_priority_score, 2),
" standard deviations), and ", bottom_name, " the smallest. ",
concentrate_clause, " at the grand-mean boundaries(importance ",
fmt_n(imp_boundary, 2), ", performance ", fmt_n(perf_boundary, 2), "); ",
"under ", alt_label, ", ",
n_things(as.integer(n_reclassified), "attribute"), " ",
vform(n_reclassified, "changes", "change"), " quadrant.",
disagree_clause,
" These are associations between attribute ratings and priority, not ",
"proof that improving an attribute will move the outcome."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "quadrant_map", "gap_ranking",
"attribute_table", "importance_comparison", "boundary_sensitivity",
"action_plan"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows, rows_removed = rows_removed,
n_imputed = n_imputed,
attr_labels = attr_labels, perf_cols = perf_use, dropped_cols = dropped_cols,
affix_stripped = affix_stripped, perf_names_raw = perf_names_raw,
have_stated = have_stated, have_derived = have_derived, have_both = have_both,
imp_source = imp_source, imp_source_h = imp_source_h,
imp_source_short = imp_source_short,
overall_name = overall_name,
scale_detected = scale_detected, scale_min = scale_min,
scale_max = scale_max, scale_mid = scale_mid,
imp_boundary = imp_boundary, perf_boundary = perf_boundary,
imp_boundary_alt = imp_boundary_alt, perf_boundary_alt = perf_boundary_alt,
alt_label = alt_label,
attrs_df = attrs_df,
quadrant_points_df = quadrant_points_df, gap_ranking_df = gap_ranking_df,
attribute_detail_df = attribute_detail_df,
importance_sources_df = importance_sources_df,
boundary_sensitivity_df = boundary_sensitivity_df,
action_plan_df = action_plan_df,
n_reclassified = n_reclassified, reclassified_names = reclassified_names,
n_source_flips = n_source_flips, source_flip_names = source_flip_names,
rank_rho = rank_rho, max_shift_name = max_shift_name, max_shift = max_shift,
top_priority_name = top_priority_name, top_priority_score = top_priority_score,
bottom_name = bottom_name, concentrate_names = concentrate_names,
n_concentrate = n_concentrate, imp_tied = imp_tied,
raw_gap_available = raw_gap_available,
max_attr_cor = max_attr_cor, top_beta = top_beta,
metrics = metrics, json_output = json_output
)
}