GuideGrowth, Marketing & Sales

How to build a customer health score that predicts churn and prompts action

A health score is only useful if it predicts a defined outcome and changes what someone does on Monday morning. This guide starts with the outcome, moves through signal selection and a rules-based score validated against past renewals, then covers when a predictive churn model is worth building, how to explain scores to account managers and how to prove the program works.

Reviewed 7 min read

On this page
  1. Choose the outcome the score predicts
  2. Candidate signals worth testing for churn risk
  3. Building and validating a rules-based score
  4. Rules-based score or predictive churn model
  5. Modeling traps that make churn models look better than they are
  6. Explaining a score so account managers will use it
  7. From risk tier to playbook to measured outcome
  8. Proving impact with a holdout instead of a trend line
  9. Questions and answers

Choose the outcome the score predicts

Different outcomes reward different signals. Pick one primary outcome and a horizon, such as the next renewal date, and write both down.

Logo churn
A customer ends the relationship entirely. Simple to label, but treats a small account the same as a large one.
Gross revenue churn
Recurring revenue lost to cancellation and downgrades, which weights accounts by what they pay.
Contraction
Fewer seats, a lower tier or removed products while the customer stays. Often an early warning of logo churn.
Non-renewal at term
The outcome observed at each contract end date. Natural for annual contracts, but labels arrive only once per term.

Candidate signals worth testing for churn risk

Treat each as a hypothesis to test against history, not as a known predictor.

0 of 7 checked

Building and validating a rules-based score

  1. Freeze a snapshot before each past renewal

    For every renewal in your history, rebuild what the signals looked like at a fixed point before the decision, such as a full quarter earlier. Using today's values for past accounts leaks the outcome into the score.

    Output
    A snapshot table with one row per past renewal
  2. Pick a handful of signals and set weights

    Choose the signals that differ most between renewed and churned accounts in the snapshot, agree weights with experienced account managers, and keep the list short enough to explain in a sentence.

    Output
    A scoring rule and its rationale
  3. Set thresholds for risk tiers

    Translate scores into red, amber and green tiers sized to what the team can act on. A red tier containing half the book is a list, not a priority.

    Output
    Tier thresholds tied to team capacity
  4. Backtest against actual outcomes

    Check what share of churned accounts the red tier would have caught and what share of red accounts actually churned. Both numbers matter: catching every churner by flagging everyone helps nobody.

    Output
    A backtest report by tier
  5. Review the misses with the people who knew the accounts

    Walk through churned accounts the score rated green and renewed accounts it rated red. Misses usually reveal a missing signal, such as a merger or budget freeze, or a data quality problem.

    Output
    A revised rule and a list of data fixes
  6. Version and document the score

    Record the definition, data sources, weights and backtest results, and change them only through a reviewed release so trends over time stay comparable.

    Output
    A versioned score specification

Rules-based score or predictive churn model

ConsiderationRules-based scorePredictive model
Data neededEnough history to sanity-check weightsEnough churn events to learn patterns and still hold out a test set
TransparencyAnyone can read the ruleNeeds per-account explanations to earn trust
Handles interactions between signalsPoorly; weights add up independentlyWell, if the data supports it
MaintenanceOccasional reweightingMonitoring for drift and regular retraining
Best fitSmaller books, few churn events, a new programLarge books with plenty of labeled history and data engineering support

Most teams should start in the highlighted column and graduate when the backtest shows clear limits.

Modeling traps that make churn models look better than they are

Label leakage

Early signalA feature such as a cancellation ticket or a renewal opportunity marked lost dominates the model.

MitigationBuild features only from data available at the snapshot date and exclude anything recorded as part of the churn process.

Accuracy that hides failure

Early signalHigh accuracy because most customers renew, while the model finds few churners.

MitigationEvaluate precision among the top-ranked accounts the team can actually contact, and check calibration so a predicted risk matches observed churn.

Survivorship in the training set

Early signalTraining only on current customers, so the accounts that churned early are missing.

MitigationInclude every account that existed at each snapshot, churned or not.

Drift after product or pricing changes

Early signalThe score's backtest performance falls on recent cohorts.

MitigationMonitor performance by cohort and retrain or reweight after major changes.

Explaining a score so account managers will use it

A number without a reason gets ignored. Show the two or three factors that moved each account's score, in plain language, next to the account record: usage of the core workflow down sharply this month, sponsor left, two unresolved high-severity tickets. For models, per-account attribution methods such as SHAP values can produce those factors, but they still need translating into business terms.

Let account managers flag a score they disagree with and give a reason. Those overrides are labeled data about what the score misses, and reviewing them regularly keeps the people closest to customers invested in improving it.

From risk tier to playbook to measured outcome

01Signals refresh02Score and tier03Playbook triggered04Owner acts05Outcome logged
  1. Signals refresh

    Usage, support, billing and relationship data update on a set schedule.

  2. Score and tier

    Each account gets a tier and its top contributing factors.

  3. Playbook triggered

    A tier and factor combination maps to a specific play, such as an adoption workshop or executive call.

  4. Owner acts

    A named account manager or specialist runs the play within an agreed time.

  5. Outcome logged

    Renewal, contraction or churn is recorded against the play that ran.

Conceptual loop linking scores to actions and outcomes; logged outcomes feed the next version of the score.

Proving impact with a holdout instead of a trend line

Churn falling after a health score launches does not prove the score worked; pricing, product releases and the economy all move at the same time. Randomly withholding the new playbook from a share of at-risk accounts, or rolling it out team by team, gives a comparison group with the same conditions.

If withholding help feels uncomfortable, compare the new playbook against the existing one rather than against nothing. The question is whether the score and its plays beat what the team was already doing.

Questions and answers

How many signals should a customer health score include?

Fewer than most teams expect. A handful of signals that clearly separate renewed and churned accounts in your own history usually beats a long list, because every extra input adds data quality risk and makes the score harder to explain. Add a signal when a review of misses shows the score is blind to something specific, not because the data happens to be available.

How often should a health score be recalculated?

Match the refresh to how fast the signals change and how quickly the team can act. Usage and support signals often justify a daily or weekly update, while billing and relationship signals change more slowly. A score that updates hourly but drives a monthly review adds noise without adding value; agree the cadence with whoever runs the playbooks.

Should NPS be part of a customer health score?

Use it carefully. Survey scores reflect the few people who answered, often not the economic buyer, and response rates vary between accounts. Test whether NPS actually separates renewed from churned accounts in your history. Many teams find that a drop in survey participation, or a detractor response from a key stakeholder, carries more signal than the average score.

How many churn events do we need before training a predictive model?

There is no fixed threshold, but you need enough churned accounts to learn patterns and still hold some back for testing. With only a few dozen events, models tend to memorize individual accounts rather than learn general rules. A backtest comparing the model with your rules-based score on held-out renewals is the honest test of whether it is ready.

More in Growth, Marketing & Sales

Back to Growth, Marketing & Sales

Next step

Send us your renewal history and the signals you already track

Tell us how many accounts and renewals you have on record and which usage, support and billing data you can export. We will suggest whether to start with a rules-based score, a model, or a data fix first.

Discuss a health score