GuideGrowth, Marketing & Sales
How to build a customer health score that predicts churn and prompts action
A health score is only useful if it predicts a defined outcome and changes what someone does on Monday morning. This guide starts with the outcome, moves through signal selection and a rules-based score validated against past renewals, then covers when a predictive churn model is worth building, how to explain scores to account managers and how to prove the program works.
On this page
- Choose the outcome the score predicts
- Candidate signals worth testing for churn risk
- Building and validating a rules-based score
- Rules-based score or predictive churn model
- Modeling traps that make churn models look better than they are
- Explaining a score so account managers will use it
- From risk tier to playbook to measured outcome
- Proving impact with a holdout instead of a trend line
- Questions and answers
Choose the outcome the score predicts
Different outcomes reward different signals. Pick one primary outcome and a horizon, such as the next renewal date, and write both down.
- Logo churn
- A customer ends the relationship entirely. Simple to label, but treats a small account the same as a large one.
- Gross revenue churn
- Recurring revenue lost to cancellation and downgrades, which weights accounts by what they pay.
- Contraction
- Fewer seats, a lower tier or removed products while the customer stays. Often an early warning of logo churn.
- Non-renewal at term
- The outcome observed at each contract end date. Natural for annual contracts, but labels arrive only once per term.
Candidate signals worth testing for churn risk
Treat each as a hypothesis to test against history, not as a known predictor.
Building and validating a rules-based score
Freeze a snapshot before each past renewal
For every renewal in your history, rebuild what the signals looked like at a fixed point before the decision, such as a full quarter earlier. Using today's values for past accounts leaks the outcome into the score.
Pick a handful of signals and set weights
Choose the signals that differ most between renewed and churned accounts in the snapshot, agree weights with experienced account managers, and keep the list short enough to explain in a sentence.
Set thresholds for risk tiers
Translate scores into red, amber and green tiers sized to what the team can act on. A red tier containing half the book is a list, not a priority.
Backtest against actual outcomes
Check what share of churned accounts the red tier would have caught and what share of red accounts actually churned. Both numbers matter: catching every churner by flagging everyone helps nobody.
Review the misses with the people who knew the accounts
Walk through churned accounts the score rated green and renewed accounts it rated red. Misses usually reveal a missing signal, such as a merger or budget freeze, or a data quality problem.
Version and document the score
Record the definition, data sources, weights and backtest results, and change them only through a reviewed release so trends over time stay comparable.
Rules-based score or predictive churn model
| Consideration | Rules-based score | Predictive model |
|---|---|---|
| Data needed | Enough history to sanity-check weights | Enough churn events to learn patterns and still hold out a test set |
| Transparency | Anyone can read the rule | Needs per-account explanations to earn trust |
| Handles interactions between signals | Poorly; weights add up independently | Well, if the data supports it |
| Maintenance | Occasional reweighting | Monitoring for drift and regular retraining |
| Best fit | Smaller books, few churn events, a new program | Large books with plenty of labeled history and data engineering support |
Most teams should start in the highlighted column and graduate when the backtest shows clear limits.
Modeling traps that make churn models look better than they are
Label leakage
Early signalA feature such as a cancellation ticket or a renewal opportunity marked lost dominates the model.
MitigationBuild features only from data available at the snapshot date and exclude anything recorded as part of the churn process.
Accuracy that hides failure
Early signalHigh accuracy because most customers renew, while the model finds few churners.
MitigationEvaluate precision among the top-ranked accounts the team can actually contact, and check calibration so a predicted risk matches observed churn.
Survivorship in the training set
Early signalTraining only on current customers, so the accounts that churned early are missing.
MitigationInclude every account that existed at each snapshot, churned or not.
Drift after product or pricing changes
Early signalThe score's backtest performance falls on recent cohorts.
MitigationMonitor performance by cohort and retrain or reweight after major changes.
Explaining a score so account managers will use it
A number without a reason gets ignored. Show the two or three factors that moved each account's score, in plain language, next to the account record: usage of the core workflow down sharply this month, sponsor left, two unresolved high-severity tickets. For models, per-account attribution methods such as SHAP values can produce those factors, but they still need translating into business terms.
Let account managers flag a score they disagree with and give a reason. Those overrides are labeled data about what the score misses, and reviewing them regularly keeps the people closest to customers invested in improving it.
From risk tier to playbook to measured outcome
- Signals refresh
Usage, support, billing and relationship data update on a set schedule.
- Score and tier
Each account gets a tier and its top contributing factors.
- Playbook triggered
A tier and factor combination maps to a specific play, such as an adoption workshop or executive call.
- Owner acts
A named account manager or specialist runs the play within an agreed time.
- Outcome logged
Renewal, contraction or churn is recorded against the play that ran.
Proving impact with a holdout instead of a trend line
Churn falling after a health score launches does not prove the score worked; pricing, product releases and the economy all move at the same time. Randomly withholding the new playbook from a share of at-risk accounts, or rolling it out team by team, gives a comparison group with the same conditions.
If withholding help feels uncomfortable, compare the new playbook against the existing one rather than against nothing. The question is whether the score and its plays beat what the team was already doing.
Questions and answers
How many signals should a customer health score include?
Fewer than most teams expect. A handful of signals that clearly separate renewed and churned accounts in your own history usually beats a long list, because every extra input adds data quality risk and makes the score harder to explain. Add a signal when a review of misses shows the score is blind to something specific, not because the data happens to be available.
How often should a health score be recalculated?
Match the refresh to how fast the signals change and how quickly the team can act. Usage and support signals often justify a daily or weekly update, while billing and relationship signals change more slowly. A score that updates hourly but drives a monthly review adds noise without adding value; agree the cadence with whoever runs the playbooks.
Should NPS be part of a customer health score?
Use it carefully. Survey scores reflect the few people who answered, often not the economic buyer, and response rates vary between accounts. Test whether NPS actually separates renewed from churned accounts in your history. Many teams find that a drop in survey participation, or a detractor response from a key stakeholder, carries more signal than the average score.
How many churn events do we need before training a predictive model?
There is no fixed threshold, but you need enough churned accounts to learn patterns and still hold some back for testing. With only a few dozen events, models tend to memorize individual accounts rather than learn general rules. A backtest comparing the model with your rules-based score on held-out renewals is the honest test of whether it is ready.