पाठशाला Pathshala · ग्राहक Grāhak, The customer · Lesson 24 · Scale

Customer health scores that actually predict churn

Most health scores are a traffic light designed in a meeting and never checked. Build one from usage, support and payment signals, then test it against the customers who actually left before anyone acts on it.

Pathshala, The Founder Library · 11 October 2026 · 6 min read

A blue blood pressure cuff and gauge lying on a white surface.
Photograph: cottonbro studio · Pexels

Most customer health scores are designed in a meeting: logins count for forty points, support tickets for thirty, an NPS answer for the rest, green above seventy. Nobody checks whether the accounts that turned red were the ones that left.

This lesson sets out how to build a score that predicts: which signals to use, how to build a first version from your own history in a spreadsheet, how to test it against the customers who actually churned, and how to turn it into a weekly list someone acts on. A figure lets you backtest a score on sixty past accounts.

Why most health scores do not predict

Three faults recur. The signals are the easy ones, not the telling ones: logins are counted because the product logs them, though a login says little about whether the job got done. The weights are opinions: whoever designed the score decided support tickets were bad, though the customers who file the most tickets are often the most engaged, and the silent ones are the risk. And the score is never validated. A number that has not been checked against who left is a guess with a colour.

The cost of getting it wrong compounds. David Skok’s analysis of SaaS churn models a business with no churn ending nearly 60 per cent larger than the same business losing 2.5 per cent of revenue a month. A score that sends the customer team to the wrong accounts every week spends that difference on the customers who were staying anyway.

Choose signals from three families

Usage, measured against the account’s own baseline rather than an absolute number. The decline matters more than the level: a clinic that used to send three hundred reminders a week and now sends ninety is at risk even if ninety is above your average. Sequoia’s guide to measuring product health notes that a drop in sessions is the earliest leading indicator of a drop in daily active users, and that engagement is the most important driver of retention. Choose the action that represents the job done, not the login.

Support strain: issues unresolved beyond your target time, the same issue raised twice, an escalation to a founder, and a fall in tickets from an account that used to raise them. Read the words as well as the count; a ticket that says “again” is worth more than three that ask how to export.

Payment trouble: a late invoice, a failed charge, a mandate that lapsed. In India this family matters more than founders expect. Under the Reserve Bank’s e-mandate rules, as Stripe’s India documentation sets them out, each recurring debit needs a pre-debit notice at least 24 hours ahead with an option to cancel, card charges above ₹15,000 need fresh authentication each time, and an off-session payment without a mandate is declined. Every step is a place where an account can slip, and a customer who has let a payment lapse once has often started to leave. Checked October 2026.

A fourth family is worth adding by hand where you can: relationship, chiefly whether your champion has left the customer’s company. It rarely shows in the data until too late, so ask account owners to record it.

Build the first version from your own history

Take every account that came up for renewal or cancelled in the last twelve months. For each, take a snapshot of the signals as they stood ninety days before the renewal or cancellation date, which is when you could still have acted. Label each account left or stayed. A spreadsheet is enough for the first version; a few hundred rows will do.

Express each signal on a scale from zero to one: usage decline as the share lost against the previous quarter, support strain from zero for none to one for an escalation, payment trouble as one if any payment was late, failed or lapsed. Start with equal weights. Score each account from zero to a hundred, where higher means more at risk. Then test it before anyone sees a dashboard.

Two cautions keep the test honest. Use only what you knew on the snapshot date: a cancellation ticket filed the week before leaving will make any score look prophetic and will be useless in practice. And respect the size of your history. With fewer than thirty leavers, keep the score to the three families and equal or near-equal weights; a score with twelve finely tuned weights fitted to twenty-five departures describes the past and predicts nothing. Where segments behave differently, a clinic and a hospital chain for example, compute usage decline against each segment’s own baseline before combining.

Validate it against who actually left

Two numbers say whether the score works. Recall is the share of customers who actually left that the score flagged. Precision is the share of flagged accounts that actually left. The scikit-learn documentation puts it plainly: precision is the ability not to label as positive a sample that is negative, and recall is the ability to find all the positive samples. A low threshold catches more leavers and raises more false alarms; a high one does the reverse.

Empty chairs in a dim waiting area photographed in black and white.
Test the score on the seats that are already empty. If it did not flag the customers who left it will not flag the next ones. Photograph: iam vumilia · Pexels

Set the threshold by capacity, not by taste. If the customer team can run fifteen serious conversations a week, the threshold should flag about fifteen accounts a week. Then adjust the weights until recall and precision are both as high as the data allows at that volume. Move the sliders below to see how much the weights matter.

Compare the score with the simplest rule you could have used instead, such as flagging every account whose usage fell by a third. If the score does not beat that rule on both numbers, use the rule; it is easier to explain to the team and harder to argue with. A score earns its complexity only by catching leavers the rule misses.

With usage and support weighted equally and payment ignored, the score flags fourteen accounts and catches two of the eleven that left: a recall of 18 per cent and a precision of 14 per cent. Weight usage at five, support at one and payment at five, and it flags eleven, catching seven of the eleven leavers. Lower the threshold to forty and it catches eight, at the cost of eight false alarms. The data in the figure is illustrative, but the pattern is common: the signal the team thought mattered was weak, and the one it ignored was strong.

A health score that has not been tested against the customers who left is an opinion with a colour.

Turn the score into a weekly action

A score earns its keep only when someone acts on it. Each Monday, list the flagged accounts with the signal that flagged them and an owner for each. Write a short playbook per family. Usage decline: a call to find what changed, often a new staff member nobody trained. Support strain: a founder or senior call that closes the open issue first and asks second. Payment trouble: a call within forty-eight hours and a payment link to re-register the mandate, as the [churn interview lesson](/library/churn-interview-learning-from-those-who-left) describes for involuntary churn.

Record the outcome of every intervention: saved, lost, or no risk after all. Those outcomes are the next quarter’s training data. Do not show the score to customers or tie it to account managers’ pay until it has held up for two quarters; a number that people are paid on stops being a measurement.

Recalibrate every quarter

Each quarter: re-run the backtest on the accounts that renewed or left in the quarter just ended, with their snapshot ninety days earlier. Record recall and precision at the current threshold, next to last quarter’s. Check each signal on its own: if one no longer separates leavers from stayers, cut its weight. Add one candidate signal from the churn interviews, test it and keep it only if it improves both numbers. Reset the threshold to the team’s capacity. Write the weights, the numbers and the date at the top of the sheet so the score has a history.

A score that is recalibrated every quarter gets better as the company grows. A score set once and left alone gets worse, because the customers, the product and the reasons for leaving all change.


The data in the figure is illustrative. Payment rules quoted here were checked in October 2026; confirm the current limits with your payment provider.

Sources

  1. David Skok, Unlocking the Path to Negative Churn, For Entrepreneurs — A 0% churn line ends nearly 60% above a 2.5% monthly churn line in his model.
  2. Sequoia Capital, Measuring Product Health — A drop in sessions is the earliest leading indicator of a drop in DAU; engagement is the most important driver of retention.
  3. Stripe Docs, India recurring payments (RBI e-mandate requirements) — Pre-debit notice at least 24 hours ahead; AFA each time above ₹15,000 for cards; off-session payments without a mandate are declined. Checked October 2026.
  4. scikit-learn, Metrics and scoring: precision, recall and F-measures