पाठशाला Pathshala · ग्राहक Grāhak, The customer · Lesson 18 · Build
Consumer personas built from data, not stock photos
A persona with a stock photo and a hobby changes nothing. Three personas built from what customers bought and how often, checked every quarter, change what a consumer brand makes and spends.
Pathshala, The Founder Library · 11 October 2026 · 6 min read

Most consumer personas are a stock photograph, a first name, an age, a city and a list of hobbies, written in a workshop and pinned to a wall. Nobody makes a different decision because of them, which is the only test a persona has to pass.
This lesson builds personas the other way round: from what customers actually bought, how often and on what terms. It covers the variables to use, how to find three groups, a worked example from a D2C brand, the one-page persona, and the quarterly check that retires a persona when the data stops matching it.
What a stock-photo persona gets wrong
The workshop persona fails because it is built from attributes that do not predict behaviour. Clayton Christensen, Scott Cook and Taddy Hall made the argument in a 2006 piece for Harvard Business School Working Knowledge: segmenting by product and customer profile and averaging what each demographic says produces a one-size-fits-none product. Their milkshake study found that 40 per cent of a chain’s milkshakes were bought in the early morning, mostly by lone commuters who wanted something to make a boring drive bearable and keep hunger away until mid-morning. No demographic profile would have found them. Their purchase pattern did.
Nielsen Norman Group’s Page Laubheimer names the workshop version precisely. Proto-personas are created with no new research, from the team’s existing knowledge and best guesses. They are useful for making assumptions visible and dangerous once they are treated as findings. The same article describes two kinds built from evidence: qualitative personas from interviews with five to thirty users, and statistical personas from a survey of at least a hundred, ideally five hundred or more, with the answers clustered. A consumer brand with an order history has something better than a survey: a record of what every customer actually did.
Start from what customers did, not who they are
Choose five or six variables that describe behaviour and that you can compute for every customer from your own data. For a D2C brand a good starting set is: orders in the first 180 days; average order value; whether the first order used a discount code; payment method, prepaid or cash on delivery; share of orders in October and November, the festive months; and category mix, the share of spend in your main product line. Add the pincode tier if delivery cost or COD behaviour differs sharply between metros and smaller towns.

Leave age, gender and income out of the clustering, even if you have them. Use them afterwards to describe the groups the behaviour found. Sequoia’s data science team makes the same point about retention: to find your most valuable users, segment them by how, and how frequently, they engage. Attributes describe a person; behaviour describes a customer.
Finding three groups in the data
Two methods work at startup scale. Rules in a spreadsheet: sort customers by the two variables that matter most, often reorder frequency and first-order discount, and draw the boundaries by eye. It is crude and transparent, and every colleague can check it. K-means clustering: standardise the variables so none dominates by its units, then let the algorithm find groups. The scikit-learn documentation states the trade-offs plainly: k-means requires the number of clusters to be specified, and its criterion assumes clusters are convex and isotropic, which is not always the case. Try three, four and five clusters and keep the number whose groups you can describe in a sentence each.
Three is the house rule of thumb for a first set. One persona is not a segmentation. Five or more are rarely remembered, and personas nobody remembers make no decisions. Whatever the method, check two numbers for each group before naming it: its share of customers and its share of revenue. A group that is a third of customers and a twentieth of revenue is a persona; it is also a warning.
Then talk to people. Pick five or six customers at random from each cluster and run the interviews the [customer interview lesson](/library/the-customer-interview-done-properly) describes, asking about the last order and the one before. The data tells you what each group does. Only the conversations tell you why, and the why is what lets a team design for them.
A worked example: a D2C tea brand
A brand selling packaged tea online has 18,000 customers who ordered in the last twelve months. Clustering on the six variables above gives three groups and a remainder. The monthly restocker: 31 per cent of customers and 58 per cent of revenue, prepaid by UPI, reorders every five to six weeks at about ₹650. The festive gifter: 19 per cent of customers and 24 per cent of revenue, 70 per cent of orders in October and November, gift boxes at about ₹2,200, usually shipped to another address. The coupon sampler: 36 per cent of customers and 9 per cent of revenue, first order on a discount code, mostly cash on delivery, one in eight ever reorders. The remaining 14 per cent of customers fit none of the three and bring 9 per cent of revenue.
The interviews add the reasons. Restockers drink the tea every day and dread running out; they would happily subscribe if they could skip a month. Gifters choose by the box and the message card, not the blend. Samplers came for the discount on a marketplace-style ad and compared the price with the supermarket. Three decisions follow at once: a subscription with a skip button for restockers, a gift catalogue live by the first week of September, and a hard look at the [unit economics](/library/d2c-unit-economics-order-that-must-make-money) of the first-order discount that is buying samplers.
Writing the persona page
One page per persona, and no stock photograph. A name that describes the behaviour, not a first name: “monthly restocker” tells a designer something, “Priya, 32” does not. The numbers: share of customers, share of revenue, order frequency, order value, payment method, retention at six months. What they are trying to get done, in one sentence from the interviews. Three verbatim quotes. What we build for them, and what we stop doing because of them. The definition: the exact rule or cluster that assigns a customer to the persona, so anyone can rerun it.
The definition is the part workshop personas never have, and it is what makes the rest honest. A persona with a definition can be counted every quarter. A persona without one can only be believed.
A persona is a definition you can rerun on next quarter’s customers. If it cannot be counted, it cannot be wrong, and a persona that cannot be wrong cannot help.
When to retire a persona
Customers change, channels change and the brand changes who it attracts. A persona built a year ago may now describe a minority of the customers arriving. Rerun the definitions on last quarter’s new customers and compare the shares with the shares when the personas were built.

At the default settings the three personas still describe 78 per cent of new customers and none has moved by more than half, so they stay. Now suppose the brand stops paying for discount ads: drag the coupon sampler down to 15 per cent and the coverage falls to 63 per cent, below the rule’s 70. The personas no longer describe the business, and the right response is to rebuild all three from fresh data, not to add a fourth. Put the sampler back, raise the restocker to 40 per cent and drag the festive gifter below 10 per cent, and the tool says to retire or merge the gifter; check first that you are not judging a festive persona on a quiet quarter.
The quarterly persona check
Once a quarter, an hour with marketing, product and the founder. Rerun each definition on the quarter’s new customers and on the whole active base. Compare shares of customers and of revenue with the last check and with the build. Read the rules: rebuild below 70 per cent coverage, retire below 10 per cent, rewrite after a move of more than half. Interview two customers from each persona and one from the unassigned group, which is where the next persona usually appears. Update each page: numbers, quotes and the decisions it is driving. Record the date of the check on every page, so nobody reads a persona that has not been tested this year.
The brand and its figures in the example are illustrative. The method works on any order history long enough to include a full year, festive season and all.
Sources
- Clayton M. Christensen, Scott Cook and Taddy Hall, What Customers Want from Your Products, HBS Working Knowledge, 16 January 2006 (the milkshake study)
- Page Laubheimer, 3 Persona Types: Lightweight, Qualitative, and Statistical, Nielsen Norman Group, 21 June 2020 — Proto-personas with no new research; qualitative from 5–30 interviews; statistical from a survey of at least 100, ideally 500 or more, clustered.
- Sequoia Capital Data Science Team, Retention
- scikit-learn, User Guide: Clustering (K-means) — Number of clusters must be specified; inertia assumes convex, isotropic clusters.