पाठशाला Pathshala · दल Dal, The team · Lesson 19 · Build
Performance reviews that do not waste a month
A review earns its hours when it decides pay, level and growth with evidence people trust. How to run a light twice-yearly cycle tied to goals and bands in two weeks.
Pathshala, The Founder Library · 11 October 2026 · 7 min read

A performance review has three jobs: decide pay, decide level and tell a person what to work on next. Most startups either skip it until the first resignation forces the question, or copy a corporate process that eats a month of every manager’s time and produces ratings nobody believes. There is a version in between that takes two weeks and holds up.
This lesson sets out the cadence, the form, the calibration meeting and the link to pay, with a figure that prices the cycle in hours before anyone starts writing.
What the heavy version costs
Deloitte counted the time its old process took and arrived at close to 2 million hours a year, spent on forms, rating meetings and consensus discussions. Its leaders also found that ratings told them more about the person giving the rating than the person rated. Adobe dropped annual reviews and ratings in 2012 for quarterly Check-in conversations about goals, feedback and growth. Neither firm stopped evaluating people. Both stopped spending most of the effort on paperwork that did not change a decision.
A forty-person startup cannot lose two million hours. It can lose a month. Every manager writes eight reviews of three pages each, every employee writes a self-review that reads like a résumé, four peers per person fill a twenty-question form, and the founders spend three evenings arguing over ratings. The work that was meant to happen in that month does not. The cost is real even though nobody invoices it, which is why the figure below puts it in hours.
Twice a year, two weeks each
Run two cycles a year at fixed dates, so nobody is surprised and nobody has to ask when. The first is tied to money: it closes in the month before the annual compensation review, so ratings feed the band move and the raises. The second, six months later, is about growth and level: the same form, a rating, and a promotion decision where one is due, but no pay change outside a promotion. GitLab, which publishes its process, runs an annual talent assessment with a recommended mid-year check-in and expects the whole matrix process to take four to six weeks at its scale. A company under a hundred people can do each cycle in two.
The two-week window has a shape. Days one to four: self-reviews and peer input. Days five to eight: manager write-ups. Day nine or ten: calibration. Days eleven to fourteen: every conversation held. Anyone who joined in the last three months is marked too new to rate, which is GitLab’s rule as well, and gets a conversation without a score.
The form: four questions and a rating
Every review, self and manager, answers the same four questions in short prose. What did you commit to, and what happened? The commitments are the person’s goals for the half, which is why a company with [quarterly goals](/library/okrs-for-team-of-ten) finds reviews easy and one without them finds reviews political. How did you work? Two or three of the company’s stated behaviours, each with one example. Are you working at your level? Against the written level descriptions, below, at or ready for the next. What is the one thing to grow? One, not five.
Then a rating on three points: below expectations, meeting expectations, above expectations. Five-point scales invite a fight over the difference between a three and a four that no evidence can settle. Peer input is two named peers chosen by the manager, three questions each (what should this person keep doing, what should they change, anything the manager should know), twenty minutes of writing. Upward input on the manager is collected in the same week with the same three questions. A self-review takes an hour. A manager write-up takes about ninety minutes per report once the goals exist. Anything longer is usually a manager reconstructing six months from memory, which is a failure of the [one-on-ones](/library/one-on-ones-and-feedback-that-lands), not of the form.
At the defaults, forty people with a three-hour self-review, four hours of manager writing per report, four peer reviews each and six hours of calibration per manager, the cycle consumes about seventy working days, and each manager gives up most of a working week. The light cycle at the same headcount needs about thirty. Move the peer count first: it is the line that grows fastest and adds least, because the fourth peer rarely says what the first two did not. Then move the manager hours, and notice that a manager with ten reports and a four-hour write-up has lost seven working days before calibration ends.
Calibration without a bell curve
Calibration is the meeting where managers check that a rating means the same thing on every team. Without it, the generous manager’s team is above expectations and the strict one’s is not, and pay follows the manager rather than the work. With a forced distribution, a manager with a strong team must mark someone down to fit the curve, and everyone learns that the rating is a quota.

Hold one meeting per group of managers who report to the same leader, ninety minutes, notes prepared in advance. Discuss only three kinds of case: every rating at either end, every promotion proposal, and anyone whose rating has moved two steps since the last cycle. Each case is argued from evidence, the goals and what happened, not from adjectives. GitLab’s guidance asks managers to use the situation, behaviour and impact model and says plainly that calibration is not stack ranking. It also publishes an expected spread of about 10 per cent developing, 60 to 65 per cent performing and 25 per cent exceeding. Use a spread like that as a smell test after the meeting, not a target before it: if two-thirds of a team are above expectations, ask whether the expectations were written down.
A review that takes a month is a failure of the six months before it. When goals and one-on-ones exist, the review is a summary, not an excavation.
Tying the rating to bands and money
People trust reviews when they can see how the rating turns into pay. That needs [compensation bands](/library/compensation-bands-for-startup) and a merit rule written before the cycle opens. The rule is a small grid: rating on one axis, position in the band on the other, a raise percentage in each cell. Someone above expectations and below the midpoint of their band gets the largest raise, because they are underpaid for their level. Someone meeting expectations and above the midpoint gets the band move and little more, because they are already paid for what they do. The total of the grid must fit the raise budget the founders set for the year, so run it on the spreadsheet before the conversations, not after.
Promotion is a separate decision from the raise. A promotion moves a person to the next level’s band because they are already working at that level, which the review evidence must show. A promotion used as a retention tool for someone not yet at the level teaches the whole company that levels are negotiable.
Tell the rating and the money in the same conversation in the pay cycle. Splitting them sounds kinder and is not: the person sits through the development discussion waiting for the number. In the growth cycle there is no number, and the whole hour can be about the next six months.
The conversation and the person who is struggling
The review conversation is an hour, held in person or on video, never by email. The manager speaks first on the rating and the reason in two sentences, then listens. Nothing in the written review should surprise the person. If it does, the feedback that should have come in a one-on-one weeks earlier did not, and the manager owes an apology before the explanation.
A below-expectations rating needs a written plan within a week: the two or three specific gaps, what meeting expectations would look like in each, and a date six to eight weeks out to review progress. Some people close the gap. For those who do not, the plan is the documented record a fair exit needs, which the [firing lesson](/library/firing-fast-and-fairly-in-india) covers. Netflix’s keeper test is a useful private question for the manager before any of this: knowing everything you know today, would you hire this person again? If the honest answer is no and the review says meets expectations, one of the two is wrong.
The review calendar
Fix the two windows in the company calendar a year ahead. Four weeks before each: confirm every person has written goals for the half, refresh the level descriptions, and in the pay cycle set the raise budget and the merit grid. In the window: self-reviews and peer input by day four, manager write-ups by day eight, calibration by day ten, every conversation held by day fourteen. In the week after: below-expectations plans written, promotions and raises confirmed in writing, and a short survey asking every person whether the review was fair and useful. Once a year: count the hours the last cycle took against the light cycle, check the rating spread by team and by manager, and cut any question on the form that changed no decision.
Hour estimates in the figure are illustrative; measure your own first cycle and use those numbers for the next.
Sources
- Marcus Buckingham and Ashley Goodall, Reinventing Performance Management, Harvard Business Review, April 2015 (Deloitte: close to 2 million hours a year)
- GitLab Handbook, Talent Assessment: annual cycle, mid-year check-in, three-point performance and growth ratings, too new to rate at three months, calibration not stack ranking (checked 10 October 2026)
- Adobe, Check-in: annual reviews and ratings replaced in 2012 with quarterly conversations
- Netflix, Culture: the keeper test