Team

Performance review frameworks

Performance review frameworks are the systems a company uses to judge how people work and tell them: annual reviews, check-ins, OKRs, 360 feedback, calibration and forced ranking.

In short

Performance review frameworks are the systems a company uses to judge individual work and feed that judgment into pay, promotion and development. The six common ones are annual reviews, continuous check-ins, OKRs, 360-degree feedback, calibration and forced ranking. They differ in who judges, how often, and what decision the result drives.

Origin
HR practice; no single originator (these six grew up separately), various
Level
301 · Advanced
Fits
Scale-up, Enterprise
Time to apply
One planning cycle to design; a quarter to run the first round
What you need
a list of the decisions the review must feed: pay, promotion, development, exits · the managers who will rate people, and the people who will be rated · one owner who can change the process after the first round

Performance review frameworks are the systems a company uses to judge how people work and tell them. Each answers three questions: who judges, how often, and which decision the result drives. Six are in common use: annual reviews, continuous check-ins, OKRs, 360-degree feedback, calibration and forced ranking. They can be combined, and most firms use two or three.

Framework Who judges Rhythm Main output Main weakness
Annual review One manager Once a year A rating and a write-up Looks backward, rater bias
Continuous check-ins Manager and employee Weekly to quarterly Priorities and feedback Needs trained managers
OKR scoring The team, against goals Quarterly A 0 to 1 score per key result Not a judgment of the person
360-degree feedback Peers, reports, manager Usually per cycle A feedback report Small average improvement
Calibration A committee of managers After ratings Adjusted ratings Compresses toward the middle
Forced ranking Managers, to a quota Yearly Fixed performance bands Quotas misfit small groups

Why are ratings so noisy?

A rating tells you more about the rater than about the person rated. Scullen, Mount and Goff’s 2000 study had 4,492 managers rated by two bosses, two peers and two subordinates. As Buckingham and Goodall report in HBR, 62 percent of the variance in ratings came from the individual rater’s peculiarities, and actual performance explained 21 percent.

Two horizontal bars compare what explains the variance in performance ratings: the rater's own tendencies at 62 percent, drawn in blue, and the person's actual performance at 21 percent.
In the 4,492-manager study, rater tendencies explained about three times as much of a rating as the work did.

Other reliability studies agree that raters often disagree. Viswesvaran, Ones and Schmidt’s meta-analysis put the mean interrater reliability of supervisor ratings of overall performance at .52. Conway and Huffcutt’s 1997 meta-analysis found .50 for supervisors, .37 for peers and .30 for subordinates. Everything below inherits this problem.

Annual reviews: what are they for?

An annual review is one manager rating a year of work, usually to set pay. Its defenders say it forces a decision and a record. Its critics point at the backward view: Cappelli and Tavis, in a 2016 HBR article, argued the focus should shift from accountability to learning, and the article’s summary says more than a third of US companies had abandoned traditional appraisals. Gallup’s 2017 post reports that only 14 percent of employees strongly agree reviews inspire them to improve and 29 percent that reviews are fair.

Continuous check-ins: how do they replace the review?

Check-ins are short, regular conversations about near-term work, and they move the conversation from last year’s score to next week’s priorities. In 2015, Pulakos and colleagues argued that formal systems reduced performance management to intermittent steps disconnected from daily work, and that real-time informal feedback drives performance.

A year-long timeline with three rows: the annual review marks one dot, Adobe's quarterly check-in marks four dots, and Deloitte's weekly check-ins, drawn in blue, mark a tick every week.
The three designs differ most in rhythm: one conversation a year, four, or one a week.

Two public cases show the design. Deloitte, according to its own authors, found its process used close to 2 million hours a year for more than 65,000 people. It replaced it with four forward-looking questions that team leaders answer at project end, such as whether they would always want the person on their team, and with weekly check-ins. Adobe dropped annual reviews and ratings in 2012, and says its conversations now happen quarterly. Rock and Jones note that Adobe and similar firms still differentiate pay. Both accounts are company-reported.

Frequency alone does not guarantee value. Kluger and DeNisi’s 1996 meta-analysis of feedback interventions found an average gain, yet over a third reduced performance.

OKR scoring: is it a performance review?

OKRs tie goals to a quarterly rhythm but are not a verdict on a person. Google’s guide says OKRs are not synonymous with performance evaluation and grades them from 0.0 to 1.0, with 0.6 to 0.7 as the expected sweet spot. The case for them rests on goal-setting research: Locke and Latham summarise that do-your-best goals give weaker results than specific ones. See the OKR page, KPI vs OKR and management by objectives for how goals differ from reviews.

360-degree feedback: when does it change behavior?

360 feedback collects views from peers, reports and managers, and it works for development more than for judgment. Smither, London and Reilly’s meta-analysis of 24 longitudinal studies found improvement in ratings over time was generally small. Bracken and Rose list four conditions for behavior change: relevant content, credible data, accountability and high participation. Deloitte left 360 tools out of its redesign.

Calibration: does it fix bias?

Calibration is a meeting where managers compare and adjust ratings. Demeré, Sedatole and Woods studied one multinational: committees adjusted 25 percent of ratings, downward more often than upward, which improved consistency and curbed leniency but worsened centrality bias, the habit of bunching ratings in the middle. Bol, de Aguiar and Lill found that supervisors who were cheaper to defend got fewer adjustments and higher ratings, even after controlling for performance.

Forced ranking: does the curve work?

Forced ranking sorts people into fixed bands. GE’s vitality curve, as Government Executive reported, fired the bottom 10 percent, and the company dropped formal forced ranking about a decade before ending annual reviews in 2015. Scullen, Bergey and Aiman-Smith’s simulation found workforce potential could improve, mostly in the first years, depending on how many were let go. Berger, Harbring and Sliwka’s experiment found 6 to 12 percent higher productivity, but harm when workers could sabotage each other.

How do the six fit together?

Match each framework to the decision it can support. Check-ins serve coaching, because they are frequent and forward-looking. OKRs serve team goals, because Google keeps their scores apart from evaluation. A 360 report serves one person’s development. Ratings, calibration and any quota serve the pay and promotion decisions, where fairness between people matters most and where the rater effect does the most damage.

Deloitte shows the split in practice. Its quarterly or project-end snapshot feeds compensation and succession discussions, while the weekly check-in is a separate conversation about near-term work. Pay is then decided by a leader who knows the person, with the snapshot data as a starting point, as Buckingham and Goodall describe. The design separates what is measured from what is discussed.

Should you drop ratings altogether?

The experts disagree. Murphy argues that systems of regular evaluation almost always fail and organizations should evaluate only where it is useful. The SIOP debate recorded the opposing view that performance is always evaluated in some way, so the question is how. We found no source that shows one framework beating the others across firms. Schleicher and colleagues’ review of research from 1980 to 2017 treats performance management as a system, not a rating form. Choose by the decision at stake, and treat any rating as a noisy input. A Growth Lab plan starts from naming which decisions a review must support.

Related reading: competency models define what to rate.

How to apply Performance review frameworks, step by step

  1. List the decisions the review feeds. Write down every decision that uses a review result: pay rises, bonuses, promotion, development plans, dismissals. Each decision needs a different kind of evidence, and one process rarely serves all of them well. Result: a short table of decisions and who makes each.
  2. Separate the pay conversation from the coaching conversation. Give coaching its own frequent, forward-looking rhythm, and keep the pay judgment as a smaller, rarer step. Google says OKR scores are not synonymous with performance evaluation, and Deloitte split a quarterly snapshot from weekly check-ins. Result: two named conversations with two calendars.
  3. Set the check-in rhythm. Pick a cadence a manager can keep: weekly for near-term work as Deloitte designed, quarterly as Adobe runs it now. Fix a length and three standing topics: priorities, feedback, obstacles. Result: recurring slots in every manager's calendar.
  4. Fix what raters are asked. Replace broad trait scales with specific questions about observable work or about what the manager would do next, such as pay, promote or reassign. Deloitte's four snapshot statements are an example. Result: a form of 3 to 5 questions with plain wording.
  5. Add 360 input only for development. If you want views from peers and reports, collect them to help the person, not to set their pay. Choose relevant content, credible raters, someone who follows up, and high participation. Result: a feedback report owned by the person, with a follow-up date.
  6. Calibrate only where one pay pool has several raters. Where several managers rate people who compete for the same budget, hold one short session to compare evidence and spot managers who are always lenient or harsh. Record the reason for every change. Result: ratings with a written audit trail.
  7. Audit the system after the first round. Compare average ratings by manager, ask people whether the process felt fair, and count the manager hours it took. Change one thing and run it again. Result: a one-page review of the review.

Examples

A 30-person support team at a payments company

Illustrative. Two team leads rate the same kind of work. Lead A's average is 4.2 out of 5 and Lead B's is 3.1. Nothing else in the data explains a gap of 1.1 points, so the gap says more about the raters than about the people, which is what the idiosyncratic rater effect predicts. A calibration session compares ticket quality and customer notes, and the team moves to quarterly check-ins with one annual pay discussion.

A 25-person clinic group

Illustrative. The practice manager drops the yearly form and holds a 20-minute monthly one-to-one with each nurse and front-desk coordinator, with three fixed topics. Once a year a short compensation review looks at the notes from those meetings. The change costs about 8 hours a month for 25 conversations and removes a form nobody read.

Deloitte's redesign

Deloitte's own authors reported that its old process consumed an enormous number of hours. It replaced year-end ratings with four statements answered by team leaders at project end, plus weekly check-ins, and dropped cascading objectives and 360 tools, as Buckingham and Goodall wrote in Harvard Business Review. The firm's claim that the new data is reliable is its own.

When to use it

Use a deliberate review framework as soon as managers disagree about what good looks like, or when pay and promotion depend on a rating. It matters most once no single founder sees everyone's work any more.

When not to use it

Skip formal ratings in a team of five to eight where the founder works beside everyone; frequent conversation is enough. Do not use one framework for every purpose: 360 results used for pay, or a forced curve applied to a small team, produce more harm than information.

Common mistakes

  • Treating a rating as a measurement of the person. Research found that rater tendencies explain far more of the variance in ratings than the work itself, so a score reflects its giver.
  • Using 360 feedback to decide pay. A meta-analysis of 24 longitudinal studies found improvement after multisource feedback was generally small, and was most likely when recipients saw a need to change.
  • Calibrating without evidence. A study of 737 employees found supervisors who defended their ratings well got fewer adjustments and higher ratings, even after controlling for performance.
  • Applying a forced curve to a small team. Ten percent of a 30-person team is three people, and chance alone can put them in the bottom band.
  • Assuming more feedback is always better. A meta-analysis of feedback interventions found that more than a third reduced performance.

FAQ

What are the main methods of performance evaluation?

Six are common: annual reviews, continuous check-ins, OKR scoring, 360-degree feedback, calibration meetings and forced ranking. They differ in who rates (one manager or many), how often, and whether the result drives pay. Most companies combine two or three, such as check-ins for coaching and a yearly pay review.

What is calibration in a performance review?

Calibration is a meeting where managers compare and adjust each other's ratings to make them consistent. In one multinational the committees adjusted about a quarter of ratings, mostly downward, which reduced leniency but compressed ratings toward the middle, according to Demeré, Sedatole and Woods (Management Science, 2019).

Is 360-degree feedback effective?

It helps development when the content is relevant, raters are credible, someone is accountable for follow-up and participation is high, according to Bracken and Rose. A meta-analysis of 24 studies found average improvement was small. Deloitte left 360 tools out of its redesign.

Do companies still use annual performance reviews?

Many do, but fewer than before. Adobe dropped annual reviews and ratings, and GE announced it was abandoning formal annual reviews. Researchers still debate whether to drop ratings altogether, as a 2015 SIOP conference debate showed. Most replacements keep some yearly summary conversation.

What is forced ranking?

Forced ranking requires managers to place employees into fixed bands, such as GE's vitality curve, where the bottom 10 percent were let go. A 2005 simulation found workforce potential could rise, mostly in early years. A 2013 experiment found productivity gains but harm when workers could sabotage each other.

Sources

  1. Steven E. Scullen, Michael K. Mount, Maynard Goff, Understanding the latent structure of job performance ratings, Journal of Applied Psychology 85(6), 2000
  2. Marcus Buckingham, Ashley Goodall, Reinventing performance management, Harvard Business Review, April 2015
  3. Adobe, Check-in (company account of its performance approach)
  4. Peter Cappelli, Anna Tavis, The performance management revolution, Harvard Business Review, October 2016
  5. David Rock, Beth Jones, Why more and more companies are ditching performance ratings, Harvard Business Review, September 2015
  6. Elaine D. Pulakos, Rose Mueller Hanson, Sharon Arad, Neta Moye, Performance management can be fixed, Industrial and Organizational Psychology 8(1), 2015
  7. Seymour Adler and colleagues, Getting rid of performance ratings: genius or folly? A debate, Industrial and Organizational Psychology 9(2), 2016
  8. Kevin R. Murphy, Performance evaluation will not die, but it should, Human Resource Management Journal 30(1), 2020
  9. Avraham N. Kluger, Angelo DeNisi, The effects of feedback interventions on performance, Psychological Bulletin 119(2), 1996
  10. Chockalingam Viswesvaran, Deniz S. Ones, Frank L. Schmidt, Comparative analysis of the reliability of job performance ratings, Journal of Applied Psychology 81(5), 1996
  11. James M. Conway, Allen I. Huffcutt, Psychometric properties of multisource performance ratings, Human Performance 10(4), 1997
  12. James W. Smither, Manuel London, Richard R. Reilly, Does performance improve following multisource feedback?, Personnel Psychology 58(1), 2005
  13. David W. Bracken, Dale S. Rose, When does 360-degree feedback create behavior change?, Journal of Business and Psychology 26(2), 2011
  14. Steven E. Scullen, Paul K. Bergey, Lynn Aiman-Smith, Forced distribution rating systems and the improvement of workforce potential, Personnel Psychology 58(1), 2005
  15. Johannes Berger, Christine Harbring, Dirk Sliwka, Performance appraisals and the impact of forced distribution, Management Science 59(1), 2013
  16. Margaret W. Demeré, Karen L. Sedatole, Alexander Woods, The role of calibration committees in subjective performance evaluation systems, Management Science 65(4), 2019
  17. Jasmijn C. Bol, Marcelle de Aguiar, Jordan B. Lill, Calibration in the performance evaluation process, Human Resource Management 64(4), 2025
  18. Google re:Work, Set goals with OKRs
  19. Gallup, Give performance reviews that actually inspire employees, 2017
  20. Government Executive (Quartz), GE had to kill its annual performance reviews after more than three decades, 2015
  21. Edwin A. Locke, Gary P. Latham, Building a practically useful theory of goal setting and task motivation, American Psychologist 57(9), 2002
  22. Charles A. Schleicher and colleagues, Putting the system into performance management systems, Journal of Management 44(6), 2018

Last updated Oct 9, 2026

Ilia PushinFounder, PUSHERS & COO Fintech ServiceIlia builds operating systems for growing companies in fintech and healthcare. Since 2021 he has run cross-border payments at ARBI Exchange, a licensed currency exchange in Thailand, including KYC and AML and the move into new jurisdictions.About the authorLinkedIn
Related frameworks
More frameworks
Want Performance review frameworks running inside your company?Request an operations audit