| name | performance-calibration |
| description | Use when running calibration sessions — covers panel composition, bias patterns, distribution norms, and facilitation. |
Performance Calibration
Purpose
Calibration ensures rubrics are applied consistently across managers. Without it, two equally-strong employees can get materially different ratings (and pay) based on which manager they happen to report to. The cost is unfairness, broken trust, and pay equity drift.
Panel Composition
- 6–10 managers per panel; same level (e.g., all engineering managers, or all directors)
- Senior facilitator (often the function head or HR partner)
- Comp partner observing for band integrity
- Optional: representative from another function for cross-perspective
Panels of 12+ become hard to facilitate. Panels of 3 become group-think.
Pre-Calibration
- Each manager submits proposed ratings with evidence
- Distribution preview shared with the panel before the session
- Anomalies flagged: outlier managers (much harsher / softer than peers); demographic patterns
- Pre-read: review rubric anchors; recent calibration norms
Session Structure (3–5 hours typical)
- Re-anchor: review rubric anchors; remind of bias patterns (15 min)
- Distribution review: where is the org tracking vs. last cycle? (10 min)
- Borderline cases: discuss employees at level transitions (the most consequential cases)
- Outlier ratings: discuss any rating that's notably higher or lower than peers
- Bias check-ins: explicit pause to ask "are we under-rating Group X?"
- Confirm: each manager updates their ratings if the discussion shifted them
- Document: rationale for any rating changes
Bias Patterns to Watch
| Pattern | Description | Counter |
|---|
| Leniency | Manager rates whole team high | Look at distribution; ask for differentiation |
| Harshness | Manager rates whole team low | Compare evidence to anchored rubric |
| Halo | One strong trait → high ratings everywhere | Force per-dimension scoring |
| Recency | Last project crowds out year | Ask about a specific quarter from earlier |
| Similarity-to-self | Rates people like themselves higher | Mixed-perspective panels |
| Maternal wall | Mothers rated lower for ambiguous reasons | Watch returning-from-leave patterns |
| Stereotype activation | Gender / race expectations shape interpretation | Demographic distribution review |
| Confidence vs. competence | Confident self-presenters over-rated | Anchor on outcomes, not affect |
Demographic Distribution Review
Before final ratings lock:
- Distribution by gender, race/ethnicity, age, tenure, level
- Cell-size minimums (n ≥ 5–8)
- Statistical significance of differences
- Investigate causes; don't assume bias or dismiss
If a manager consistently rates one demographic group lower with weak evidence: coach the manager; don't just adjust the rating (the underlying pattern persists).
Forced Distribution
Most companies should not force. Cost: corrodes collaboration; demoralizes; arbitrarily punishes; can violate fairness norms.
If using guidance distributions (e.g., "expect roughly 70% Meets, 20% Exceeds, 10% Below"):
- Treat as guidance, not requirement
- Allow exceptions with rationale
- Watch for managers who hit the distribution by manipulating rather than rating honestly
Calibration Theater (Avoid)
Signs:
- Discussion happens; no ratings change
- Managers defend their ratings; group acquiesces
- "We can't change Maya's rating; her manager promised her"
- Distribution review noted but not acted on
- Demographic gaps observed but not addressed
If calibration produces no ratings changes, it's not working.
Post-Calibration
- Comp partner re-runs band integrity and pay equity check
- Final ratings communicated to employees by their managers
- Documentation: rationale for changes; manager scorecards (who's calibrated well; who needs coaching)
Common Failures
- Single-manager ratings without calibration
- Calibration limited to "the bottom and the top" (the middle has bias too)
- Ratings discussed without rubric in front of the panel
- No demographic distribution review
- "Make manager Y rate harder" without coaching the underlying issue
- Calibration after manager has already told the employee the rating
Cross-References
performance-management-systems skill
career-leveling skill — anchors ratings
pay-equity-analysis skill — adjacent audit
dei-strategist agent — distribution review
Key References
- Bock, L. (2015). Work Rules!
- Industry practitioner work on calibration (Lattice, Culture Amp guides)