- name
- performance-calibration
- description
- Use when running calibration sessions — covers panel composition, bias patterns, distribution norms, and facilitation.
# Performance Calibration
## Purpose
Calibration ensures rubrics are applied consistently across managers. Without it, two equally-strong employees can get materially different ratings (and pay) based on which manager they happen to report to. The cost is unfairness, broken trust, and pay equity drift.
## Panel Composition
- 6–10 managers per panel; same level (e.g., all engineering managers, or all directors)
- Senior facilitator (often the function head or HR partner)
- Comp partner observing for band integrity
- Optional: representative from another function for cross-perspective
Panels of 12+ become hard to facilitate. Panels of 3 become group-think.
## Pre-Calibration
- Each manager submits proposed ratings with evidence
- Distribution preview shared with the panel before the session
- Anomalies flagged: outlier managers (much harsher / softer than peers); demographic patterns
- Pre-read: review rubric anchors; recent calibration norms
## Session Structure (3–5 hours typical)
1. **Re-anchor**: review rubric anchors; remind of bias patterns (15 min)
2. **Distribution review**: where is the org tracking vs. last cycle? (10 min)
3. **Borderline cases**: discuss employees at level transitions (the most consequential cases)
4. **Outlier ratings**: discuss any rating that's notably higher or lower than peers
5. **Bias check-ins**: explicit pause to ask "are we under-rating Group X?"
6. **Confirm**: each manager updates their ratings if the discussion shifted them
7. **Document**: rationale for any rating changes
## Bias Patterns to Watch
| Pattern | Description | Counter |
|---------|-------------|---------|
| **Leniency** | Manager rates whole team high | Look at distribution; ask for differentiation |
| **Harshness** | Manager rates whole team low | Compare evidence to anchored rubric |
| **Halo** | One strong trait → high ratings everywhere | Force per-dimension scoring |
| **Recency** | Last project crowds out year | Ask about a specific quarter from earlier |
| **Similarity-to-self** | Rates people like themselves higher | Mixed-perspective panels |
| **Maternal wall** | Mothers rated lower for ambiguous reasons | Watch returning-from-leave patterns |
| **Stereotype activation** | Gender / race expectations shape interpretation | Demographic distribution review |
| **Confidence vs. competence** | Confident self-presenters over-rated | Anchor on outcomes, not affect |
## Demographic Distribution Review
Before final ratings lock:
- Distribution by gender, race/ethnicity, age, tenure, level
- Cell-size minimums (n ≥ 5–8)
- Statistical significance of differences
- Investigate causes; don't assume bias *or* dismiss
If a manager consistently rates one demographic group lower with weak evidence: coach the manager; don't just adjust the rating (the underlying pattern persists).
## Forced Distribution
Most companies should not force. Cost: corrodes collaboration; demoralizes; arbitrarily punishes; can violate fairness norms.
If using guidance distributions (e.g., "expect roughly 70% Meets, 20% Exceeds, 10% Below"):
- Treat as guidance, not requirement
- Allow exceptions with rationale
- Watch for managers who hit the distribution by manipulating rather than rating honestly
## Calibration Theater (Avoid)
Signs:
- Discussion happens; no ratings change
- Managers defend their ratings; group acquiesces
- "We can't change Maya's rating; her manager promised her"
- Distribution review noted but not acted on
- Demographic gaps observed but not addressed
If calibration produces no ratings changes, it's not working.
## Post-Calibration
- Comp partner re-runs band integrity and pay equity check
- Final ratings communicated to employees by their managers
- Documentation: rationale for changes; manager scorecards (who's calibrated well; who needs coaching)
## Common Failures
- Single-manager ratings without calibration
- Calibration limited to "the bottom and the top" (the middle has bias too)
- Ratings discussed without rubric in front of the panel
- No demographic distribution review
- "Make manager Y rate harder" without coaching the underlying issue
- Calibration after manager has already told the employee the rating
## Cross-References
- `performance-management-systems` skill
- `career-leveling` skill — anchors ratings
- `pay-equity-analysis` skill — adjacent audit
- `dei-strategist` agent — distribution review
## Key References
- Bock, L. (2015). *Work Rules!*
- Industry practitioner work on calibration (Lattice, Culture Amp guides)
Auf GitHub ansehen