Skip to main content

performance-calibration

Use when running calibration sessions — covers panel composition, bias patterns, distribution norms, and facilitation.

Jump to install

Source facts

Repository
ccashwell/agentic-hr
Last source activity
April 25, 2026 at 23:49
Detected SKILL.md language
English
Stars
0
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
performance-calibration
description
Use when running calibration sessions — covers panel composition, bias patterns, distribution norms, and facilitation.
# Performance Calibration ## Purpose Calibration ensures rubrics are applied consistently across managers. Without it, two equally-strong employees can get materially different ratings (and pay) based on which manager they happen to report to. The cost is unfairness, broken trust, and pay equity drift. ## Panel Composition - 6–10 managers per panel; same level (e.g., all engineering managers, or all directors) - Senior facilitator (often the function head or HR partner) - Comp partner observing for band integrity - Optional: representative from another function for cross-perspective Panels of 12+ become hard to facilitate. Panels of 3 become group-think. ## Pre-Calibration - Each manager submits proposed ratings with evidence - Distribution preview shared with the panel before the session - Anomalies flagged: outlier managers (much harsher / softer than peers); demographic patterns - Pre-read: review rubric anchors; recent calibration norms ## Session Structure (3–5 hours typical) 1. **Re-anchor**: review rubric anchors; remind of bias patterns (15 min) 2. **Distribution review**: where is the org tracking vs. last cycle? (10 min) 3. **Borderline cases**: discuss employees at level transitions (the most consequential cases) 4. **Outlier ratings**: discuss any rating that's notably higher or lower than peers 5. **Bias check-ins**: explicit pause to ask "are we under-rating Group X?" 6. **Confirm**: each manager updates their ratings if the discussion shifted them 7. **Document**: rationale for any rating changes ## Bias Patterns to Watch | Pattern | Description | Counter | |---------|-------------|---------| | **Leniency** | Manager rates whole team high | Look at distribution; ask for differentiation | | **Harshness** | Manager rates whole team low | Compare evidence to anchored rubric | | **Halo** | One strong trait → high ratings everywhere | Force per-dimension scoring | | **Recency** | Last project crowds out year | Ask about a specific quarter from earlier | | **Similarity-to-self** | Rates people like themselves higher | Mixed-perspective panels | | **Maternal wall** | Mothers rated lower for ambiguous reasons | Watch returning-from-leave patterns | | **Stereotype activation** | Gender / race expectations shape interpretation | Demographic distribution review | | **Confidence vs. competence** | Confident self-presenters over-rated | Anchor on outcomes, not affect | ## Demographic Distribution Review Before final ratings lock: - Distribution by gender, race/ethnicity, age, tenure, level - Cell-size minimums (n ≥ 5–8) - Statistical significance of differences - Investigate causes; don't assume bias *or* dismiss If a manager consistently rates one demographic group lower with weak evidence: coach the manager; don't just adjust the rating (the underlying pattern persists). ## Forced Distribution Most companies should not force. Cost: corrodes collaboration; demoralizes; arbitrarily punishes; can violate fairness norms. If using guidance distributions (e.g., "expect roughly 70% Meets, 20% Exceeds, 10% Below"): - Treat as guidance, not requirement - Allow exceptions with rationale - Watch for managers who hit the distribution by manipulating rather than rating honestly ## Calibration Theater (Avoid) Signs: - Discussion happens; no ratings change - Managers defend their ratings; group acquiesces - "We can't change Maya's rating; her manager promised her" - Distribution review noted but not acted on - Demographic gaps observed but not addressed If calibration produces no ratings changes, it's not working. ## Post-Calibration - Comp partner re-runs band integrity and pay equity check - Final ratings communicated to employees by their managers - Documentation: rationale for changes; manager scorecards (who's calibrated well; who needs coaching) ## Common Failures - Single-manager ratings without calibration - Calibration limited to "the bottom and the top" (the middle has bias too) - Ratings discussed without rubric in front of the panel - No demographic distribution review - "Make manager Y rate harder" without coaching the underlying issue - Calibration after manager has already told the employee the rating ## Cross-References - `performance-management-systems` skill - `career-leveling` skill — anchors ratings - `pay-equity-analysis` skill — adjacent audit - `dei-strategist` agent — distribution review ## Key References - Bock, L. (2015). *Work Rules!* - Industry practitioner work on calibration (Lattice, Culture Amp guides)
View on GitHub