performance-calibration
Use when running calibration sessions — covers panel composition, bias patterns, distribution norms, and facilitation.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Use when running calibration sessions — covers panel composition, bias patterns, distribution norms, and facilitation.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Use when designing or applying progressive discipline — covers the escalating sequence, documentation, gross misconduct exceptions, and consistency.
Use when conducting or scoping workplace investigations — covers intake, scope, interview protocols, evidence, findings, and outcomes.
Use when auditing or designing the candidate's journey — covers communication, time, transparency, and the link between candidate experience and employer brand.
Use when designing career frameworks, level rubrics, IC and management track parity, and promotion processes.
Use when planning organizational change — covers communication architecture, sequencing, manager enablement, and predictable failure modes.
Use when training managers in coaching skills — covers GROW, listening, asking vs. telling, and growth mindset framing.
| name | performance-calibration |
| description | Use when running calibration sessions — covers panel composition, bias patterns, distribution norms, and facilitation. |
Calibration ensures rubrics are applied consistently across managers. Without it, two equally-strong employees can get materially different ratings (and pay) based on which manager they happen to report to. The cost is unfairness, broken trust, and pay equity drift.
Panels of 12+ become hard to facilitate. Panels of 3 become group-think.
| Pattern | Description | Counter |
|---|---|---|
| Leniency | Manager rates whole team high | Look at distribution; ask for differentiation |
| Harshness | Manager rates whole team low | Compare evidence to anchored rubric |
| Halo | One strong trait → high ratings everywhere | Force per-dimension scoring |
| Recency | Last project crowds out year | Ask about a specific quarter from earlier |
| Similarity-to-self | Rates people like themselves higher | Mixed-perspective panels |
| Maternal wall | Mothers rated lower for ambiguous reasons | Watch returning-from-leave patterns |
| Stereotype activation | Gender / race expectations shape interpretation | Demographic distribution review |
| Confidence vs. competence | Confident self-presenters over-rated | Anchor on outcomes, not affect |
Before final ratings lock:
If a manager consistently rates one demographic group lower with weak evidence: coach the manager; don't just adjust the rating (the underlying pattern persists).
Most companies should not force. Cost: corrodes collaboration; demoralizes; arbitrarily punishes; can violate fairness norms.
If using guidance distributions (e.g., "expect roughly 70% Meets, 20% Exceeds, 10% Below"):
Signs:
If calibration produces no ratings changes, it's not working.
performance-management-systems skillcareer-leveling skill — anchors ratingspay-equity-analysis skill — adjacent auditdei-strategist agent — distribution review