- name
- hiring-rubrics-scorecards
- description
- Use when designing scorecards, rubrics, and debrief mechanics for structured hiring.
# Hiring Rubrics and Scorecards
## Purpose
Scorecards capture independent judgments per interviewer before debrief. They turn intuition into evidence and prevent the loudest opinion from carrying the room.
## Scorecard Structure
Each interviewer's scorecard:
1. **Competencies assigned** (1–2 per interviewer, max)
2. **Specific anchors per competency** (what 4/4 looks like, what 1/4 looks like)
3. **Evidence cited**: quotes, observed behaviors, output reviewed
4. **Rating per competency** (4-point scale recommended)
5. **Overall recommendation**: Strong Yes / Yes / No / Strong No
6. **Notes for debrief**
## 4-Point Scale (Recommended)
- **4 — Strong Hire**: clear evidence of excellence at level; would make a noticeable positive impact
- **3 — Hire**: solid evidence; meets bar for level; room to grow
- **2 — No Hire**: gaps that would slow team or create coaching burden disproportionate to value
- **1 — Strong No Hire**: significant concerns; pattern of weakness in core competency
Avoid 5-point with neutral middle (everyone lands at 3, no information). Avoid 10-point (false precision).
## Behavioral Anchor Examples
### "Drives Outcomes" (Senior IC)
- 4: Owned a meaningful outcome end-to-end with ambiguity; concrete examples of scope changes or roadblocks navigated; evidence of follow-through
- 3: Owned outcomes with structure; clear examples of execution; some ambiguity navigated
- 2: Executed assigned tasks; less clarity on independent ownership
- 1: Difficulty articulating outcomes owned; deference patterns to others' direction
### "Collaborates Across Boundaries"
- 4: Examples of building trust with skeptical partners; conflict navigated openly; outcomes that required the partner's help
- 3: Worked productively with cross-functional partners; some examples of give-and-take
- 2: Worked alongside partners but limited evidence of joint outcomes
- 1: Patterns of friction; difficulty naming partners' perspectives
## Debrief Mechanics
### Before
- All scorecards submitted (a hard constraint)
- Hiring manager reviews scorecards
- Decision criteria stated up-front
### During (45–60 min)
1. Hiring manager opens with the role and decision criteria
2. Each interviewer presents (5–7 min): rating per competency + specific evidence
3. Discussion of disagreements; ground in evidence, not impression
4. Hiring manager summarizes the picture
5. Decision
### Anti-Patterns
- Decision before debrief (loudest voice wins)
- Vague verdicts ("good vibes" / "no")
- "Culture fit" as a tiebreaker without operationalization
- Over-weighting one impressive moment vs. the body of evidence
- Anchoring on the first opinion (rotate who speaks first)
## Verdict Language Guidance
Strong: "I see clear evidence of [competency] at this level — specifically [example]. I have one concern about [thing], but it's manageable. Recommend hire."
Weak: "Seemed fine. Smart. I think we should hire."
The first is replicable, auditable, debate-able. The second is opaque.
## Rubric for Levels
Each level (per `career-leveling`) gets a rubric for hiring purposes too. Hiring at L4 requires evidence of L4-level scope, ambiguity, and impact. Hiring at L5 requires evidence of L5-level scope. If the evidence supports L3 but the role is L4, the right call is "no hire" or "level the role to L3."
## Bias Patterns to Watch in Scoring
- **Halo**: one strong trait → high ratings everywhere
- **Recency**: last 5 minutes of a 60-min interview dominates
- **Similarity-to-self**: candidates resembling the interviewer rated higher
- **Confidence**: confident candidates rated higher than equally-skilled diffident ones
- **Style over substance**: polish weighed too heavily
- **Maternal wall**: parents rated lower for ambiguous reasons
- **Stereotype activation**: gender/race expectations shaping interpretation
Train interviewers on these. Calibration mitigates more than awareness alone.
## Calibration Sessions
Every 6 months at scale:
- Review scoring distributions by interviewer
- Identify outliers (consistently harsher / softer)
- Walk through anonymized scorecards together
- Re-anchor on what each rating means
- Update rubrics if drift evident
## Auditing Scorecards
Quarterly:
- Distribution of ratings by interviewer
- Quality of evidence (lots of "good vibes"? coach for specificity)
- Demographic patterns in scoring
- Outcome correlation: do the people we hired with strong scorecards perform better at 90 days, 12 months?
## Pitfalls
- "Strong yes" with one sentence of evidence
- Same template across role families (each family has different competencies)
- Rubrics that read like aspirational adjectives ("entrepreneurial," "passionate") without behavioral anchors
- Rubric created once, never revisited
- Scorecards filed and never analyzed
## Key References
- Bock, L. (2015). *Work Rules!*
- Smart, G., & Street, R. (2008). *Who: The A Method for Hiring*.
- Bohnet, I. (2016). *What Works*.
GitHubで見る