| name | hiring-rubrics-scorecards |
| description | Use when designing scorecards, rubrics, and debrief mechanics for structured hiring. |
Hiring Rubrics and Scorecards
Purpose
Scorecards capture independent judgments per interviewer before debrief. They turn intuition into evidence and prevent the loudest opinion from carrying the room.
Scorecard Structure
Each interviewer's scorecard:
- Competencies assigned (1–2 per interviewer, max)
- Specific anchors per competency (what 4/4 looks like, what 1/4 looks like)
- Evidence cited: quotes, observed behaviors, output reviewed
- Rating per competency (4-point scale recommended)
- Overall recommendation: Strong Yes / Yes / No / Strong No
- Notes for debrief
4-Point Scale (Recommended)
- 4 — Strong Hire: clear evidence of excellence at level; would make a noticeable positive impact
- 3 — Hire: solid evidence; meets bar for level; room to grow
- 2 — No Hire: gaps that would slow team or create coaching burden disproportionate to value
- 1 — Strong No Hire: significant concerns; pattern of weakness in core competency
Avoid 5-point with neutral middle (everyone lands at 3, no information). Avoid 10-point (false precision).
Behavioral Anchor Examples
"Drives Outcomes" (Senior IC)
- 4: Owned a meaningful outcome end-to-end with ambiguity; concrete examples of scope changes or roadblocks navigated; evidence of follow-through
- 3: Owned outcomes with structure; clear examples of execution; some ambiguity navigated
- 2: Executed assigned tasks; less clarity on independent ownership
- 1: Difficulty articulating outcomes owned; deference patterns to others' direction
"Collaborates Across Boundaries"
- 4: Examples of building trust with skeptical partners; conflict navigated openly; outcomes that required the partner's help
- 3: Worked productively with cross-functional partners; some examples of give-and-take
- 2: Worked alongside partners but limited evidence of joint outcomes
- 1: Patterns of friction; difficulty naming partners' perspectives
Debrief Mechanics
Before
- All scorecards submitted (a hard constraint)
- Hiring manager reviews scorecards
- Decision criteria stated up-front
During (45–60 min)
- Hiring manager opens with the role and decision criteria
- Each interviewer presents (5–7 min): rating per competency + specific evidence
- Discussion of disagreements; ground in evidence, not impression
- Hiring manager summarizes the picture
- Decision
Anti-Patterns
- Decision before debrief (loudest voice wins)
- Vague verdicts ("good vibes" / "no")
- "Culture fit" as a tiebreaker without operationalization
- Over-weighting one impressive moment vs. the body of evidence
- Anchoring on the first opinion (rotate who speaks first)
Verdict Language Guidance
Strong: "I see clear evidence of [competency] at this level — specifically [example]. I have one concern about [thing], but it's manageable. Recommend hire."
Weak: "Seemed fine. Smart. I think we should hire."
The first is replicable, auditable, debate-able. The second is opaque.
Rubric for Levels
Each level (per career-leveling) gets a rubric for hiring purposes too. Hiring at L4 requires evidence of L4-level scope, ambiguity, and impact. Hiring at L5 requires evidence of L5-level scope. If the evidence supports L3 but the role is L4, the right call is "no hire" or "level the role to L3."
Bias Patterns to Watch in Scoring
- Halo: one strong trait → high ratings everywhere
- Recency: last 5 minutes of a 60-min interview dominates
- Similarity-to-self: candidates resembling the interviewer rated higher
- Confidence: confident candidates rated higher than equally-skilled diffident ones
- Style over substance: polish weighed too heavily
- Maternal wall: parents rated lower for ambiguous reasons
- Stereotype activation: gender/race expectations shaping interpretation
Train interviewers on these. Calibration mitigates more than awareness alone.
Calibration Sessions
Every 6 months at scale:
- Review scoring distributions by interviewer
- Identify outliers (consistently harsher / softer)
- Walk through anonymized scorecards together
- Re-anchor on what each rating means
- Update rubrics if drift evident
Auditing Scorecards
Quarterly:
- Distribution of ratings by interviewer
- Quality of evidence (lots of "good vibes"? coach for specificity)
- Demographic patterns in scoring
- Outcome correlation: do the people we hired with strong scorecards perform better at 90 days, 12 months?
Pitfalls
- "Strong yes" with one sentence of evidence
- Same template across role families (each family has different competencies)
- Rubrics that read like aspirational adjectives ("entrepreneurial," "passionate") without behavioral anchors
- Rubric created once, never revisited
- Scorecards filed and never analyzed
Key References
- Bock, L. (2015). Work Rules!
- Smart, G., & Street, R. (2008). Who: The A Method for Hiring.
- Bohnet, I. (2016). What Works.