| name | playtest-protocol-designer |
| description | Structure playtests to generate actionable data, not just "it was fun" or "it wasn't." Use this skill when: (1) planning a playtest session with clear objectives, (2) designing observation checklists and post-play questionnaires, (3) choosing between playtest types (blind, facilitated, A/B, stress test), (4) building data collection plans (quantitative vs. qualitative), (5) creating analysis frameworks to interpret results, (6) structuring feedback sessions to extract useful information, (7) validating whether the intended experience is landing with real players. Medium-agnostic โ works for digital playtests, tabletop playtests, and hybrid testing.
|
Playtest Protocol Designer
Playtesting without protocol is just playing. You'll have fun. You'll learn nothing. This skill structures playtests to generate the specific data needed to improve the game.
Playtest Types
Blind Playtest
Players receive the game with no explanation from the designer. They read the rules (or tutorial) and play.
- Tests: Rulebook clarity, onboarding, independent learnability
- When: Rules are drafted, UI/tutorial exists
- Who: Players who have never seen the game
- Designer role: Silent observer. Do not help. Do not explain. Take notes.
Facilitated Playtest
Designer teaches the game and observes play, answering questions.
- Tests: Core mechanics, balance, engagement, flow
- When: Prototype stage, mechanics are functional but rules may be rough
- Who: Any players, including familiar ones
- Designer role: Teach, then observe. Note every question asked โ those are design failures.
A/B Playtest
Two groups play different versions of the same game element.
- Tests: Specific design decisions (version A vs. version B)
- When: A specific design question has two viable answers
- Who: Comparable skill-level groups
- Designer role: Set up both, observe both, compare data.
Stress Test
Push the game to breaking points deliberately.
- Tests: Edge cases, degenerate strategies, scaling limits
- When: Core mechanics work but robustness is unknown
- Who: Experienced/adversarial players who will try to break things
- Designer role: Encourage breaking. "Try to win in the cheapest way possible."
Focus Test
Players experience a specific element and provide targeted feedback.
- Tests: Specific systems, narrative moments, aesthetic impact
- When: A particular element needs validation in isolation
- Who: Target audience for that element
- Designer role: Present the element, ask targeted questions.
Core Deliverables
1. Playtest Objective
Every playtest must have a clear purpose:
PLAYTEST ID: [Identifier โ e.g., PT-007]
TYPE: [Blind / Facilitated / A/B / Stress / Focus]
HYPOTHESIS: [What we're testing โ connected to prototype hypothesis]
QUESTIONS: [Specific questions this playtest should answer]
NOT TESTING: [What we're explicitly NOT evaluating this session]
Anti-pattern: "Let's playtest and see what happens." If you don't know what you're looking for, you won't find it.
2. Observation Checklist
What to watch for during play:
Engagement indicators:
Disengagement indicators:
Confusion indicators:
Balance indicators:
3. Post-Play Questionnaire
Targeted questions, not generic. Design these per-playtest.
Universal questions (use every time):
- What was the most fun moment? (identifies peaks)
- What was the most frustrating moment? (identifies friction)
- Would you play again? Why or why not? (overall verdict)
- What confused you? (identifies clarity problems)
- Did you feel like your decisions mattered? (agency check)
Hypothesis-specific questions (custom per test):
- Design 3-5 questions that directly probe the hypothesis being tested
- Use specific language, not vague ("Did you enjoy the random character?" not "Did you like it?")
- Include one question that tests the negative ("What would you change about X?")
Question design rules:
- No leading questions โ "Wasn't the combat exciting?" vs. "How did the combat feel?"
- No compound questions โ One question, one topic
- Scale + open-ended โ "Rate combat satisfaction 1-5" THEN "Why that rating?"
- Behavioral over opinion โ "What did you do when X happened?" vs. "Did you like X?"
4. Data Collection Plan
What data to capture and how:
Quantitative data:
METRIC HOW CAPTURED WHAT IT TELLS US
โโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโ
Game duration Timer Pacing health
Turns to completion Counter Loop efficiency
Resource at end Snapshot Economy balance
Win margin Score diff Competitiveness
Choice distribution Tally marks Option viability
Qualitative data:
DATA TYPE HOW CAPTURED WHAT IT TELLS US
โโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโ
Player reactions Observer notes Emotional experience
Questions asked Written log Clarity problems
Strategies used Observer notes Design space utilization
Table talk Audio/notes Social dynamics
Post-game comments Questionnaire Reflective assessment
5. Analysis Framework
How to interpret results after the playtest:
Pattern identification:
- Did multiple players encounter the same problem? (Systematic issue)
- Did one player have a unique problem? (Edge case or skill gap)
- Did observed behavior match self-reported experience? (If not, dig deeper)
Severity classification:
FINDING SEVERITY ACTION
โโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโ
[Description] [Critical/Major/ [Immediate fix / Next
Minor/Cosmetic] iteration / Backlog /
Monitor]
Critical: Prevents play. Game breaks or stops making sense.
Major: Significantly harms experience. Players notice and complain.
Minor: Small friction. Players notice but continue.
Cosmetic: Polish issue. Doesn't affect gameplay.
Signal vs. noise: Not every player complaint requires a design change. Look for patterns across multiple playtests, not individual reactions. One player hating a mechanic is data. Five players hating it is a signal.
Facilitator Guidelines
Before the test:
- Set up the game completely before players arrive
- Prepare note-taking materials (paper, forms, or digital)
- Brief observers on what to watch for
- Decide: will you intervene if players are stuck? At what threshold?
During the test:
- Resist the urge to explain or defend design choices
- Note the TIME when interesting things happen (for cross-referencing)
- Watch faces as much as the game board/screen
- If you must intervene, note what triggered it โ that's a design problem
After the test:
- Run the questionnaire immediately (memory fades fast)
- Ask open-ended questions before specific ones
- Don't argue with feedback ("but that's intentional" is not useful)
- Thank players genuinely โ their time is valuable
Workflow
Planning a playtest
- Define the objective โ What question must this test answer?
- Choose playtest type โ What type matches the objective?
- Build observation checklist โ What to watch for, connected to the objective
- Design questionnaire โ Universal + hypothesis-specific questions
- Set up data collection โ What metrics, how captured
- Recruit players โ Match player profile to test needs
- Prepare the prototype (Skill 15) โ Ensure it's ready
- Brief observers โ Everyone knows what to watch for
- Run the test
- Analyze results โ Pattern identification, severity classification
- Report findings โ Actionable items, not just observations
- Flag to iteration (Skill 17) โ Log what was found and what changes
Interpreting ambiguous results
- Is the sample size sufficient? (1 playtest is anecdotal; 5+ is a pattern)
- Are players representative of the target audience?
- Did the prototype fidelity affect results? (Testing the wrong thing)
- Is this a design problem or a communication problem?
- Would watching one more test clarify, or is it time to iterate?
Outputs
This skill produces:
- Playtest protocol โ objective, type, checklist, questionnaire, data plan
- Observation checklist โ engagement, disengagement, confusion, balance
- Post-play questionnaire โ universal + custom questions
- Data collection plan โ metrics and capture methods
- Analysis framework โ pattern identification and severity classification
These outputs feed into:
- Skill 11 (Balance) โ Balance hypotheses get tested
- Skill 17 (Iteration) โ Playtest results drive design changes
- Skill 4 (Experience) โ Did the intended experience land?
- Skill 0 (Coherence) โ Results may trigger re-evaluation cascades