Structure playtests to generate actionable data, not just "it was fun" or "it wasn't." Use this skill when: (1) planning a playtest session with clear objectives, (2) designing observation checklists and post-play questionnaires, (3) choosing between playtest types (blind, facilitated, A/B, stress test), (4) building data collection plans (quantitative vs. qualitative), (5) creating analysis frameworks to interpret results, (6) structuring feedback sessions to extract useful information, (7) validating whether the intended experience is landing with real players. Medium-agnostic — works for digital playtests, tabletop playtests, and hybrid testing.
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
A direct command skips the review prompt. Inspect the source before running it.
Structure playtests to generate actionable data, not just "it was fun" or "it wasn't." Use this skill when: (1) planning a playtest session with clear objectives, (2) designing observation checklists and post-play questionnaires, (3) choosing between playtest types (blind, facilitated, A/B, stress test), (4) building data collection plans (quantitative vs. qualitative), (5) creating analysis frameworks to interpret results, (6) structuring feedback sessions to extract useful information, (7) validating whether the intended experience is landing with real players. Medium-agnostic — works for digital playtests, tabletop playtests, and hybrid testing.
Playtest Protocol Designer
Playtesting without protocol is just playing. You'll have fun. You'll learn nothing. This skill structures playtests to generate the specific data needed to improve the game.
Playtest Types
Blind Playtest
Players receive the game with no explanation from the designer. They read the rules (or tutorial) and play.
When: Core mechanics work but robustness is unknown
Who: Experienced/adversarial players who will try to break things
Designer role: Encourage breaking. "Try to win in the cheapest way possible."
Focus Test
Players experience a specific element and provide targeted feedback.
Tests: Specific systems, narrative moments, aesthetic impact
When: A particular element needs validation in isolation
Who: Target audience for that element
Designer role: Present the element, ask targeted questions.
Core Deliverables
1. Playtest Objective
Every playtest must have a clear purpose:
PLAYTEST ID: [Identifier — e.g., PT-007]
TYPE: [Blind / Facilitated / A/B / Stress / Focus]
HYPOTHESIS: [What we're testing — connected to prototype hypothesis]
QUESTIONS: [Specific questions this playtest should answer]
NOT TESTING: [What we're explicitly NOT evaluating this session]
Anti-pattern: "Let's playtest and see what happens." If you don't know what you're looking for, you won't find it.
2. Observation Checklist
What to watch for during play:
Engagement indicators:
Players are making active decisions (not just going through motions)
Players are discussing strategy with each other
Players show emotional reactions (excitement, tension, laughter)
Players lean forward / show focused body language
Players ask "can I do X?" (exploring the possibility space)
Disengagement indicators:
Players check phones or look away
Players ask "whose turn is it?" (lost track of game state)
Players make choices without consideration (just picking randomly)
Players ask "how much longer?" or "are we almost done?"
Silence during what should be exciting moments
Confusion indicators:
Players ask the same question twice
Players make illegal moves without realizing
Players misunderstand a rule and play incorrectly
Players stare at their options for extended periods
Players ask each other for rule clarification instead of the rulebook
Balance indicators:
One player/strategy pulls ahead early and stays ahead
A player feels hopeless before the game ends
All players converge on the same strategy
A mechanic is used by nobody (dead feature)
A mechanic is used by everybody (possibly dominant)
3. Post-Play Questionnaire
Targeted questions, not generic. Design these per-playtest.
Universal questions (use every time):
What was the most fun moment? (identifies peaks)
What was the most frustrating moment? (identifies friction)
Would you play again? Why or why not? (overall verdict)
What confused you? (identifies clarity problems)
Did you feel like your decisions mattered? (agency check)
Hypothesis-specific questions (custom per test):
Design 3-5 questions that directly probe the hypothesis being tested
Use specific language, not vague ("Did you enjoy the random character?" not "Did you like it?")
Include one question that tests the negative ("What would you change about X?")
Question design rules:
No leading questions — "Wasn't the combat exciting?" vs. "How did the combat feel?"
No compound questions — One question, one topic
Scale + open-ended — "Rate combat satisfaction 1-5" THEN "Why that rating?"
Behavioral over opinion — "What did you do when X happened?" vs. "Did you like X?"
4. Data Collection Plan
What data to capture and how:
Quantitative data:
METRIC HOW CAPTURED WHAT IT TELLS US
────────────────── ──────────────── ──────────────────────────
Game duration Timer Pacing health
Turns to completion Counter Loop efficiency
Resource at end Snapshot Economy balance
Win margin Score diff Competitiveness
Choice distribution Tally marks Option viability
Qualitative data:
DATA TYPE HOW CAPTURED WHAT IT TELLS US
────────────────── ──────────────── ──────────────────────────
Player reactions Observer notes Emotional experience
Questions asked Written log Clarity problems
Strategies used Observer notes Design space utilization
Table talk Audio/notes Social dynamics
Post-game comments Questionnaire Reflective assessment
5. Analysis Framework
How to interpret results after the playtest:
Pattern identification:
Did multiple players encounter the same problem? (Systematic issue)
Did one player have a unique problem? (Edge case or skill gap)
Did observed behavior match self-reported experience? (If not, dig deeper)
Critical: Prevents play. Game breaks or stops making sense.
Major: Significantly harms experience. Players notice and complain.
Minor: Small friction. Players notice but continue.
Cosmetic: Polish issue. Doesn't affect gameplay.
Signal vs. noise: Not every player complaint requires a design change. Look for patterns across multiple playtests, not individual reactions. One player hating a mechanic is data. Five players hating it is a signal.
Facilitator Guidelines
Before the test:
Set up the game completely before players arrive
Prepare note-taking materials (paper, forms, or digital)
Brief observers on what to watch for
Decide: will you intervene if players are stuck? At what threshold?
During the test:
Resist the urge to explain or defend design choices
Note the TIME when interesting things happen (for cross-referencing)
Watch faces as much as the game board/screen
If you must intervene, note what triggered it — that's a design problem
After the test:
Run the questionnaire immediately (memory fades fast)
Ask open-ended questions before specific ones
Don't argue with feedback ("but that's intentional" is not useful)
Thank players genuinely — their time is valuable
Workflow
Planning a playtest
Define the objective — What question must this test answer?
Choose playtest type — What type matches the objective?
Build observation checklist — What to watch for, connected to the objective