pm-evaluation-framework
pm-evaluation-framework contains 23 collected skills from kalyvask, with repository-level occupation coverage and site-owned skill detail pages.
Skills in this repository
Walk a founder or early-stage PM through finding product-market fit on a specific bet. Use when the user is pre-PMF (no validated value hypothesis), has a candidate insight + audience + business model, and needs help deciding what to test next, how to interpret a failed experiment, and whether to pivot (change who/how) or restart (change what). Different from pm-value-hypothesis-tester which pressure-tests a single hypothesis statement statically; this skill runs the multi-step PMF process iteratively. Built on the Lean Startup synthesis (Rachleff, Ries, Blank, Moore, Christensen, Marks, Fitzpatrick, Cook).
Critically review the user-facing surface of a PRD, design spec, or feature proposal against behavioral and UX principles. Use when a PM has a solution in hand and wants the design layer pressure-tested — defaults, friction placement, choice architecture, information density, AI surface decisions, peak-and-end moments. Distinct from pm-red-team (strategy adversary) and pm-evaluator (rubric scoring); this skill stays at the design layer and asks whether the design works with human cognition or against it. Returns the three load-bearing design holes with specific re-writes.
Audit a status update, exec review, board email, all-hands talking point, or dashboard callout for credibility leaks before sending. Flags claims that overstate ("shipped" when 5 users are in beta), cherry-picked windows ("+15% w/w" off a holiday), vanity metrics used as validation, attribution claims with obvious uncontrolled counterfactuals, and "on track" forecasts without a threshold. Use when a PM is about to communicate progress upward and wants every claim pressure-tested. Built on the principle that goodwill is a finite budget — every overclaim debits the account, and once drained the real wins stop being believed.
Pick a single North Star metric for a product. Weighs candidates across behavioral (MAU, sessions), value-delivered (messages sent, orders completed), and financial (ARR, revenue), forcing the comparison most teams skip — explainability, adoption + retention coverage, lead/lag, gameability. Use when a PM is defining a North Star, debating between MAU/DAU vs. value/revenue, or revisiting an existing one. Pushes back on jumping to a sophisticated composite or to revenue too early. Different from pm-metrics-critic, which reviews whole dashboards; this skill picks the one number.
Given a feature scope description, recommend the appropriate design-process intensity. Routes large features (M/L/XL) through the full Research → PRD → Design Concept → Detailed Design → Working Code pipeline; routes small features and bug fixes (XS/S) through an Express Lane (verbal sign-off, no Figma) that skips stages without value. Catches both over-processing of small work and under-processing of large work. Use at kickoff when scoping a new feature, or when triaging a backlog where the design team is over-loaded and you need to decide where to invest the process.
Stress-test a design, flow, or PRD section by simulating how a specific persona would walk through it step by step. Takes the design plus a target persona (defined inline or pulled from pm-state stakeholders) and outputs a walkthrough from the persona's point of view — what they see, what they think, where they get stuck, the first failure mode that would make them bounce. Designed to surface obvious failures in 10 minutes before committing engineering time or scheduling real user research. Distinct from pm-design-critic (which applies behavioral principles to the artifact) and pm-customer-interview-coach (which helps run interviews with real users); this skill simulates the user.
Load the PM agent's persistent context — personal style and preferences from `you.md`, the active project list, cross-project stakeholders, and the relevant project's state. Use as the first step in any PM workflow when you need the agent to understand the user's current work before doing anything else. The other operate-stage skills (pm-morning-brief, pm-meeting-prep, pm-meeting-debrief, pm-weekly-review, pm-stakeholder-tracker) chain to this skill first.
Pull recent Gmail threads, classify them into action categories (substantive response needed / quick reply / FYI / can ignore), and draft replies for the substantive ones. Surfaces stale threads where the user is awaiting a response and stale threads where someone is awaiting them. Used as part of pm-morning-brief or invoked directly when the inbox feels out of control. Drafts only — never auto-sends.
Given a Granola meeting transcript, extract commitments (yours and theirs), identify decisions made, draft follow-up messages, and update the relevant project's `todos.md` and `stakeholders.md`. Use right after a meeting to close the loop before the context fades. Surfaces commitments that need to be captured into the project state, drafts the follow-ups, and asks the user to approve writes.
Given a calendar event (or upcoming meeting context), draft a one-page meeting brief. Identifies attendees, pulls relevant project state and stakeholder context, surfaces last interaction and open commitments, and drafts an opinionated prep doc. Use 15-30 minutes before any meeting that matters — exec review, stakeholder check-in, customer call, decision point.
Run the morning PM briefing. Pulls today's calendar, identifies meetings needing prep, surfaces commitments due today across projects, surfaces inbox threads needing attention, and identifies anything that slipped from yesterday's commitments. Writes the brief to `pm-state/inbox/YYYY-MM-DD-morning-brief.md`. Use first thing in the morning, before opening email or calendar.
Cross-project stakeholder tracking — who you owe responses to, who owes you, last interaction date, open commitments per person. Reads cross-project and per-project stakeholder files plus recent Granola and Gmail signal. Surfaces relationships drifting toward neglect and commitments hanging in the gap between meetings. Use weekly or before any big stakeholder push (board prep, exec review, customer escalation cycle).
Run the Friday PM weekly review — what got done, what slipped, what's worth renegotiating, what's on next week. Reads each active project's state, pulls the week's Granola transcripts and calendar, surfaces patterns across projects (stakeholders going cold, decisions deferred, todos rotting). Writes the review to `pm-state/inbox/YYYY-MM-DD-weekly-review.md`. Use late Friday afternoon before logging off.
Critique an activation funnel, conversion funnel, trial-to-paid mechanic, or growth dashboard against the repo's activation and conversion content. Use when a PM has funnel data, an onboarding flow, a paywall design, or a growth dashboard and wants the funnel layer pressure-tested — drop-off drivers, the binding stage, whether the right model (Enterprise vs PLG) is being applied, whether the activation event is real or vanity. Distinct from pm-metrics-critic (metric layer) and pm-design-critic (surface layer); this skill sits between them at the funnel layer. Returns the three load-bearing funnel holes with specific experiments and kill thresholds.
Walk a PM through a real product decision step by step — diagnosis, framing, framework application, validation plan, kill criteria. Use when the user is actively wrestling with a specific product decision (kill / continue / pivot, build vs. buy, scope a launch, prioritize a backlog, respond to flat metrics) and wants help thinking it through end-to-end. Different from pm-framework-selector — this skill *runs* the decision; the selector just points at the right framework.
Grade a PM's written analysis, strategy memo, PRD, or proposal against the five-criterion PM evaluation rubric. Use when the user shares a PM artifact (a memo, a deck draft, a PRD, an analysis of a real product situation) and wants honest critique — or when reviewing your own draft before sending it up the chain. Returns a score, the strongest sections, the weakest sections, and specific re-work recommendations.
Recommend which product-management framework(s) to apply to a specific decision. Use when the user describes a product decision and asks "what framework should I use?" / "how should I think about this?" / "which lens applies?" — or when you spot a PM decision in a conversation that would benefit from explicit framework grounding (prioritization, MVP scoping, validation, growth, risk, problem-framing). Anchored in the lifecycle-phase frameworks library in this repo.
Draft or critique a PRD against the repo's PRD template and problem-framing rubric. Use when the user is starting a PRD, sharing a PRD draft for review, or stuck on a specific PRD section (problem statement, success metrics, scope, kill criteria). Pushes back on vague problem statements, generic success metrics, and feature-laundry-list scope. Returns either a structured PRD draft or a section-by-section critique with specific re-writes.
Adversarially re-review a PM artifact, recommendation, or AI-generated critique that already exists. Use as a second pass after another skill (pm-evaluator, pm-prd-drafter, pm-decision-coach, pm-value-hypothesis-tester) has produced output, or on any external AI output the user wants pressure-tested before deferring to it. Plays the role of a hostile exec, skeptical board member, or competing PM — looking for what the first pass missed, what bias it brought, and what would not survive a real review. Returns the three load-bearing holes, what's already strong enough to keep, and the specific re-writes that would close the gaps.
Plan, critique, or stress-test a customer-discovery interview against the Mom Test rules. Use when the user is preparing an interview script, reviewing a transcript, debriefing what they "learned," or deciding what to do with a batch of customer conversations. Catches leading questions, fluff-inducing phrasing, compliments mistaken for data, premature zoom into a hypothesized problem, and missed commitment asks. Returns a question-by-question rewrite, a list of what was actually learned vs. imagined, and the next conversation to run.
Pressure-test a value hypothesis (the what / who / how / why-now) before resources are committed, design the smallest experiment that would falsify it, and pre-commit kill criteria. Use when the user is about to launch, fundraise, or scale a product and wants the bet stress-tested — or when they're stuck on a fuzzy "we'll figure it out" hypothesis. Catches insight-free products, segments that are needy but not desperate, hedged "who" decisions, and skipped early-adopter beachheads. Returns a sharpened hypothesis, the falsifying experiment, and the kill threshold.
Review a launch plan against the repo's launch-criteria template and failure-management playbook. Use when the user is preparing a launch (closed beta, limited GA, full GA) and wants a pre-flight review, or is staring at a checklist of "ready" gates that all happen to be owned by the same person. Surfaces missing gates, unowned criteria, vague readiness bars, and the one or two failure modes most likely to show up post-launch. Returns a gate-by-gate verdict and a launch-or-defer recommendation.
Critique a metrics dashboard, success-criteria section, or proposed North Star metric against the repo's metrics guide. Use when the user shares a dashboard, a list of KPIs, a success-criteria slide, or asks "are these the right metrics?" Surfaces the three most common metric failures: aggregate metrics that hide segment-level failure, vanity metrics decoupled from strategy assumptions, and standalone metrics with no falsifiable counter. Returns a metric-by-metric verdict and a recommended replacement set.