Skip to main content
GitHub 仓库

product-eval

product-eval 收录了来自 sparkline-ventures 的 14 个 skills,并提供仓库级职业覆盖和站内 skill 详情页。

已收集 skills
14
Stars
14
更新
2026-06-29
Forks
0
职业覆盖
2 个职业分类 · 已分类 100%
仓库浏览

这个仓库中的 skills

calibrate
软件开发工程师

Tune and validate the scoring constants (evidence-weight base, diversity factors, gate threshold, crispness bars) against past decisions with known outcomes. Use when the user says "calibrate", "tune the thresholds", "validate the scoring", "the confidence numbers feel off", or has a set of past decisions to check the model against. Reports mismatches and recommended constant changes; never silently changes them.

2026-06-29
critique
项目管理专家

Use when someone wants existing thinking attacked, graded, or handed to fresh eyes, not written or summarized. The point is to break the artifact, not produce it. It owns three intents: a skeptical outside read of conclusions, findings, themes, a synthesis, a brief, or a FAQ, from someone who was not in the room; running those findings past hostile lenses (a wary customer, CFO, engineer, competitor) to see where they collapse; and grading a problem statement or one-liner (right altitude, names a real persona, or vague mush) with a score and what is missing. Reach for it on 'give this a fresh adversarial read', 'put this in front of a skeptical X', 'score how crisp this problem statement is', 'poke holes', 'red-team this'. Do not use it to create or restate the artifact. Attacking a whole plan, roadmap, or PRD routes through pressure-test; this skill owns findings, syntheses, problem statements, and panels.

2026-06-29
decision-readiness
项目管理专家

Judge whether there is enough evidence to make a build decision on a specific problem or bet, and return one of three outcomes: Decide now, Run a research sprint, or Do not commit yet. Use when the user asks "do we have enough data to decide?", "is this worth building?", "are we ready to commit to X?", "what's the evidence for this?", or "decision readiness". This is the core gate, Confidence band drives the decision; signal strength is assessed per evidence item.

2026-06-29
design-test
项目管理专家

Turn a "Run a research sprint" verdict (promising but under-evidenced) into the cheapest experiment that would close the gap, a concrete validating test with a hypothesis, method, sample, effort, and pass/fail threshold. Use when decision-readiness returns Run a research sprint, or when the user asks "how do we validate this", "what's the cheapest test", "how do we de-risk this bet", or has a promising idea with thin first-party evidence. Designs the test; its results feed back into gather-evidence. Do NOT use to pull existing data (that is gather-evidence) or to judge sufficiency (that is decision-readiness).

2026-06-29
discover-sources
项目管理专家

Map what product-data sources a user can draw on, connectors, tools, files, offline systems, and the decision-confidence ceiling those sources allow. Use when the intent is discovery and setup of the evidence base itself: "what data do I have / could connect", "scan/inventory my sources", "is X hooked up", "where could product insights come from", "how confident could any decision be given my data", or starting a new project before any framing or prioritizing. Do NOT use when the user wants you to actually work the data they already have, pulling, summarizing, grading, or reporting on tickets, reviews, surveys, metrics, scorecards, or trends. Those are analysis and reporting tasks, not source discovery. The trigger is "what evidence exists and how good can it get," not "analyze the evidence." Produces a source map plus a confidence ceiling.

2026-06-29
drift-check
项目管理专家

Compare an investigation's CURRENT state against an EARLIER SNAPSHOT of itself and report what moved between those two points in time: problem-statement drift, evidence added or retired, Value/Confidence score movement, theme or narrative reframing, and whether confidence is genuinely accumulating or just churning. Use when the user says "what changed since last month", "has this drifted", "compare this to the prior version", "is the confidence trajectory improving", or wants to audit a scope before re-deciding. Reads scoped snapshots and git history and runs the cross-stage consistency checks. Do NOT use for finding patterns or trends WITHIN the current evidence set, that is synthesize; nor for reviewing a product-metrics dashboard, that is a metrics review. This is specifically a then-vs-now diff of the investigation. No daemon; runs on demand or on a schedule.

2026-06-29
gather-evidence
项目管理专家

Pull product evidence from connected sources, CSV uploads, and the web; then normalize, identity-resolve, strength-rate, and weight it into scoped evidence items. Use when the user says "gather evidence", "pull the tickets/reviews/analytics", "find signal for this problem", "collect data on X", or as part of Rank opportunities / Decide on one bet. Writes evidence items and a resolved identity map that synthesize and decision-readiness consume.

2026-06-29
pressure-test
项目管理专家

Use when the user wants an existing plan attacked rather than built, they have a roadmap, plan, PRD, bet, or strategy and want you to find what's wrong with it before they commit. Triggers on any adversarial-critique intent: tear it apart, find the holes, poke holes, play devil's advocate, stress-test, pressure-test, challenge the assumptions, what are we missing / not seeing, where will this break, why might this fail, talk me out of this, red-team it. The defining signal is skeptical scrutiny of a decision the user is about to make. Surfaces thin evidence, hidden or unexamined assumptions, wrong scope, and drift, then returns blocking issues, suggestions, strengths, and a ready/needs-fixes/not-ready verdict. Do NOT use for generating new ideas, brainstorming directions, writing or updating a plan/PRD/roadmap, or reviewing metrics and status updates, this only critiques something that already exists.

2026-06-29
prioritize-backlog
项目管理专家

Rank opportunities or backlog items by evidence-backed Value and Confidence bands, storing exact 0-100 scores for audit. Use when the user says "prioritize my backlog", "rank these", "what should we build next", "score these opportunities", or hands over a list of problems/features to sequence. This is the Rank opportunities outcome, distinct from updating a roadmap's Now/Next/Later. Falls back to proxies when data is thin and flags lower confidence.

2026-06-29
scorecard
项目管理专家

Render the results of a product-eval investigation as a self-contained interactive HTML page: ranked opportunities with their Value, Confidence, verdict, evidence weight, customer quotes, and open gaps, laid out as KPI cards, a Value by Confidence quadrant, and per-opportunity cards. Use when the user wants to visualize, render, share, or make a dashboard or scorecard from an existing product-eval scope, that is, the output of prioritize-backlog or decision-readiness, or a scope's scores.md, themes.md, and decisions-log.md. Triggers include 'decision dashboard', 'product dashboard', 'scorecard from the ranked opportunities', 'turn the prioritization into something visual to share'. Do NOT use for general metrics or KPI dashboards, north-star or quarterly metric trends, grading raw signal strength, checking connected data sources, or drafting briefs; this only renders an already-scored opportunity ranking, it does not gather data or compute scores.

2026-06-29
start
项目管理专家

Use as the front door when the user is getting oriented rather than requesting a finished deliverable. Trigger on bare openers ('start', 'where do I begin', 'help me get going', 'just installed this'), questions about what this plugin or tool does, requests to survey or kick off an investigation into a space or area, or any broad product goal stated without a single concrete action ('find our next bet', 'what should we work on'). The intent is 'point me in the right direction', not 'do this specific task'. Route them toward one of four outcomes: rank opportunities, decide on one bet, pressure-test a plan, or write an output. Do NOT use when the user names a concrete deliverable and the inputs for it: synthesize provided evidence, rank a given list, judge one named bet, pressure-test a pasted plan, or draft a specific brief, memo, or spec. When in doubt, if there is a clear single action to perform, stay out.

2026-06-29
synthesize
项目管理专家

Use when turning gathered evidence (research, tickets, reviews, interview notes, analytics, or weighted evidence items in a folder) into a small set of problem-framed themes or clusters. Also use when someone has a vague observation, complaint, or signal (e.g. 'users keep dropping off in signup') and wants it shaped into a real, crisp problem statement at the right altitude. Triggers on: cluster or group evidence into problems, find the themes or recurring patterns, see what is trending over time, frame this up, is this the right problem, what should we actually build, or feeding candidates into prioritization. Reads scoped evidence, clusters by underlying problem (not topic), detects trends and corroboration, reasons about confidence, names gaps, and writes buildable problem statements that prioritize-backlog can score. Not for gathering or scoring evidence, only for making sense of evidence already collected, or framing a loose pain into a structured problem.

2026-06-29
to-build-brief
软件开发工程师

Transform an approved decision into an agent-ready build spec, --code for Claude Code/Codex, --design for Figma/wireframe/Gemini. Use when the user says "turn this into a spec", "make it buildable", "build brief", "hand this to Claude Code / Figma", or after a brief is approved. Strips narrative voice; outputs structured, testable markdown.

2026-06-29
write-brief
项目管理专家

Use when the user wants the human-facing decision document for a problem or bet that has been investigated, to align people on it or to force a choice. Covers three shapes: a narrative or report (problem, why now, impact, and the ask), a hard-questions FAQ, or a decision or options memo that weighs several options (including do-nothing), recommends one, and states the ask. Trigger on phrasing like 'write the brief', 'draft the brief', 'opportunity brief', 'the narrative', 'the FAQ', 'decision memo', 'options memo', 'recommendation memo', 'weigh options and recommend one', or after a decision passes the readiness gate. This is the leadership decision artifact, rendered from existing evidence and scores, NOT an engineering PRD, spec, acceptance criteria, status update, competitive battlecard, or metrics review. Refuses to invent metrics.

2026-06-29