| name | prompt-reliability-score |
| description | Use this skill when the user wants a pre-launch reliability score or readiness audit of a Claude prompt across multiple quality dimensions, without adversarial testing. Triggers on phrases like "score this prompt", "is this ready to ship", "rate my prompt", "reliability audit", "launch readiness check", or requests for a dimensional breakdown of prompt quality. Do NOT use when the user wants adversarial failure-finding (use prompt-stress-test) or when they want the prompt rewritten (use prompt-refinement). |
Prompt Reliability Score
Rate a prompt across 5 dimensions. Flag launch risks. Do not rewrite.
Inputs required
- The prompt (verbatim)
- The use case it serves (one sentence)
Procedure
- Score across 5 dimensions, 1–10 each:
- Instruction clarity
- Output format specificity
- Constraint strength
- Edge case handling
- Tone consistency
- For each score, cite a specific phrase from the prompt as evidence.
- Calculate overall score as the arithmetic mean, rounded to one decimal.
- Flag every dimension scoring below 7 as a launch risk.
- Deliver the report. Do not suggest rewrites.
Rules
- Score against the stated use case, not in the abstract.
- Every score must be justified with a direct quote from the prompt.
- Dimensions below 7 are launch risks, not suggestions.
- Overall score is calculated, not estimated.
- If the use case is missing, ask before scoring — the same prompt scores differently across use cases.
Output format
- Dimension Scores — table: dimension | score | evidence quote
- Overall Reliability Score — X.X/10
- Launch Risk Flags — bulleted list of dimensions <7 with one-line rationale each
- Verdict — ship / hold / rework