一键导入
evaluate-content
Use when judging content quality OR editing/improving existing copy: shareability, readability, voice, cuttability, angle, copy sweeps.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Use when judging content quality OR editing/improving existing copy: shareability, readability, voice, cuttability, angle, copy sweeps.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Verify a code change actually works by building/running the app and observing it at its real surface (CLI, API, UI, library, agent), capturing runtime evidence rather than trusting tests. Make sure to use this skill whenever the user has changed code and wants to know it works, is about to merge/push and wants confidence, says "did this actually work", "verify this works", "prove it works", "confirm the change", "make sure it works", wants runtime evidence, or is re-running tests / importing-and-calling just to check behavior, even if they never say the word "verify". When in doubt after any code change, reach for this. For post-deploy production health checks use verify-deploy; for static correctness/quality review use simplify or code-review.
Test an interactive lesson/course (or any "instructions to an AI" skill) by self-play. An agent plays BOTH the instructor following the lesson script AND a calibrated student persona, producing full turn-by-turn transcripts of every lesson, then publishes the raw transcripts to a single static page. Use when asked to "run lesson transcripts", "test the course end to end", "self-play the lessons", "publish raw test transcripts", "walk a synthetic student through every lesson", or to QA an interactive-instruction skill by actually running it rather than just reviewing findings. Distinct from dogfood and adversarial-ux-test (web-app browser QA) and synthetic-userstudies (findings plus a few cherry-picked transcripts). This one captures the COMPLETE run of every lesson and ships them all raw.
Daily memory garbage collection for MEMORY.md / USER.md. Apply decay rules, drain .pending.md, consolidate near-duplicates, maintain canonical theme tags, prune old episode and session files. Invoke when asked to "run memory GC", "clean up memory", "apply memory decay", or from the scheduled cron job.
Monitor a PR until it's ready to merge. Watches CI, reads reviews, checks scope, fixes blocking issues, opens follow-up PRs for low-priority comments, and repeats. Use when: babysit this PR, watch this PR, monitor PR, fix and watch PR, keep this PR green.
Review code changes by analyzing git diffs, leaving inline comments on PRs, and performing thorough pre-push review. Works with the gh CLI or falls back to git + GitHub REST API via curl. Use when: review this PR, review a pull request, look at this PR, code review, review my changes before I push, pre-push review, leave inline comments on a pull request.
Run iterative self-referential development loops using the Ralph Wiggum technique. Use when tasks need repeated iteration, TDD cycles, greenfield builds, or autonomous refinement until tests pass or completion criteria are met. Triggers on ralph loop, ralph mode, iterative loop, autonomous loop, /goal.
| name | evaluate-content |
| description | Use when judging content quality OR editing/improving existing copy: shareability, readability, voice, cuttability, angle, copy sweeps. |
Honest, brutal content evaluation. Run this on anything before it goes live — articles, threads, tweets, headlines, outlines.
| Task | Framework |
|---|---|
| Evaluating content quality | The Six Questions |
| Editing/improving existing copy | The Seven Sweeps |
| Quick polish pass | Quick-Pass Editing Checks |
| Monthly content review | Pillar-Level Evaluation |
| Video hooks | Short-Form Video Hook Checklist |
Every piece of content gets judged on these six questions. Score each 1-5 and explain why.
What you're really asking: Is this interesting enough that someone would text it to a friend unprompted?
Score 5: "I'm sending this to three people right now." It has a surprising fact, a spicy take, or it articulates something people feel but haven't been able to say. Score 3: "Hm, interesting." They'd read it but not forward it. It's informative but not remarkable. Score 1: "Why would I share this?" Generic, nothing new, reads like a summary of things people already know.
Tests:
Common failures:
What you're really asking: Does the reader get something from this — either knowledge they didn't have, or entertainment they enjoyed?
Score 5: Reader learned something specific they can use, OR they genuinely enjoyed reading it. Ideally both. Score 3: Some useful info but presented in a way that's easy to skim-and-forget. Or mildly entertaining but no substance. Score 1: Neither. It's a slog to get through and you don't learn anything new.
Tests:
Common failures:
What you're really asking: Does this sound like a human wrote it, or does it sound like AI / a corporate blog / a college essay?
Score 5: You can hear the author's voice. Irregular sentences. Personal asides. Opinions stated directly. You'd recognize this writer's style blind. Score 3: Mostly natural but has some stiff patches. Occasional "it's important to note" or "in this article we'll explore" slipping through. Score 1: Full robot. Uniform sentence length. No personality. Could've been written by any AI or any corporate comms team.
Tests:
~/marketing/WRITING-STYLE.md. Any hits = automatic deduction. This is the ground truth for voice — not generic "sounds human" criteria.~/marketing/WRITING-STYLE.md: does it lead with a concrete moment (not a thesis)? Does it use specific numbers? Does it show the work rather than claim it? If none of these are present, it's not Eric's voice.Common failures:
What you're really asking: Where does the reader's attention drift? What's there for the writer, not the reader?
Score 5 (tight): Every sentence earns its place. You couldn't cut anything without losing meaning. Score 3 (some fat): 10-20% could go. Some repeated points, some throat-clearing, some sections that don't advance the argument. Score 1 (bloated): 30%+ is filler. Repeats the same point in different words. Long preambles. Obvious padding.
How to evaluate:
Provide a specific cut list:
CUT: [paragraph/sentence] — reason
CUT: [paragraph/sentence] — reason
TRIM: [paragraph/sentence] → [shorter version]
What you're really asking: Why does THIS person need to write THIS article? What's the angle only they can bring?
Score 5: Crystal clear point of view. You know exactly what the author believes and why. There's a driving emotion — frustration with the status quo, excitement about a discovery, pride in something they built. Score 3: Has an angle but doesn't commit to it fully. Hedges too much. Or the emotion is there but buried under too much analysis. Score 1: No angle. This could be a Wikipedia summary. No emotion. No stakes. No reason this person specifically needed to write this.
Identify:
Common failures:
What you're really asking: Who is this for? Can you picture them? What did they Google or scroll past that led them here?
Score 5: You can describe the reader in one sentence. Their problem is clear. The article directly addresses their situation. Score 3: General audience. "People interested in investing." Not wrong, but not specific enough to drive strong resonance. Score 1: Nobody in particular. Or the article serves two different audiences and does neither well.
Identify:
When invoked by the editor-in-chief skill, use classification labels instead of numeric scores. This maps the 6 questions to 6 dimensions with STRONG / NEEDS WORK / WEAK labels:
| Question | Dimension | STRONG | NEEDS WORK | WEAK |
|---|---|---|---|---|
| Q1 (Shareable?) | Shareability | 2+ screenshot moments, reader would forward | Hook/insight exists but buried or undersold | Nothing worth sharing |
| Q2 (Informative/fun?) | Substance | Every claim backed by evidence | Mix of evidence and assertions | Vague claims, no proof |
| Q3 (Human voice?) | Voice | Sounds like the author, irregular rhythm, opinionated | Mostly human, some stiff patches or AI tells | Robot cadence, banned patterns present |
| Q4 (Cuttable fat?) | Leanness | Every sentence earns its place | 10-20% filler | 30%+ chaff |
| Q5 (Unique angle?) | Emotion | Driving emotion clear and felt throughout | Emotion exists but buried or inconsistent | Flat, no emotional throughline |
| Q6 (Target reader?) | Reader Fit | Clear reader profile, article directly serves them | General audience, not specific enough | Nobody in particular |
Output format for classification mode:
Shareability: [STRONG|NEEDS WORK|WEAK] — [1-2 sentence explanation with specific examples]
Substance: [STRONG|NEEDS WORK|WEAK] — [explanation]
Voice: [STRONG|NEEDS WORK|WEAK] — [explanation]
Leanness: [STRONG|NEEDS WORK|WEAK] — [explanation]
Emotion: [STRONG|NEEDS WORK|WEAK] — [explanation]
Reader Fit: [STRONG|NEEDS WORK|WEAK] — [explanation]
When used standalone (not via editor-in-chief), use the original numeric scoring:
| Score | Meaning |
|---|---|
| 28-30 | Ship it. This is great. |
| 22-27 | Good but needs polish. Address the weakest scores. |
| 16-21 | Needs significant rework. Focus on scores below 3. |
| Below 16 | Start over or fundamentally rethink the angle. |
# Content Evaluation: [title or description]
## Scores
| Question | Score | One-line verdict |
|----------|-------|-----------------|
| Shareable? | X/5 | ... |
| Informative/fun? | X/5 | ... |
| Human voice? | X/5 | ... |
| Cuttable fat? | X/5 | ... |
| Unique angle? | X/5 | ... |
| Target reader? | X/5 | ... |
| **Total** | **XX/30** | |
## Detailed Feedback
### Shareable? (X/5)
[explanation + specific examples from the content]
### Informative/fun? (X/5)
[explanation + specific examples]
### Human voice? (X/5)
[explanation + flagged slop patterns]
### Cuttable fat? (X/5)
[specific cut list]
### Unique angle? (X/5)
[thesis + driving emotion + "only I" factor]
### Target reader? (X/5)
[reader profile + trigger + search query + awareness level]
## Top 3 Changes That Would Improve This Most
1. ...
2. ...
3. ...
Before scoring, check these. Any hit is a red flag that should pull the relevant score down:
Use this when evaluating TikTok slideshows, YouTube Shorts scripts, or any short-form video content. Run before the 6 questions above.
A solo dev with zero audience, zero budget got 18M views in 28 days using these principles. Each is a binary pass/fail.
Hook: Is it about the viewer, not the product?
Angle: Is it broad enough for people outside your niche?
Curiosity gap: Is there something they need to see by the end? The hook must create a question the viewer needs answered. Tease a result, reveal, or test early — hold the answer until the final slide. If there's no unresolved tension, there's nothing keeping them to the end.
15-second silent demo test: Does it work with sound off? Watch the first 15 seconds muted. If a stranger couldn't understand what the content is about or feel the hook's tension without audio, the visual hook is broken — and no amount of copy fixes a broken visual hook. Captions help but don't substitute for a self-explanatory visual sequence.
Line discipline: Does every line earn the next? Read the script line by line. For each line, ask: does this build curiosity, escalate tension, or deliver on a promise? If it explains, teaches, or provides context without earning it — cut it. Information is the reward at the end, not the scaffolding throughout.
Before delivering any evaluation or suggested rewrites, scan your own output for kill phrases. You are not exempt from the standards you're applying. Specifically:
If your feedback contains any of these patterns, rewrite the feedback before sending.
Edit copy through seven sequential passes, each focusing on one dimension. After each sweep, loop back to check previous sweeps aren't compromised (the cascading re-check pattern).
Cascading re-check: After completing Sweep N, re-check Sweeps 1 through N-1 to ensure your edits didn't introduce new issues in already-cleared dimensions. This compounds: after Sweep 7, do one final pass through all six prior sweeps.
Focus: Can the reader understand what you're saying?
What to check:
Common issues:
Process:
Focus: Is the copy consistent in how it sounds?
What to check:
Common issues:
Process:
Focus: Does every claim answer "why should I care?"
What to check:
The So What test: For every statement, ask "Okay, so what?" If the copy doesn't answer with a deeper benefit, it needs work.
Process:
Focus: Is every claim supported with evidence?
What to check:
Types of proof: Testimonials with names, case study references, statistics and data, third-party validation, guarantees and risk reversals, customer logos, review scores.
Common issues:
Process:
Focus: Is the copy concrete enough to be compelling?
What to check:
Specificity upgrades:
| Vague | Specific |
|---|---|
| Save time | Save 4 hours every week |
| Many customers | 2,847 teams |
| Fast results | Results in 14 days |
| Improve your workflow | Cut your reporting time in half |
| Great support | Response within 2 hours |
Process:
Focus: Does the copy make the reader feel something?
What to check:
Techniques: Paint the "before" state vividly, use sensory language, tell micro-stories, reference shared experiences, ask questions that prompt reflection.
Process:
Focus: Have we removed every barrier to action?
What to check:
Risk reducers: Money-back guarantees, free trials, "No credit card required," "Cancel anytime," social proof near CTA, clear expectations of what happens next, privacy assurances.
Process:
Use these for faster reviews when a full seven-sweep process isn't needed.
Cut these words:
Replace these:
| Weak | Strong |
|---|---|
| Utilize | Use |
| Implement | Set up |
| Leverage | Use |
| Facilitate | Help |
| Innovative | New |
| Robust | Strong |
| Seamless | Smooth |
| Cutting-edge | New/Modern |
Replace complex or pompous words with simpler ones. Source: Plain English Campaign, plainlanguage.gov.
| Complex | Plain |
|---|---|
| absence of | no, none |
| accomplish | do, finish |
| additional | extra, more |
| advise | tell, say |
| allocate | give, share |
| anticipate | expect |
| approximately | about |
| ascertain | find out |
| assistance | help |
| at the present time | now |
| cease | stop, end |
| commence | start, begin |
| communicate | tell, talk |
| consequently | so |
| currently | now |
| demonstrate | show, prove |
| determine | decide |
| discontinue | stop |
| disseminate | spread |
| due to the fact that | because |
| endeavour | try |
| establish | set up, show |
| expedite | speed up |
| facilitate | help |
| for the purpose of | to, for |
| furthermore | also, and |
| implement | carry out, do |
| in accordance with | under |
| in conjunction with | with |
| in order to | to |
| in the event of | if |
| indicate | show, suggest |
| initiate | start, begin |
| moreover | also, and |
| notify | tell |
| obtain | get |
| on behalf of | for |
| owing to | because |
| permit | let, allow |
| prior to | before |
| procure | get |
| provide | give |
| purchase | buy |
| regarding | about |
| reimburse | repay |
| require | need |
| retain | keep |
| subsequently | later |
| sufficient | enough |
| terminate | end, stop |
| utilise | use |
Phrases to remove entirely: "a total of," "absolutely," "actually," "at the end of the day," "at this moment in time," "basically," "I am of the opinion that" (use "I think"), "in the final analysis," "it should be understood," "last but not least," "obviously," "of course," "quite," "really," "the fact of the matter is," "to all intents and purposes," "very."
Watch for:
| Problem | Symptom | Fix |
|---|---|---|
| Wall of Features | List of what the product does without why it matters | Add "which means..." after each feature to bridge to benefits |
| Corporate Speak | "Leverage synergies to optimize outcomes" | Ask "How would a human say this?" and use those words |
| Weak Opening | Starting with company history or vague statements | Lead with the reader's problem or desired outcome |
| Buried CTA | The ask comes after too much buildup, or isn't clear | Make the CTA obvious, early, and repeated |
| No Proof | "Customers love us" with no evidence | Add specific testimonials, numbers, or case references |
| Generic Claims | "We help businesses grow" | Specify who, how, and by how much |
| Mixed Audiences | Copy tries to speak to everyone, resonates with no one | Pick one audience and write directly to them |
| Feature Overload | Listing every capability, overwhelming the reader | Focus on 3-5 key benefits that matter most to the audience |
Load on-demand:
references/pillar-evaluation.mdfor monthly pillar-level evaluation: scorecard, verdicts (SCALE/KEEP/ELEVATE/ROTATE OUT), recap format, and pillar-vs-post-level comparison.
~/marketing/WRITING-STYLE.md — ground truth for voice evaluation. Read the Kill Phrases list and Voice Fingerprint section before scoring Question 3. Generic "sounds human" is not the bar; Eric's specific fingerprint is.