| name | extract-creator-voice |
| description | Build a complete voice profile for any creator by pulling and analyzing their top-performing transcripts. Extracts all six layers of authentic voice — lexicon, syntax, rhythm & cadence, emotional register, storytelling architecture, and listener relationship — so AI-generated content is practically indistinguishable from the creator's real voice. Use before writing any scripts, captions, emails, or VSLs for a specific person. Triggers on "build a voice profile," "extract [name]'s voice," "capture how [name] speaks," "voice analysis," "speaking style," or any request to write content that sounds like a specific real person. |
Produce a complete, layered Voice Profile document for a specific creator — grounded entirely in their real transcript data — that enables AI to generate content practically indistinguishable from their authentic voice.
"Practically indistinguishable" means:
- An audience member who follows this creator would not pause and think "that doesn't sound like them"
- The content contains their actual fingerprints: rhythm, phrases, emotional texture, storytelling patterns
- It does not read as AI-generated, generically polished, or borrowed from anyone else's brand
Surface-level voice capture (vocabulary lists, topic summaries) is insufficient. This skill extracts all six layers of voice because indistinguishability requires all six. The layer most commonly missed — and most responsible for AI content that feels "off" — is rhythm and cadence.
Most voice profiles fail at the same place: they capture WHAT the creator says but not HOW they say it.
Word choice alone doesn't create authenticity. What creates authenticity is the rhythm underneath the words — the specific beat pattern of short declaratives and longer elaborations, the places they pause, the way they repeat for emphasis, the emotional gear shifts between teaching and vulnerability and humor. These are felt by the listener before they're consciously registered.
The six layers, from surface to deep:
| Layer | What It Is | Why It Matters |
|---|
| 1. Lexicon | Actual words, phrases, terms they use | Wrong words make everything sound off |
| 2. Syntax | How sentences are built and structured | Over-polished AI syntax doesn't match real cadence |
| 3. Rhythm & Cadence | The music of how they speak | The most missed — rhythm is felt, not just heard |
| 4. Emotional Register | When they're vulnerable, excited, funny, urgent | Flat content feels fake regardless of word choice |
| 5. Storytelling Architecture | How they structure and deliver stories | Stories are the primary trust vehicle |
| 6. Listener Relationship | How they speak TO their audience | Intimacy gaps are what make AI feel robotic |
Do not skip layers. Do not treat vocabulary as a proxy for voice. The profile is not complete until all six are documented with evidence from actual transcripts.
Before starting, collect the following. Ask all at once if anything is missing.
Required:
- Creator name (full name as they're known publicly)
- YouTube channel handle or URL (primary content source)
- Client context if applicable (which vault client folder this profile belongs to)
Optional but useful:
- Any other platforms where they publish long-form content (podcast feed, Instagram, website)
- Known top-performing videos (if the user already knows which ones resonated)
- Co-hosts or guests who regularly appear with them (needed for diarization)
- Format context: solo content? interview? co-hosted podcast? family dynamic?
If the user provides a YouTube URL directly — skip intake and start immediately with transcript acquisition for that video, then expand from there.
<source_selection>
Step 1: Find the right source material
The voice is most authentic in top-performing content. High view/engagement = content that resonated most deeply = the creator at their most effective and natural.
Priority sources (pull in this order):
- Top-performing long-form content (20+ minutes, highest view counts) — fullest voice sample, broadest emotional range, highest density of signature phrases and stories
- Solo camera content (creator speaks alone to lens) — reveals the teaching/monologue voice without conversational crutches; critical for generating solo scripts, VSLs, emails
- Conversational content with a trusted recurring partner (co-host, family member, close collaborator) — reveals relational voice, humor, and vulnerability; often MORE authentic than polished solo content
- Older high-performing content (2+ years old) — surfaces phrases and patterns that have persisted over time, indicating core voice vs. transient style
- Written content (newsletters, social posts) — shows how their spoken voice translates to written form; often differs in ways that matter for copy
Minimum viable set: 3 long-form transcripts + 1 solo piece + at least 2 distinct topics covered.
How to find top content using TranscriptAPI:
# Search for the channel first
mcp__claude_ai_TranscriptAPI__search_youtube
query: "[creator name] [brand/topic]"
search_type: "channel"
# List channel videos to find top performers
mcp__claude_ai_TranscriptAPI__list_channel_videos
channel: "@[handle]"
# Or search within the channel for specific content types
mcp__claude_ai_TranscriptAPI__search_channel_videos
channel: "@[handle]"
query: "[topic or format]"
Sort by view count. Select the 3–5 highest-performing videos that feature the creator as a primary speaker. If there are multiple content formats (podcast episodes, solo explainers, interviews), select across formats.
Note: Some videos show hasCaptions: false in search results but still have auto-generated captions accessible. Always attempt the transcript pull regardless — don't skip based on the search result flag.
</source_selection>
<transcript_acquisition>
Step 2: Pull transcripts in parallel
Pull all selected transcripts simultaneously in a single message with multiple tool calls:
mcp__claude_ai_TranscriptAPI__get_youtube_transcript
video_url: [video_id_or_url]
format: "text"
include_timestamp: true
Timestamp data is important — it reveals speaking pace, pause patterns, and how long the creator spends on each point.
If a video has no accessible captions:
- Skip it unless it's critical
- Note it in the profile's Gap Notes section
- Alternative: request that the user download the audio and run it through a diarization service
Volume check: If any transcript returns "Result too long, truncated" — immediately re-request with shorter segments (use timestamp offsets). Never analyze truncated data.
</transcript_acquisition>
<speaker_diarization>
Step 3: Identify who is speaking
YouTube auto-transcripts mark speaker changes (usually with >>) but do not label names. Resolve this before analysis — analyzing the wrong speaker's lines will contaminate the profile.
Method A: Contextual name-drop identification (use first)
Scan transcripts for speakers calling each other by name ("Mike, what do you think?" → the next turn is Michael). Build a mapping table:
SPEAKER A [NAME]: Identifying signals
SPEAKER B [NAME]: Identifying signals
Signals to look for:
- Opening intro ("My name is..." or "We're your hosts, [Name] and [Name]")
- Direct name address ("Great point, [Name].")
- Role-based cues ("As the dad in this conversation...", "From the student's perspective...")
- Topic ownership (who introduces the topic vs. who responds?)
- Recurring phrases that belong to one voice
Method B: Content-position identification
When names aren't used:
- The host/main speaker usually opens, frames topics, and closes
- The guest/collaborator usually reacts, asks questions, or challenges
- Subject matter expertise patterns (who speaks with more authority on which subtopics?)
Method C: Solo content
If the creator speaks alone to camera, no diarization is needed. Note this explicitly — solo content is the cleanest voice signal.
Document the diarization table in the profile. Future users of the profile need to know which attribution method was used.
</speaker_diarization>
<layer_analysis>
Step 4: Run the six-layer extraction
Work through every layer. Do not skip layers or merge them. Evidence for each layer must come from actual transcript lines, not inference.
LAYER 1: LEXICON
Extract from transcripts:
Signature phrases (appear across 2+ transcripts — these are fingerprints):
- Transitional connectors: how they bridge ideas ("And so," "But here's the thing," "Once again")
- Teaching signals: how they enter instruction mode ("Let me share," "Here's the thing," "Think about that")
- Affirmations: how they validate ("I love that," "That's incredible," "Right?")
- Story anchors: how they open a story ("I remember when," "Years ago," "Let me take you back")
- Self-deprecation phrases: how they poke fun at themselves
- Closing signals: how they end a thought or episode
Vocabulary inventory:
- Business/domain terms they use naturally (with precision, not incorrectly)
- Character/values vocabulary (legacy, mission, stewardship, etc.)
- What they call their audience, their community, their method
The DON'T list (equally important):
What vocabulary would immediately break the voice? Look for what's conspicuously absent across all transcripts. Common AI-defaults that real people often don't use: "hustle," "grind," "crush it," "level up," "game changer," "unlock," "leverage," corporate buzzwords, filler superlatives.
Delivery: Signature phrases organized by function (20+ items). DO list. DON'T list.
LAYER 2: SYNTAX
Extract from transcripts:
Sentence length pattern: Read 10 consecutive sentences. Are they short bursts (under 10 words), medium (10–20 words), or long runs (20+)? What's the dominant mode? What triggers the secondary mode?
The thesis-evidence-restatement pattern: Many speakers state a claim → build supporting detail in staccato sentences → restate the claim with variation. Document if this exists and what it looks like.
Question types and frequency:
- Rhetorical questions to themselves (vulnerability, self-examination)
- Questions directed at a co-host (collaborative exploration)
- Questions directed at the audience (engagement, application)
- How often? What topics trigger questions?
Sentence fragments: Do they use fragments for emphasis? ("Not always." "Years ago." "Right.") These are rhythm devices. Replicating them in generated content closes the uncanny valley gap.
Qualifier ratio: Do they hedge often ("I think," "maybe," "sort of") or speak declaratively? Getting the ratio wrong makes content feel either too timid or too authoritative.
Opening phrase pattern: How does this person start a new thought? What are the 3–5 most common sentence-openers across transcripts?
Delivery: Document the dominant sentence structure with 2–3 verbatim examples from transcripts.
LAYER 3: RHYTHM & CADENCE (most important layer)
This is what makes content feel like them rather than just sound-alike. Rhythm is felt before it's consciously heard.
The beat pattern:
Select a 4–6 sentence passage from their transcript. Count syllables per sentence. What's the pattern?
Example:
"That to me is not an investment." (10 syl) ← thesis
"They're not creating long-term wealth." (9 syl) ← evidence
"They're looking to buy a home." (7 syl) ← evidence
"That to me is speculation." (9 syl) ← restatement
Pattern: declining → rebound (the "rocking" pattern)
Name the pattern. Document it with syllable counts. Find 2–3 more examples of the same pattern. This is the rhythm template for generated content.
Secondary patterns: Most speakers have 2–3 rhythm patterns for different emotional contexts. Find the one they use for vulnerability, for humor, for urgency. Document each.
The triple: Does this person naturally group things in threes? ("values, principles, stories" / "provide, protect, teach") The triple is the most human rhythmic device and the one AI most consistently underuses. Document frequency and typical contexts.
The callback: Do they return to an earlier phrase or image to close a loop? ("Remember when I said X? Well, now you see why...") This creates the feeling of a complete, satisfying thought. Document if it exists.
Repetition for emphasis: Where do they repeat the same phrase back-to-back or with slight variation? ("That to me is not an investment. That to me is speculation.") This is deliberate and rhythmically important.
Pause indicators: Look for: fragments, mid-sentence check-ins ("right?" "you know?" "isn't it?"), self-interruptions ("I mean..."), double repetition of a word ("saving saving saving"). These indicate natural breath points. In written content, replicate as paragraph breaks or em-dashes.
Pacing by emotion: Map how rhythm changes by emotional state:
- Teaching/calm: what's the typical sentence length?
- Excited: do sentences get longer? More "and" chains?
- Vulnerable: do sentences get shorter? More fragments?
- Humorous: where does the punchline land in the sentence structure?
Calibration sample: Include a 5–6 sentence verbatim passage that captures their dominant rhythm. This gets read aloud before generating content to "tune" the writer's ear.
Delivery: Named beat pattern + syllable-count example + calibration passage.
LAYER 4: EMOTIONAL REGISTER
Map the distinct emotional gears this creator uses. Most speakers have 4–6 distinct registers. Name them plainly.
For each register, document:
- Name (e.g., "Vulnerable-Reflective," "Direct-Teaching," "Warm-Proud," "Playful")
- Trigger topics (what subjects cause this register to appear?)
- Linguistic markers (what changes in their language when they shift into this register? Sentence length? Hedging? Pace? Vocabulary?)
- Typical entry phrase (what do they say when transitioning into this register?)
The vulnerability pattern (document specifically if it exists):
- What triggers it (topics, admissions)?
- What language appears (hedged, fragment-heavy, past tense, self-questioning)?
- How they recover (usually toward a lesson or gratitude)?
The humor signature: Every person's humor has a fingerprint. Is it self-deprecating? Callback-based? At whose expense? Warm or edgy? Document so generated content includes humor that actually sounds like theirs — not generic AI humor.
Delivery: Table of registers with triggers and linguistic markers.
LAYER 5: STORYTELLING ARCHITECTURE
Story inventory: List every specific story referenced in the transcripts by name. Mark which ones appear across multiple transcripts — these are signature parables central to the brand identity.
For each signature story, note:
- The core narrative arc (3 sentences max)
- The teaching point it's attached to
- The entry phrase they use to begin it
- Specific details that make it memorable (exact dollar amount, year, person's name)
Story structure template: What is their default narrative arc?
- How do they open a story? (verbatim entry phrase examples)
- Where does conflict appear?
- Where does the turning point fall?
- How do they deliver the lesson?
- Do they turn the lesson back to the listener?
Detail density: Do they name the specific dollar amount? The exact year? The person's exact words? High-detail storytellers sound wrong when generated content uses vague approximations. Document the standard.
The lesson delivery: How do they get from story to teaching point? Abruptly? With a bridge phrase? By turning the question back to the listener? This transition is often where AI content falls apart — it either rushes the lesson or over-explains it.
Co-protagonist pattern: Many creators make another person (child, spouse, mentor, student) the co-hero of their stories. This deflects ego while keeping content personal. Document recurring co-protagonists and their role in the story structure.
Delivery: Story inventory with names + structure template + detail density standard.
LAYER 6: LISTENER RELATIONSHIP
The intimacy model: How does this creator treat the listener? As a student? A peer? A younger version of themselves? A friend? This determines the entire register of address.
Direct address frequency and patterns:
- How often do they directly address the listener ("you," "if you're listening," "I want you to...")?
- What phrases do they use to bring the listener into the content?
- Do they address the listener by a specific archetype ("If you're a dad hearing this...")?
The authority style: Do they challenge the listener ("I challenge you to...") or invite them ("What if you tried...")? High-authority vs. collaborative peer — which is this person's natural mode?
Parasocial intimacy: Do they share personal details that create closeness? (Name their kids. Share specific dollar amounts from their own finances. Reference where they live.) Document the density and type — generated content must replicate this at the same level.
The close pattern: Document exactly how they end episodes, segments, or posts. Closings are often formulaic and immediately noticeable when wrong. Include the verbatim formula if one exists.
Delivery: Intimacy model statement + direct address phrases + closing formula.
</layer_analysis>
<output_format>
Step 5: Write the Voice Profile Document
Write to: vault/business/clients/[client-slug]/profiles/[speaker-firstname]-voice-profile.md
If this is not a vault client (standalone use), write to: [working-directory]/[speaker-firstname]-voice-profile.md
Document structure:
---
type: reference
source: [transcript-source]
domain: marketing
status: active
use-when: "writing any content in [Name]'s voice — [list formats: scripts, captions, emails, VSLs]"
created: [YYYY-MM-DD]
tags: [client-slug, speaker-name, voice-profile, speaking-style, content-generation]
---
# [Full Name] — Voice Profile
**Built from:** [list sources with view counts]
---
## SPEAKER DIARIZATION NOTES
[Diarization method used. Mapping table of speakers and identification signals.]
---
## 1. VOCAL IDENTITY
[One paragraph capturing the essence of this voice in plain language.
Who does it feel like you're talking to? What's their relationship to the listener?
Write this last — it synthesizes everything above.]
---
## 2. SIGNATURE PHRASES & VERBAL TICS
[Organized by function: connectors, affirmations, teaching signals, story anchors, humor, self-deprecation, closing signals. 20+ items. DO list. DON'T list.]
---
## 3. RHYTHM & CADENCE
[Named beat pattern with syllable-count example. Secondary patterns.
Triple usage. Callback pattern. Repetition for emphasis. Pause indicators.
Pacing-by-emotion table. Calibration sample passage.]
---
## 4. SENTENCE STRUCTURE & SYNTAX
[Dominant sentence length. Thesis-evidence-restatement pattern if it exists.
Question types and frequency. Fragment usage. Qualifier ratio. Common sentence openers.]
---
## 5. STORYTELLING ARCHITECTURE
[Story inventory (named, with arcs). Structure template. Story entry phrases.
Detail density standard. Lesson delivery method. Co-protagonist pattern.]
---
## 6. EMOTIONAL REGISTER MAP
[Table of 4–6 named registers with trigger topics and linguistic markers.
Vulnerability pattern documented. Humor signature documented.]
---
## 7. VOCABULARY & REGISTER
[DO list: domain terms, character/values vocabulary, community language.
DON'T list: absent vocabulary that would immediately break the voice.]
---
## 8. LISTENER RELATIONSHIP
[Intimacy model. Direct address patterns with example phrases.
Authority style. Parasocial intimacy density. Closing formula verbatim.]
---
## 9. BOOK & MENTOR REFERENCES
[Every author, book, quote source, or mentor referenced across transcripts.
These signal intellectual identity — include them in content at appropriate density.]
---
## 10. QUICK REFERENCE — DO/DON'T
[Bullet list for fast-checking generated content before delivery.
15–20 DOs. 10–15 DON'Ts.]
---
## 11. AUTHENTIC QUOTE BANK
[8–12 verbatim quotes from transcripts organized by topic/emotion.
These are anchor texts — read them before generating content to tune the voice.]
---
[What this profile does NOT yet capture. Recommended next sources.
Format-specific voice differences if detected.]
</output_format>
**Step 6: The 5-Point Voice Check**
Before delivering the profile, verify it passes all five:
1. Phrase authenticity
Are the signature phrases actually from transcripts — not inferred or paraphrased? Every item on the DO list should trace to at least one verbatim source.
2. Rhythm documented with evidence
Does the rhythm section include actual syllable counts? A named pattern? A calibration passage? If it's described abstractly without evidence, it will fail in production.
3. Emotional range represented
Are there at least 4 distinct emotional registers documented? If the profile describes only one emotional mode, it's incomplete — content generated from it will be tonally flat.
4. The DON'T list is substantive
Does the DON'T list have 10+ items? If not, it's underdeveloped. The DON'T list prevents AI drift back toward generic output — it's not optional.
5. The quote bank is verbatim
Every quote in the bank must be pulled directly from transcript text, word-for-word. No paraphrasing. No cleaning up. Real speech includes hesitations, self-corrections, and informal constructions — these are voice signals, not errors.
The side-by-side test (final check):
Place 2–3 sentences from an actual transcript next to 2–3 sections from the profile's documented patterns. Ask: do these feel like they were produced by the same person? If not, identify which layer is misrepresenting the voice and correct it.
<file_management>
After writing the profile, create the symlink in the business skills directory for Obsidian access:
ln -sf ~/.claude/skills/extract-creator-voice/SKILL.md \
~/vault/business/skills/extract-creator-voice.md
If the profile is for a vault client, confirm it lands in the correct client folder:
vault/business/clients/[client-slug]/profiles/[speaker]-voice-profile.md
If a voice profile already exists for this creator, read it first before running the extraction — append new findings to existing sections rather than overwriting validated data.
</file_management>
<quality_standards>
Done means:
- Minimum 3 long-form transcripts analyzed (5 preferred)
- Speaker diarization resolved and documented
- All 6 layers covered with transcript evidence (not inference)
- Rhythm documented with syllable-count examples and a calibration passage
- Emotional register map has 4+ distinct registers with triggers
- Story inventory names 5+ signature stories
- DO list has 20+ items; DON'T list has 10+ items
- Quote bank has 8+ verbatim extracts
- File written to the correct vault location
What failure looks like:
- Vocabulary list only, no rhythm or syntax analysis → will fail the indistinguishability test
- Emotional register described as one flat mode → generated content will feel tonally robotic
- Quote bank contains paraphrased speech → destroys the calibration function
- DON'T list is empty or has fewer than 5 items → AI will drift toward generic corporate language
- Profile built primarily from low-performing content → captures an inauthentic version of the voice (creators often sound off on topics they're not passionate about or formats they're uncomfortable in)
The indistinguishability standard:
If an audience member who follows this creator reads generated content and thinks "that sounds like them" — the profile works. If they think "that's close but something's off" — one of the six layers is underdocumented. Use the 5-point check to identify which one.
</quality_standards>