Use when a PERSONAL brand needs AI avatar video at scale — three tool tiers, four workflows for single avatar, translation, batch, and hybrid, reference image intake, face, style, logo, and palette replacement, voice clone pairing, anti-detection, and a QA score, with disclosure-law variants for US FTC, EU AI Act, SEA, and LATAM, covering HeyGen and Synthesia. Trigger on 'AI avatar', 'HeyGen video', 'Synthesia', 'talking head AI video', 'translate my videos with AI', 'I cannot be on camera every day'. Not for — the words the avatar says, see `04-script-video-global`; audio-only voice clone and podcast, see `25-voice-clone-podcast-global`; a company product video edit, see `44-video-editor-brief-global`.
Use when a PERSONAL brand needs AI avatar video at scale — three tool tiers, four workflows for single avatar, translation, batch, and hybrid, reference image intake, face, style, logo, and palette replacement, voice clone pairing, anti-detection, and a QA score, with disclosure-law variants for US FTC, EU AI Act, SEA, and LATAM, covering HeyGen and Synthesia. Trigger on 'AI avatar', 'HeyGen video', 'Synthesia', 'talking head AI video', 'translate my videos with AI', 'I cannot be on camera every day'. Not for — the words the avatar says, see `04-script-video-global`; audio-only voice clone and podcast, see `25-voice-clone-podcast-global`; a company product video edit, see `44-video-editor-brief-global`.
metadata
{"version":"1.1.1","category":"content"}
license
MIT
triggers
["AI avatar","HeyGen","Synthesia","avatar AI video","talking head AI","AI video translate","batch AI video","avatar reference image","AI avatar prompt","replace avatar face"]
AI Avatar Production (Global) — Pipeline 3-Tier, 4 Workflows, QA Score 100
Flagship skill of the AI Content cluster. Covers the full pipeline from zero to publish, voice clone, anti-detection, and region-specific disclosure law.
For newbies
What is an AI Avatar?
An AI Avatar is a video that shows your face (or a stand-in) but uses AI-generated voice and motion. You provide one photo or a short selfie video; the AI produces a final video with natural-looking speech, gestures, and expressions. No filming crew, no studio, no actor required.
Budget tier? Free ($0) / Pro ($30-100/mo) / Enterprise ($200+/mo)?
Videos per month target? 1-5 / 10-30 / 30+?
Based on the 4 answers, auto-select Tier + Workflow.
If the user has already uploaded reference images, do not ask a long intake form first; classify the images, create the setup/prompt, then ask only for missing assets.
Tier decision — Tools and pricing
Tier
Suggested tool
Price/month
Quality
Limit
Fits
Free
Captions Free, HeyGen Trial, D-ID Trial
$0
6/10 — watermark, limited duration
1-5 videos, max 60s/video
Personal test, new freelancers
Pro
HeyGen Creator ($29), Synthesia Starter ($29), ElevenLabs Pro ($22)
$30-100
8/10 — no watermark, HD
10-30 videos, max 5 min/video
SME, small agency, content creator
Enterprise
HeyGen Business ($89+), Synthesia Enterprise (custom)
$200-500+
9.5/10 — custom avatar, API, priority render
30+ videos, unlimited
Large agency, large brand, e-learning
Quick recommendations:
Just starting: HeyGen Trial (1 video free, full experience)
Serious but budget-limited: Captions Pro ($10/mo) for lipsync + ElevenLabs Starter ($5) for voice
Scale fast: HeyGen Creator + ElevenLabs Pro = best price/quality combo
Repeat 3 weeks = 30 videos. Or scale Days 1-2 to 15 scripts/week.
Cost estimate batch 30 videos/month
Tier
Tool combo
Monthly cost
Per-video cost
Free
HeyGen Trial + Captions Free
$0 (limited 3-5 videos)
$0 (watermark)
Pro
HeyGen Creator + ElevenLabs Pro
~$51
~$1.70
Enterprise
HeyGen Business + ElevenLabs Scale
~$189
~$6.30
Batch optimization tips
Templated scripts: 3-5 frameworks, swap the core content
Voice consistency: One voice clone for the entire series
Off-peak rendering: Queue overnight to skip the queue
QA checklist: Print the QA Score, check videos like an assembly line
Workflow 4: Hybrid Real + AI
Real face for trust + AI body for speed.
Use cases
Real face intro 5s + AI body 55s (save filming time)
AI video weekdays + Real video weekly (balance quality/effort)
Real talking head + AI B-roll (studio-grade output)
Assembly + tools
Film real intro 5-10s (eye contact, natural greeting); use Captions for lipsync fixes
Create AI for the rest with same outfit/background (HeyGen / Synthesia)
Edit in CapCut / Premiere (precise cuts, smooth transitions)
Color match AI to real footage (LUT or DaVinci Resolve free)
Trust gain: Real face up front -> 20-35% more engagement than full-AI.
Voice Clone Protocol
Voice sample requirements
Criterion
Requirement
Duration
3-5 minutes
Quality
WAV/FLAC, 44.1kHz+, mono, quiet room
Script content
Phonetically varied passages (all vowels, hard consonants)
Emotion
Read normal, natural, not acted
Tool comparison
Tool
Price
Quality
Notes
ElevenLabs
From $5/mo
9/10
Best overall, 30+ languages
HeyGen Voice
Included Creator+
6/10
Convenient if using HeyGen
Resemble AI
From $99/mo
7/10
Strong API
PlayHT
From $39/mo
7/10
Good for narration
Consent form template
MANDATORY before cloning anyone's voice.
VOICE USAGE CONSENT
I, [FULL NAME], consent to [COMPANY] using my voice for: [SPECIFIC PURPOSE].
Term: [X months / Until revoked]
Date: [YYYY-MM-DD]
Signature: _______________
Reference: See references/voice-clone-prompts-global.md
Avatar Setup Checklist
Before recording / uploading photo or video for an AI avatar:
Lighting: Natural light or softbox; no harsh shadows on the face
Background: Solid (white / gray) or real environment (office, store)
Wardrobe: On-brand; avoid small busy patterns (AI moire)
Framing: Chest up; eyes on the upper-third line
Eye contact: Look directly at the lens (not the screen)
Gestures: Natural; hands can rest or do light gestures
File format: MP4 (H.264) for video, PNG / JPG for photo
Backup: Keep originals on cloud (Google Drive / OneDrive) before uploading to the tool
Reference Image -> Avatar Prompt Director
Use this when the user drops one or more reference images and wants to create an avatar, replace a face, adapt brand colors, add a logo, or create the prompt before uploading assets into a tool.
Classify Input Images
Image type
Role
Requirement
Style ref
Mood, lighting, background, outfit, camera angle
Do not use as identity unless requested
Face ref
Identity preservation / face replacement
1-3 clear face images, no filter, front + 3/4 angle
Selfie video
Better custom avatar / natural lipsync
30s-2 min, looking at camera, speaking naturally
Logo/palette
Personal/company brand adaptation
PNG/SVG logo + 2-4 hex colors
Product/location
Prop or avatar environment
Clear product label or location/background image
Multiple Images = Multiple Flows
## Avatar Flows
| Flow | Input image | Role | Suggested tool | Missing assets |
|------|-------------|------|----------------|----------------|
| A | style-01 | style/background | Design Master -> HeyGen | face ref, logo |
| B | face-01 | identity | HeyGen custom avatar | script, voice sample |
If every image is a different style direction, create a separate prompt for each flow.
If images support one avatar, group by role: style + face + logo + palette + product.
Ask for each next asset explicitly: face image, selfie video, logo, hex colors, script, voice sample.
LinkedIn: Practically no detection — best fit for AI avatars
YouTube: Strict for monetized videos — must disclose per YPP policy
CRITICAL: NEVER use AI avatars to impersonate real people without consent. This is illegal in most jurisdictions and grounds for permanent platform bans.
Ethics and Disclosure — Region selector
Disclosure laws differ dramatically by region. Pick the matching variant:
Region
Variant file
Key law
US / Canada
variants/01-us.md
FTC Endorsement Guides (16 CFR Part 255), 2023 update
EU / EEA / UK
variants/02-eu.md
EU AI Act Article 50 (always disclose) + UCPD + GDPR
Southeast Asia
variants/03-sea.md
Per-country: ASAS (SG), AKARI (ID), DTI (PH), MCMC (MY), TH
ALWAYS read the matching variant BEFORE publishing AI avatar content in that region. Penalties range from warning to multi-thousand-USD fines per influencer (US) and can stack under EU AI Act + GDPR.
Universal disclosure rule of thumb
When in doubt, disclose. Disclosure is rarely penalized; non-disclosure can be.
"This video uses AI Avatar technology for visuals and voice."
Placement: video description, first 3 seconds on-screen text, OR platform "AI-generated" tag (where available — Meta, TikTok, YouTube all now support this).
QA Score — 100 points
Scorecard
#
Criterion
Points
Description
1
Lipsync
/10
Mouth tracks speech within 0.2s
2
Voice match
/10
Voice sounds like the speaker (if clone) or natural (if TTS)
3
Visual quality
/10
Sharp image, no artifacts, no blur
4
Background
/10
Background suits context, no render glitches
5
Lighting
/10
Even light, no harsh shadows, matches background
6
Gesture
/10
Natural, no jitters, hand/head movement present
7
Script flow
/10
Hook -> Problem -> Solution -> CTA
8
Disclosure
/10
AI disclosure compliant with region (see variant)
9
Platform fit
/10
Correct aspect ratio, duration, format for platform