| name | youtube-transcript |
| description | Download YouTube video transcripts when user provides a YouTube URL or asks to download/get/fetch a transcript from YouTube. Also use when user wants to transcribe or get captions/subtitles from a YouTube video. |
| allowed-tools | Bash,Read,Write |
YouTube Transcript Downloader
This skill helps download transcripts (subtitles/captions) from YouTube videos using yt-dlp.
When to Use This Skill
Activate this skill when the user:
- Provides a YouTube URL and wants the transcript
- Asks to "download transcript from YouTube"
- Wants to "get captions" or "get subtitles" from a video
- Asks to "transcribe a YouTube video"
- Needs text content from a YouTube video
How It Works
Priority Order:
- Check if yt-dlp is installed - install if needed
- List available subtitles - see what's actually available
- Try manual subtitles first (
--write-sub) - highest quality
- Fallback to auto-generated (
--write-auto-sub) - usually available
- Last resort: Whisper transcription - if no subtitles exist (requires user confirmation)
- Confirm the download and show the user where the file is saved
- Optionally clean up the VTT format if the user wants plain text
Installation Check
IMPORTANT: Always check if yt-dlp is installed first:
which yt-dlp || command -v yt-dlp
If Not Installed
Attempt automatic installation based on the system:
macOS (Homebrew):
brew install yt-dlp
Linux (apt/Debian/Ubuntu):
sudo apt update && sudo apt install -y yt-dlp
Alternative (pip - works on all systems):
pip3 install yt-dlp
python3 -m pip install yt-dlp
If installation fails: Inform the user they need to install yt-dlp manually and provide them with installation instructions from https://github.com/yt-dlp/yt-dlp#installation
Check Available Subtitles
ALWAYS do this first before attempting to download:
yt-dlp --list-subs "YOUTUBE_URL"
This shows what subtitle types are available without downloading anything. Look for:
- Manual subtitles (better quality)
- Auto-generated subtitles (usually available)
- Available languages
Download Strategy
Option 1: Manual Subtitles (Preferred)
Try this first - highest quality, human-created:
yt-dlp --write-sub --skip-download --output "OUTPUT_NAME" "YOUTUBE_URL"
Option 2: Auto-Generated Subtitles (Fallback)
If manual subtitles aren't available:
yt-dlp --write-auto-sub --skip-download --output "OUTPUT_NAME" "YOUTUBE_URL"
Both commands create a .vtt file (WebVTT subtitle format).
Option 3: Whisper Transcription (Last Resort)
ONLY use this if both manual and auto-generated subtitles are unavailable.
Step 1: Show File Size and Ask for Confirmation
yt-dlp --print "%(filesize,filesize_approx)s" -f "bestaudio" "YOUTUBE_URL"
yt-dlp --print "%(duration)s %(title)s" "YOUTUBE_URL"
IMPORTANT: Display the file size to the user and ask: "No subtitles are available. I can download the audio (approximately X MB) and transcribe it using Whisper. Would you like to proceed?"
Wait for user confirmation before continuing.
Step 2: Check for Whisper Installation
command -v whisper
If not installed, ask user: "Whisper is not installed. Install it with pip install openai-whisper (requires ~1-3GB for models)? This is a one-time installation."
Wait for user confirmation before installing.
Install if approved:
pip3 install openai-whisper
Step 3: Download Audio Only
yt-dlp -x --audio-format mp3 --output "audio_%(id)s.%(ext)s" "YOUTUBE_URL"
Step 4: Transcribe with Whisper
whisper audio_VIDEO_ID.mp3 --model base --output_format vtt
whisper audio_VIDEO_ID.mp3 --model base --language en --output_format vtt
Model Options (stick to base for now):
tiny - fastest, least accurate (~1GB)
base - good balance (~1GB) ← USE THIS
small - better accuracy (~2GB)
medium - very good (~5GB)
large - best accuracy (~10GB)
Step 5: Cleanup
After transcription completes, ask user: "Transcription complete! Would you like me to delete the audio file to save space?"
If yes:
rm audio_VIDEO_ID.mp3
Getting Video Information
Extract Video Title (for filename)
yt-dlp --print "%(title)s" "YOUTUBE_URL"
Use this to create meaningful filenames based on the video title. Clean the title for filesystem compatibility:
- Replace
/ with -
- Replace special characters that might cause issues
- Consider using sanitized version:
$(yt-dlp --print "%(title)s" "URL" | tr '/' '-' | tr ':' '-')
Post-Processing
Convert to Plain Text (Recommended)
YouTube's auto-generated VTT files contain duplicate lines because captions are shown progressively with overlapping timestamps. Always deduplicate when converting to plain text while preserving the original speaking order.
python3 -c "
import sys, re
seen = set()
with open('transcript.en.vtt', 'r') as f:
for line in f:
line = line.strip()
if line and not line.startswith('WEBVTT') and not line.startswith('Kind:') and not line.startswith('Language:') and '-->' not in line:
clean = re.sub('<[^>]*>', '', line)
clean = clean.replace('&', '&').replace('>', '>').replace('<', '<')
if clean and clean not in seen:
print(clean)
seen.add(clean)
" > transcript.txt
Complete Post-Processing with Video Title
VIDEO_TITLE=$(yt-dlp --print "%(title)s" "YOUTUBE_URL" | tr '/' '_' | tr ':' '-' | tr '?' '' | tr '"' '')
VTT_FILE=$(ls *.vtt | head -n 1)
python3 -c "
import sys, re
seen = set()
with open('$VTT_FILE', 'r') as f:
for line in f:
line = line.strip()
if line and not line.startswith('WEBVTT') and not line.startswith('Kind:') and not line.startswith('Language:') and '-->' not in line:
clean = re.sub('<[^>]*>', '', line)
clean = clean.replace('&', '&').replace('>', '>').replace('<', '<')
if clean and clean not in seen:
print(clean)
seen.add(clean)
" > "${VIDEO_TITLE}.txt"
echo "✓ Saved to: ${VIDEO_TITLE}.txt"
rm "$VTT_FILE"
echo "✓ Cleaned up temporary VTT file"
Output Formats
- VTT format (
.vtt): Includes timestamps and formatting, good for video players
- Plain text (
.txt): Just the text content, good for reading or analysis
Tips
- The filename will be
{output_name}.{language_code}.vtt (e.g., transcript.en.vtt)
- Most YouTube videos have auto-generated English subtitles
- Some videos may have multiple language options
- If auto-subtitles aren't available, try
--write-sub instead for manual subtitles
Complete Workflow Example
VIDEO_URL="https://www.youtube.com/watch?v=dQw4w9WgXcQ"
VIDEO_TITLE=$(yt-dlp --print "%(title)s" "$VIDEO_URL" | tr '/' '_' | tr ':' '-' | tr '?' '' | tr '"' '')
OUTPUT_NAME="transcript_temp"
if ! command -v yt-dlp &> /dev/null; then
echo "yt-dlp not found, attempting to install..."
if command -v brew &> /dev/null; then
brew install yt-dlp
elif command -v apt &> /dev/null; then
sudo apt update && sudo apt install -y yt-dlp
else
pip3 install yt-dlp
fi
fi
echo "Checking available subtitles..."
yt-dlp --list-subs "$VIDEO_URL"
yt-dlp --write-sub --skip-download --output 2>/dev/null;
-lh .*
yt-dlp --write-auto-sub --skip-download --output 2>/dev/null;
-lh .*
FILE_SIZE=$(yt-dlp -- -f )
DURATION=$(yt-dlp -- )
TITLE=$(yt-dlp -- )
-r RESPONSE
[[ =~ ^[Yy]$ ]];
! -v whisper &> /dev/null;
-r INSTALL_RESPONSE
[[ =~ ^[Yy]$ ]];
pip3 install openai-whisper
1
yt-dlp -x --audio-format mp3 --output
AUDIO_FILE=$( audio_*.mp3 | -n 1)
whisper --model base --output_format vtt
-r CLEANUP_RESPONSE
[[ =~ ^[Yy]$ ]];
-lh *.vtt
0
VTT_FILE=$( *.vtt 2>/dev/null || *.vtt | -n 1)
[ -f ];
python3 -c >
Note: This complete workflow handles all scenarios with proper error checking and user prompts at each decision point.
Error Handling
Common Issues and Solutions:
1. yt-dlp not installed
- Attempt automatic installation based on system (Homebrew/apt/pip)
- If installation fails, provide manual installation link
- Verify installation before proceeding
2. No subtitles available
- List available subtitles first to confirm
- Try both
--write-sub and --write-auto-sub
- If both fail, offer Whisper transcription option
- Show file size and ask for user confirmation before downloading audio
3. Invalid or private video
- Check if URL is correct format:
https://www.youtube.com/watch?v=VIDEO_ID
- Some videos may be private, age-restricted, or geo-blocked
- Inform user of the specific error from yt-dlp
4. Whisper installation fails
- May require system dependencies (ffmpeg, rust)
- Provide fallback: "Install manually with:
pip3 install openai-whisper"
- Check available disk space (models require 1-10GB depending on size)
5. Download interrupted or failed
- Check internet connection
- Verify sufficient disk space
- Try again with
--no-check-certificate if SSL issues occur
6. Multiple subtitle languages
- By default, yt-dlp downloads all available languages
- Can specify with
--sub-langs en for English only
- List available with
--list-subs first
Best Practices:
- ✅ Always check what's available before attempting download (
--list-subs)
- ✅ Verify success at each step before proceeding to next
- ✅ Ask user before large downloads (audio files, Whisper models)
- ✅ Clean up temporary files after processing
- ✅ Provide clear feedback about what's happening at each stage
- ✅ Handle errors gracefully with helpful messages