Skip to main content

audio-transcriber

Extracts audio from dashcam MP4 files and produces GPU-accelerated timestamped transcripts with optional speaker diarization. This skill should be used when users request audio transcription from video files, mention dashcam audio/transcribe MP4/extract speech, want to analyze conversations from video footage, need timestamped transcripts with speaker identification, or ask to process video folders with audio extraction.

معلومات المصدر

المستودع
majiayu000/claude-skill-registry-data
آخر نشاط في المصدر
٢١ أبريل ٢٠٢٦ في ٠٢:٢١
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٢٢
التفرعات
٨

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

مستكشف الملفات
2 ملفات

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
audio-transcriber
description
Extracts audio from dashcam MP4 files and produces GPU-accelerated timestamped transcripts with optional speaker diarization. This skill should be used when users request audio transcription from video files, mention dashcam audio/transcribe MP4/extract speech, want to analyze conversations from video footage, need timestamped transcripts with speaker identification, or ask to process video folders with audio extraction.
# Audio Transcriber **Skill Type:** Media Processing & Analysis **Domain:** Audio Transcription, Speech Recognition, GPU Acceleration **Version:** 2.0 **Last Updated:** 2025-10-26 --- ## Description Extracts audio from dashcam MP4 files and produces GPU-accelerated timestamped transcripts with optional speaker diarization. Uses faster-whisper with CUDA for efficient processing, organizing outputs by date with comprehensive metadata and quality metrics. **When to Use This Skill:** - User requests audio transcription from video files - User mentions "dashcam audio", "transcribe MP4", or "extract speech" - User wants to analyze conversations from video footage - User needs timestamped transcripts with speaker identification - User asks to process video folders with audio extraction --- ## Quick Start ### User Trigger Phrases - "Transcribe audio from my dashcam videos" - "Extract and transcribe speech from [folder/date]" - "Generate transcripts for [MP4 files/date range]" - "Process dashcam audio with speaker identification" - "Create subtitles from video files" ### Expected Inputs 1. **Video Folder Path** (required) - Path to MP4 files or date-organized folders 2. **Date Range** (optional) - Single day, range, or "all available" 3. **Output Directory** (optional) - Default: parallel to input with `_transcripts` suffix 4. **Processing Options** (optional) - Model size, formats, diarization, GPU settings ### Expected Outputs - Audio extracts (WAV files) organized by date - Transcripts in multiple formats (TXT, JSON, SRT, VTT) - Global INDEX.csv with searchable segment metadata - Results JSON with GPU metrics and processing statistics - Quality reports with confidence scores and coverage --- ## Core Capabilities ### 1. User Input Acquisition (Section 0 Protocol) **CRITICAL: Always follow the Section 0 protocol before processing.** #### Step 1: Parse User Request ``` User: "Transcribe my dashcam videos from September 3rd" AI Detection: ✅ Date detected: 2025-09-03 ❌ Folder path not specified → Try auto-discovery ``` #### Step 2: Auto-Discovery Search these locations in order: 1. `C:\Users\[user]\Desktop\CARDV\Movie_F\YYYYMMDD\*.MP4` 2. `G:\My Drive\PROJECTS\INVESTIGATION\DASHCAM\Movie_F\YYYYMMDD\*.MP4` 3. Current working directory + `\*.mp4` 4. User-provided explicit path #### Step 3: Validate Files ```python def validate_video_folder(folder_path): # Check folder exists # Count MP4 files # Estimate total audio duration # Check disk space for outputs # Verify FFmpeg installation ``` #### Step 4: Prompt for Missing Inputs If files not found: ``` 🎥 I need video files to transcribe. I couldn't find MP4s in the expected location. Please provide ONE of the following: 1. **Direct folder path**: e.g., `C:\Users\yousu\Desktop\CARDV\Movie_F\20250903` 2. **Parent directory**: I'll search for date subfolders 3. **Single video file**: Process just one MP4 💡 Tip: Dashcam folders are usually organized by date (YYYYMMDD) ``` #### Step 5: Configuration Summary & Confirmation ``` 📋 **Transcription Configuration Summary** **Input:** 📁 Folder: C:\Users\yousu\Desktop\CARDV\Movie_F\20250903 📹 Videos found: 47 MP4 files ⏱️ Estimated audio: ~7.8 hours 💾 Estimated output size: ~450 MB (transcripts only) **Processing:** 🖥️ GPU: NVIDIA GeForce RTX 4080 (detected) 🧠 Model: faster-whisper base (FP16, CUDA) 🎯 Segmentation: fixed 30s chunks 🗣️ Diarization: disabled (opt-in) 📝 Formats: txt, json, srt **Output:** 💾 Audio extracts: C:\Users\yousu\Desktop\CARDV\Movie_F\20250903\audio\ 📄 Transcripts: C:\Users\yousu\Desktop\CARDV\Movie_F\20250903\transcripts\ 📊 INDEX.csv: C:\Users\yousu\Desktop\CARDV\Movie_F\20250903\transcripts\INDEX.csv Ready to proceed? (Yes/No) ``` **NEVER begin processing without user confirmation.** --- ### 2. Audio Processing Pipeline #### A. Audio Extraction (FFmpeg with Retry Matrix) ```python # Primary extraction command ffmpeg -i video.mp4 -vn -acodec pcm_s16le -ar 16000 -ac 1 audio.wav # Retry sequence on failure: # 1. Codec fallback: pcm_s16le → flac # 2. Add demuxer args: -fflags +genpts -rw_timeout 30000000 # 3. Extended probe: -analyzeduration 100M -probesize 100M ``` **Quality Checks:** - Verify audio stream exists (ffprobe preflight) - Check duration matches video duration - Detect silent/corrupted audio - Log extraction errors to `_FAILED.json` #### B. Segmentation (Two Modes) **Fixed Mode (Default):** - Split audio into 30-second chunks - Predictable processing time - No external VAD required - Best for continuous speech **VAD Mode (Advanced):** - Use Silero VAD to detect speech regions - Variable-length segments (2-60s) - Skip long silences - Best for sparse audio (parking mode) **Mutual Exclusion:** Only one mode active at a time. #### C. GPU Transcription (faster-whisper) ```python # Load model with GPU optimization model = WhisperModel( "base", device="cuda", compute_type="float16" ) # Transcribe with word-level timestamps segments, info = model.transcribe( audio_path, beam_size=5, word_timestamps=True, vad_filter=True ) ``` **GPU Metrics Captured:** - Device name, VRAM, utilization - CUDA version, driver version - Average GPU % during run (sampled at 1-2 Hz) - Memory usage peaks #### D. Speaker Diarization (Optional) **Backends:** - **pyannote**: State-of-the-art (requires HF token + VRAM) - **speechbrain**: Good performance (no auth required) **Label Normalization:** - Different backends → unified `spkA`, `spkB`, etc. - Consistent across INDEX.csv and JSON outputs **Fallback Behavior:** - If HF token missing → skip diarization, log warning - If OOM error → disable diarization, continue transcription --- ### 3. Output Generation #### A. File Organization (Per-Day Structure) ``` C:\Users\yousu\Desktop\CARDV\Movie_F\ └── 20250903\ ├── audio\ │ ├── 20250903133516_059495B.wav │ ├── 20250903134120_059496B.wav │ └── ... (47 files) ├── transcripts\ │ ├── 20250903133516_059495B.txt │ ├── 20250903133516_059495B.json │ ├── 20250903133516_059495B.srt │ └── ... (47 × 3 = 141 files) └── INDEX.csv ``` #### B. Format Details **TXT (Plain Text):** ``` [00:00:15] Speaker A: Hey, where are we going? [00:00:18] Speaker B: Just heading to the mall. [00:00:22] Speaker A: Okay, sounds good. ``` **JSON (Complete Metadata):** ```json { "video_file": "20250903133516_059495B.MP4", "audio_duration_sec": 60, "language": "en", "language_confidence": 0.95, "segments": [ { "start": 15.2, "end": 17.8, "text": "Hey, where are we going?", "confidence": 0.89, "speaker": "spkA", "words": [ {"word": "Hey", "start": 15.2, "end": 15.4, "confidence": 0.92}, {"word": "where", "start": 15.5, "end": 15.8, "confidence": 0.88} ] } ] } ``` **SRT (SubRip Subtitles):** ``` 1 00:00:15,200 --> 00:00:17,800 [spkA] Hey, where are we going? 2 00:00:18,000 --> 00:00:20,500 [spkB] Just heading to the mall. ``` **VTT (WebVTT):** ``` WEBVTT 00:00:15.200 --> 00:00:17.800 <v spkA>Hey, where are we going? 00:00:18.000 --> 00:00:20.500 <v spkB>Just heading to the mall. ``` #### C. INDEX.csv (Global Search Index) Composite key: `(video_rel, seg_idx)` | Column | Description | |--------|-------------| | `dataset` | Movie_F / Movie_R / Park_F / Park_R | | `date` | YYYYMMDD | | `video_rel` | Relative path from root | | `video_stem` | Filename without extension | | `seg_idx` | 0-based segment index | | `ts_start_ms` | Segment start milliseconds | | `ts_end_ms` | Segment end milliseconds | | `text` | Transcript text (truncated to 512 chars) | | `text_len` | Full text length | | `lang` | ISO language code | | `lang_conf` | Language detection confidence | | `conf_avg` | Average token confidence | | `speaker` | Normalized speaker label | | `format_mask` | Files generated (txt/json/srt/vtt) | | `transcript_file` | Basename | | `audio_file` | Basename | | `engine` | e.g., `faster-whisper:base:fp16` | | `cuda_version` | CUDA version | | `driver_version` | Driver version | | `created_utc` | ISO 8601 timestamp | #### D. Results JSON (Single Source of Truth) ```json { "status": "ok", "summary": { "videos_processed": 47, "segments": 1847, "hours_audio": 7.8, "gpu_detected": true, "device_count": 1, "devices": [ { "index": 0, "name": "NVIDIA GeForce RTX 4080", "total_mem_mb": 16384, "free_mem_mb": 14200 } ], "utilization": { "gpu_pct": 35, "mem_pct": 42, "sampling_hz": 2 }, "cuda_version": "12.1", "driver_version": "546.01", "torch_version": "2.2.0+cu121", "errors": 0, "failed_files": [] }, "artifacts": { "index_csv": "C:\\Users\\yousu\\Desktop\\CARDV\\Movie_F\\20250903\\INDEX.csv", "output_dir": "C:\\Users\\yousu\\Desktop\\CARDV\\Movie_F\\20250903\\transcripts" } } ``` --- ### 4. Quality & Error Handling #### A. Resume Safety - Skip existing transcripts unless `--force` flag - Idempotent: re-running is safe - Checkpoint support for long runs #### B. Error Types & Recovery **Per-Video Failures** (`{video_stem}_FAILED.json`): ```json { "video_path": "C:\\...\\video.mp4", "error_type": "ffmpeg_err", "error_message": "Failed to decode audio stream", "ffprobe_metadata": {"duration": null, "codec": "h264"}, "timestamp": "2025-09-03T14:30:00Z" } ``` Error types: - `ffmpeg_err`: Audio extraction failed - `decode_err`: Whisper decode failed - `OOM`: Out of GPU memory - `corrupted`: Container/stream corrupted - `no_audio`: No audio stream detected #### C. SRT/VTT Validation - Strictly monotonic timestamps - No overlapping segments - Clamp gaps <50ms - Proper timecode formatting (comma vs period) --- ## Implementation Guide ### Phase 1: Input Acquisition ```python # 1. Parse user request inputs = parse_user_request(user_message) # 2. Auto-discover video files if not inputs['video_folder']: inputs['video_folder'] = auto_discover_videos() # 3. Validate inputs validate_video_folder(inputs['video_folder']) check_ffmpeg_available() check_gpu_available() # 4. Estimate resource requirements estimate_processing_time(inputs) estimate_disk_space(inputs) # 5. Present configuration summary show_configuration_summary(inputs) # 6. Wait for confirmation if not user_confirms(): return # Do not proceed ``` ### Phase 2: Audio Extraction ```python for video_file in video_files: # FFprobe preflight check metadata = ffprobe(video_file) if not has_audio_stream(metadata): log_failed(video_file, "no_audio") continue # Extract audio with retry try: audio_path = extract_audio_ffmpeg( video_file, output_dir=audio_output_dir, sample_rate=16000, channels=1 ) except FFmpegError as e: # Retry with fallback codec
عرض على GitHub
ملف SKILL.md هذا كبير جدا، لذلك يعرض SkillsMP القسم الاول فقط هنا. عرض على GitHub