Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

asr-transcribe-to-text

النجوم١٬٢٧٧
التفرعات٢١١
آخر تحديث١٦ يوليو ٢٠٢٦ في ٠٨:٠٥

Transcribes audio and video to speaker-labeled text: Qwen3-ASR full-audio transcription + whisper word timing + pyannote diarization, aligned into a who-said-what transcript BY DEFAULT (decoupled WhisperX-style — the audio is never cut before ASR). Handles local files, direct media URLs, and podcast/web pages; local MLX inference on macOS Apple Silicon or remote OpenAI-compatible ASR endpoints. Use when the user wants to transcribe recordings, podcasts, lectures, interviews, meetings, screen recordings, or any audio/video file; also use for ASR, Qwen ASR, speech-to-text, 转录, 语音转文字, 录音转文字, speaker diarization, who said what, 说话人分离, 说话人识别, 谁在说话 — speaker labels are the default, plain text is the opt-out. Also covers word-level timestamps via mlx-whisper for subtitles and audio-visual alignment (字幕, 时间戳, 音画对齐) and CAM++ voiceprint speaker identification. Also preprocesses: merging multi-segment recorder dumps (多段录音合并/拼接) and pitch-preserved speedup for metered-ASR quota uploads (飞书妙记/Feishu Minutes).

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
19 ملفات
SKILL.md
readonly