Skip to main content
在 Manus 中运行任何 Skill
一键导入

audio-ml-pipeline

星标2
分支0
更新时间2026年6月25日 11:58

Audio / speech ML pipeline design — task framing (ASR/speech-to-text, TTS, speaker ID/verification, sound-event classification, keyword spotting, diarization/who-spoke-when, audio anomaly), representation selection (raw waveform / mel-spectrogram / log-mel / MFCC) with sample-rate + framing + VAD decisions, model family (Whisper / wav2vec2 / HuBERT / AST / pyannote, pretrained-then-finetune), augmentation (SpecAugment / noise / reverb / speed-pitch), task-specific metrics (WER/CER, DER, event-F1/mAP, EER, MOS), and the audio-specific leakage failure modes. Use when building any audio or speech ML pipeline, transcription, voice, or acoustic-monitoring system.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

SKILL.md
readonly