Skip to main content
Run any Skill in Manus
with one click

audio-ml-pipeline

Stars2
Forks0
UpdatedJune 25, 2026 at 11:58

Audio / speech ML pipeline design — task framing (ASR/speech-to-text, TTS, speaker ID/verification, sound-event classification, keyword spotting, diarization/who-spoke-when, audio anomaly), representation selection (raw waveform / mel-spectrogram / log-mel / MFCC) with sample-rate + framing + VAD decisions, model family (Whisper / wav2vec2 / HuBERT / AST / pyannote, pretrained-then-finetune), augmentation (SpecAugment / noise / reverb / speed-pitch), task-specific metrics (WER/CER, DER, event-F1/mAP, EER, MOS), and the audio-specific leakage failure modes. Use when building any audio or speech ML pipeline, transcription, voice, or acoustic-monitoring system.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly