Skip to main content
Run any Skill in Manus
with one click

hearing-stt

Stars0
Forks0
UpdatedJune 14, 2026 at 14:48

Use when choosing, wiring, swapping, or benchmarking a speech-to-text (STT/ASR) engine behind the hearing project's transcription facade — the pluggable layer that turns audio into segments so the rest of the architecture never names a specific engine. Triggers on transcribe(audio)->segments, ASR engine selection, Whisper / whisper.cpp / faster-whisper / WhisperX / distil-whisper, NVIDIA Parakeet-TDT / Canary, Moonshine, Apple-Silicon MLX paths (lightning-whisper-mlx, mlx-audio, Lightning-SimulWhisper), cloud STT APIs (OpenAI gpt-4o-transcribe / -mini / gpt-realtime-whisper / gpt-4o-transcribe-diarize, Deepgram Nova-3, AssemblyAI Universal, Google Chirp, Azure Speech, Speechmatics, Rev), per-minute STT pricing, local-vs-cloud / on-device decision, WER tolerance, concurrency, self-hosting break-even (~100k min/month), CTranslate2 / int8 quantization, model-name versioning and the ~June-2026 OpenAI transcribe retirements, and "which engine should I use for meeting transcription". Diarization (who-spoke-when) is

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
3 files
SKILL.md
readonly