Skip to main content
Run any Skill in Manus
with one click

speech-to-text-master

Stars106
Forks11
UpdatedJune 4, 2026 at 11:17

语音转文字 (ASR) (语音转文字 (ASR / 语音识别 / Speech-to-Text) 应用工程 (从业者视角) — 选型、集成、优化语音转文字能力,尤其面向移动端、低成本、快速识别的场景。覆盖: (a) 引擎/API 地图 — 云端 API(OpenAI gpt-4o-transcribe / Whisper API、Deepgram、AssemblyAI、Google / Azure / AWS Transcribe、讯飞、字节火山、阿里、腾讯、百度) vs 开源模型(Whisper / faster-whisper / whisper.cpp / distil-whisper、NVIDIA NeMo Parakeet / Canary、阿里 FunASR / Paraformer / SenseVoice、Moonshine、Vosk、Kaldi / k2 / icefall) vs 端侧·移动 SDK(whisper.cpp + CoreML / Metal、iOS Speech framework、Android SpeechRecognizer、Picovoice、SenseVoice 端侧); (b) 准确率(WER / CER) × 延迟(RTF / 流式) × 成本 三角权衡与选型决策树; (c) 成本优化 playbook — 端侧免费 / 批量折扣 / VAD 裁静音 / 量化(int8 / ggml) / 蒸馏 / 自托管 break-even; (d) 移动端集成 — 端侧 vs 云、流式 vs 批量、断点检测(endpointing)、隐私 / 离线; (e) 后处理 — 标点 / 数字规整(ITN) / 说话人分离(diarization) / 时间戳。学派分歧: 云 API vs 端侧自托管、通用大模型(Whisper) vs 专用流式(RNN-T / Conformer)、闭源 API vs 开源、准确率派 vs 成本派、英文优先 vs 中文 ASR(FunASR / SenseVoice / 讯飞)。不含: 文字转语音(TTS / 语音合成,方向相反)、声纹识别 / 说话人验证为主业、语音 agent / 对话式 AI、ASR 模型训练科研深水区。) Master OS — automated mastery of 语音转文字 (ASR / 语音识别 / Speech

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
8 files
SKILL.md
readonly