用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/crazynomad/skills --skill tts命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Generate images (Nano Banana Pro / Imagen) and videos (Veo 3.1) through Google Flow using the account's Ultra/Pro SUBSCRIPTION credits instead of the metered Gemini/Vertex API — no per-call API cost. Use when the user wants to create a thumbnail, cover, poster, b-roll clip, image, or short video with AI and wants to avoid API billing, or explicitly mentions Flow / gflow / Veo / Nano Banana via subscription. Wraps the gflow-cli tool; encodes this machine's Flow new-UI quirk (PREFER_CLASSIC) and credit-safe retry rules learned the hard way. NOT for the paid Gemini-API nano-banana-pro path (that costs money) — this is the subscription path.
Decide whether a task in the current project is worth running as an agent loop. Analyzes the repo for evidence first (tests/CI/bench scripts = available verifiers; issue+PR queues = recurring work; module boundaries), then interviews the user one question at a time on the genuine decisions, and returns one of three verdicts — don't loop (stay in the loop yourself), timer loop (/loop, /schedule), or goal loop (/goal) — with cited evidence and, when looping is warranted, a drafted four-part contract (goal / verification / boundary / stop) bound to real commands found in the repo. Recommending NO loop is a first-class outcome. Use when the user asks 值不值得 loop / should I loop this / 要不要上 /goal / 这个活能不能挂个循环自动跑, or wants to apply loop engineering to a project.
GitHub backlog governance manager-loop. Triage open issues (type + routing labels), complete thin descriptions, maintain a bounded ready queue (Todo ≤ 5) on a GitHub Projects board, and repair board drift (closed issue still "In Progress" etc.). Config-driven — reads .claude/backlog-manager.yaml from the target repo; runs an init flow to generate it if missing. DRY-RUN by default, pass "apply" to execute writes. Use for recurring backlog grooming / issue triage of any GitHub repo, standalone or driven by /loop. Requires gh CLI with repo + project scopes.
基于 SOC 职业分类
正在显示 SKILL.md
| name | tts |
| description | Text-to-speech and speech-to-text on Apple Silicon using Vox CLI (Qwen3-TTS + MLX) |
Local text-to-speech, speech-to-text, and voice cloning powered by Qwen3-TTS/ASR + MLX on Apple Silicon.
A powerful local TTS/STT skill based on Vox CLI that runs entirely on Apple Silicon Macs. Features speech synthesis with preset and custom voices, voice cloning from audio samples, speech recognition with subtitle generation, and batch processing. All models run locally via MLX - your data never leaves your machine.
Use this skill when users:
python scripts/vox_tts.py speak "Hello, this is a test" [OPTIONS]
python scripts/vox_tts.py transcribe audio.wav [OPTIONS]
python scripts/vox_tts.py clone "Text to speak" --ref voice_sample.wav [OPTIONS]
Speak text with default voice:
python scripts/vox_tts.py speak "Hello world" -o ./output
Speak text and play immediately:
python scripts/vox_tts.py speak "Hello world" --play
Speak with a specific voice:
python scripts/vox_tts.py speak "Hello world" --voice Chelsie -o ./output
Speak with emotion/style instruction:
python scripts/vox_tts.py speak "I can't believe it!" --instruct "excited and surprised" -o ./output
Speak Chinese text:
python scripts/vox_tts.py speak "你好,世界" --voice Chelsie -o ./output
Clone a voice and speak:
python scripts/vox_tts.py clone "Text in the cloned voice" --ref sample.wav -o ./output
Register a cloned voice for reuse:
python scripts/vox_tts.py clone --ref sample.wav --register my-voice
python scripts/vox_tts.py speak "Now using my custom voice" --voice my-voice
Design a voice from description:
python scripts/vox_tts.py design "Hello everyone" --desc "A warm, friendly female voice with a slight British accent" -o ./output
Transcribe audio to text:
python scripts/vox_tts.py transcribe recording.wav
Transcribe with subtitles:
python scripts/vox_tts.py transcribe recording.wav --subtitle srt -o ./output
Batch TTS from file (one line per utterance):
python scripts/vox_tts.py batch texts.txt --voice Chelsie -o ./output
Use large model for higher quality:
python scripts/vox_tts.py speak "High quality speech" --model large -o ./output
List available voices:
python scripts/vox_tts.py voices
speak - Text-to-Speech| Argument | Description | Default |
|---|---|---|
text | Text to synthesize (required) | - |
-o, --output | Output directory | Current directory |
-v, --voice | Voice name (see voices command) | Chelsie |
-m, --model | Model size: small, large, large-hq | small |
-s, --speed | Speech speed multiplier | 1.0 |
-i, --instruct | Emotion/style instruction | None |
--play | Play audio after generation | False |
--subtitle | Generate subtitle: srt or vtt | None |
transcribe - Speech-to-Text| Argument | Description | Default |
|---|---|---|
audio | Audio file path (required) | - |
-o, --output | Output directory | Current directory |
--subtitle | Output format: srt, vtt, json | plain text |
--language | Source language | auto-detect |
clone - Voice Cloning| Argument | Description | Default |
|---|---|---|
text | Text to synthesize | None |
--ref | Reference audio file (3+ seconds, required) | - |
--register | Register as reusable voice name | None |
-o, --output | Output directory | Current directory |
design - Voice Design| Argument | Description | Default |
|---|---|---|
text | Text to synthesize (required) | - |
--desc | Voice description in natural language (required) | - |
-o, --output | Output directory | Current directory |
batch - Batch Processing| Argument | Description | Default |
|---|---|---|
file | Text file, one utterance per line (required) | - |
-v, --voice | Voice name | Chelsie |
-o, --output | Output directory | Current directory |
# Requirements: Apple Silicon Mac (M1/M2/M3/M4), Python 3.10+, macOS 13+
# Install via pipx (recommended, global CLI)
brew install pipx
pipx ensurepath
git clone https://github.com/3Craft/tts.git /tmp/vox-tts
cd /tmp/vox-tts && pipx install .
# Or install via pip in a venv
pip install -e /path/to/tts
# Optional: Chinese text support
pipx inject vox-cli 'misaki[zh]'
# Optional: Japanese text support
pipx inject vox-cli 'misaki[ja]'
OutputDir/
├── output.wav # Generated audio
└── output.srt # Subtitle (if --subtitle srt)
OutputDir/
├── 001_first_line.wav
├── 002_second_line.wav
└── ...
OutputDir/
├── recording.txt # Plain text transcript
└── recording.srt # Subtitle (if --subtitle srt)
When user requests TTS, STT, or voice cloning:
Read skill documentation:
view("/mnt/skills/user/tts/SKILL.md")
Check if vox is installed:
python /mnt/skills/user/tts/scripts/vox_tts.py check
Install if needed:
brew install pipx && pipx ensurepath
git clone https://github.com/3Craft/tts.git /tmp/vox-tts
cd /tmp/vox-tts && pipx install .
Execute command:
# TTS
python /mnt/skills/user/tts/scripts/vox_tts.py speak "TEXT" \
--voice Chelsie -o /mnt/user-data/outputs
# STT
python /mnt/skills/user/tts/scripts/vox_tts.py transcribe audio.wav \
-o /mnt/user-data/outputs
# Voice cloning
python /mnt/skills/user/tts/scripts/vox_tts.py clone "TEXT" \
--ref sample.wav -o /mnt/user-data/outputs
Present files to user:
present_files(["/mnt/user-data/outputs/..."])
Run vox voices to see all preset voices. Some commonly used:
| Voice | Description |
|---|---|
| Chelsie | Default female voice |
| Ethan | Male voice |
Custom voices can be registered via clone --register.
large-hq for highest qualityvox serve) keeps models in memory for sub-second responseQ: First run is slow?
A: Models are downloaded on first use (1-3 GB). Subsequent runs are much faster. Use daemon mode (vox serve) for instant response.
Q: "Not Apple Silicon" error? A: Vox CLI requires Apple Silicon (M1/M2/M3/M4). It uses MLX which only runs on Apple's Neural Engine.
Q: Chinese/Japanese text sounds wrong?
A: Install language support: pipx inject vox-cli 'misaki[zh]' or pipx inject vox-cli 'misaki[ja]'
Q: Out of memory?
A: Use the small model (default) instead of large. Close other apps to free memory.
Q: How to use a custom voice persistently?
A: Register it: vox clone --ref sample.wav --register my-voice, then use --voice my-voice.
User: "帮我把这段文字转成语音:今天天气真好"
Claude:
python /mnt/skills/user/tts/scripts/vox_tts.py speak "今天天气真好" \
--voice Chelsie -o /mnt/user-data/outputs --play
User: "Transcribe this audio file to SRT subtitles"
Claude:
python /mnt/skills/user/tts/scripts/vox_tts.py transcribe recording.wav \
--subtitle srt -o /mnt/user-data/outputs
User: "用这段录音克隆一个声音,然后用它朗读一段话"
Claude:
# Register the cloned voice
python /mnt/skills/user/tts/scripts/vox_tts.py clone \
--ref sample.wav --register custom-voice
# Speak with the cloned voice
python /mnt/skills/user/tts/scripts/vox_tts.py speak "这是用克隆声音朗读的文字" \
--voice custom-voice -o /mnt/user-data/outputs --play
v1.0 (Current)