ソース情報
- リポジトリ
- crazynomad/skills
- ソースの最終更新活動
- 2026年3月29日 00:48
- 検出された SKILL.md の言語
- 英語
- スター
- 27
- フォーク
- 7
インストール方法
デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。
ソースファイルを確認
インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。
メニュー
デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。
インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
直接コマンドでは確認用 Prompt が省略されます。実行前にソースを確認してください。
npx skills add https://github.com/crazynomad/skills --skill ttsコマンドは1行のまま表示されます。コピー前に横へスクロールして全体を確認してください。
ローカルで確認しますか?SkillsMP が現在取得できるファイルをダウンロードできます。
SKILL.md を表示中
SOC 職業分類に基づく
| name | tts |
| description | Text-to-speech and speech-to-text on Apple Silicon using Vox CLI (Qwen3-TTS + MLX) |
Local text-to-speech, speech-to-text, and voice cloning powered by Qwen3-TTS/ASR + MLX on Apple Silicon.
A powerful local TTS/STT skill based on Vox CLI that runs entirely on Apple Silicon Macs. Features speech synthesis with preset and custom voices, voice cloning from audio samples, speech recognition with subtitle generation, and batch processing. All models run locally via MLX - your data never leaves your machine.
Use this skill when users:
python scripts/vox_tts.py speak "Hello, this is a test" [OPTIONS]
python scripts/vox_tts.py transcribe audio.wav [OPTIONS]
python scripts/vox_tts.py clone "Text to speak" --ref voice_sample.wav [OPTIONS]
Speak text with default voice:
python scripts/vox_tts.py speak "Hello world" -o ./output
Speak text and play immediately:
python scripts/vox_tts.py speak "Hello world" --play
Speak with a specific voice:
python scripts/vox_tts.py speak "Hello world" --voice Chelsie -o ./output
Speak with emotion/style instruction:
python scripts/vox_tts.py speak "I can't believe it!" --instruct "excited and surprised" -o ./output
Speak Chinese text:
python scripts/vox_tts.py speak "你好,世界" --voice Chelsie -o ./output
Clone a voice and speak:
python scripts/vox_tts.py clone "Text in the cloned voice" --ref sample.wav -o ./output
Register a cloned voice for reuse:
python scripts/vox_tts.py clone --ref sample.wav --register my-voice
python scripts/vox_tts.py speak "Now using my custom voice" --voice my-voice
Design a voice from description:
python scripts/vox_tts.py design "Hello everyone" --desc "A warm, friendly female voice with a slight British accent" -o ./output
Transcribe audio to text:
python scripts/vox_tts.py transcribe recording.wav
Transcribe with subtitles:
python scripts/vox_tts.py transcribe recording.wav --subtitle srt -o ./output
Batch TTS from file (one line per utterance):
python scripts/vox_tts.py batch texts.txt --voice Chelsie -o ./output
Use large model for higher quality:
python scripts/vox_tts.py speak "High quality speech" --model large -o ./output
List available voices:
python scripts/vox_tts.py voices
speak - Text-to-Speech| Argument | Description | Default |
|---|---|---|
text | Text to synthesize (required) | - |
-o, --output | Output directory | Current directory |
-v, --voice | Voice name (see voices command) | Chelsie |
-m, --model | Model size: small, large, large-hq | small |
-s, --speed | Speech speed multiplier | 1.0 |
-i, --instruct | Emotion/style instruction | None |
--play | Play audio after generation | False |
--subtitle | Generate subtitle: srt or vtt | None |
transcribe - Speech-to-Text| Argument | Description | Default |
|---|---|---|
audio | Audio file path (required) | - |
-o, --output | Output directory | Current directory |
--subtitle | Output format: srt, vtt, json | plain text |
--language | Source language | auto-detect |
clone - Voice Cloning| Argument | Description | Default |
|---|---|---|
text | Text to synthesize | None |
--ref | Reference audio file (3+ seconds, required) | - |
--register | Register as reusable voice name | None |
-o, --output | Output directory | Current directory |
design - Voice Design| Argument | Description | Default |
|---|---|---|
text | Text to synthesize (required) | - |
--desc | Voice description in natural language (required) | - |
-o, --output | Output directory | Current directory |
batch - Batch Processing| Argument | Description | Default |
|---|---|---|
file | Text file, one utterance per line (required) | - |
-v, --voice | Voice name | Chelsie |
-o, --output | Output directory | Current directory |
# Requirements: Apple Silicon Mac (M1/M2/M3/M4), Python 3.10+, macOS 13+
# Install via pipx (recommended, global CLI)
brew install pipx
pipx ensurepath
git clone https://github.com/3Craft/tts.git /tmp/vox-tts
cd /tmp/vox-tts && pipx install .
# Or install via pip in a venv
pip install -e /path/to/tts
# Optional: Chinese text support
pipx inject vox-cli 'misaki[zh]'
# Optional: Japanese text support
pipx inject vox-cli 'misaki[ja]'
OutputDir/
├── output.wav # Generated audio
└── output.srt # Subtitle (if --subtitle srt)
OutputDir/
├── 001_first_line.wav
├── 002_second_line.wav
└── ...
OutputDir/
├── recording.txt # Plain text transcript
└── recording.srt # Subtitle (if --subtitle srt)
When user requests TTS, STT, or voice cloning:
Read skill documentation:
view("/mnt/skills/user/tts/SKILL.md")
Check if vox is installed:
python /mnt/skills/user/tts/scripts/vox_tts.py check
Install if needed:
brew install pipx && pipx ensurepath
git clone https://github.com/3Craft/tts.git /tmp/vox-tts
cd /tmp/vox-tts && pipx install .
Execute command:
# TTS
python /mnt/skills/user/tts/scripts/vox_tts.py speak "TEXT" \
--voice Chelsie -o /mnt/user-data/outputs
# STT
python /mnt/skills/user/tts/scripts/vox_tts.py transcribe audio.wav \
-o /mnt/user-data/outputs
# Voice cloning
python /mnt/skills/user/tts/scripts/vox_tts.py clone "TEXT" \
--ref sample.wav -o /mnt/user-data/outputs
Present files to user:
present_files(["/mnt/user-data/outputs/..."])
Run vox voices to see all preset voices. Some commonly used:
| Voice | Description |
|---|---|
| Chelsie | Default female voice |
| Ethan | Male voice |
Custom voices can be registered via clone --register.
large-hq for highest qualityvox serve) keeps models in memory for sub-second responseQ: First run is slow?
A: Models are downloaded on first use (1-3 GB). Subsequent runs are much faster. Use daemon mode (vox serve) for instant response.
Q: "Not Apple Silicon" error? A: Vox CLI requires Apple Silicon (M1/M2/M3/M4). It uses MLX which only runs on Apple's Neural Engine.
Q: Chinese/Japanese text sounds wrong?
A: Install language support: pipx inject vox-cli 'misaki[zh]' or pipx inject vox-cli 'misaki[ja]'
Q: Out of memory?
A: Use the small model (default) instead of large. Close other apps to free memory.
Q: How to use a custom voice persistently?
A: Register it: vox clone --ref sample.wav --register my-voice, then use --voice my-voice.
User: "帮我把这段文字转成语音:今天天气真好"
Claude:
python /mnt/skills/user/tts/scripts/vox_tts.py speak "今天天气真好" \
--voice Chelsie -o /mnt/user-data/outputs --play
User: "Transcribe this audio file to SRT subtitles"
Claude:
python /mnt/skills/user/tts/scripts/vox_tts.py transcribe recording.wav \
--subtitle srt -o /mnt/user-data/outputs
User: "用这段录音克隆一个声音,然后用它朗读一段话"
Claude:
# Register the cloned voice
python /mnt/skills/user/tts/scripts/vox_tts.py clone \
--ref sample.wav --register custom-voice
# Speak with the cloned voice
python /mnt/skills/user/tts/scripts/vox_tts.py speak "这是用克隆声音朗读的文字" \
--voice custom-voice -o /mnt/user-data/outputs --play
v1.0 (Current)