소스 정보
- 저장소
- crazynomad/skills
- 최근 소스 활동
- 2026년 3월 29일 00:48
- 감지된 SKILL.md 언어
- 영어
- 스타
- 27
- 포크
- 7
설치 방법
기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.
소스 파일 검토
설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.
메뉴
기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.
설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/crazynomad/skills --skill tts명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
SOC 직업 분류 기준
SKILL.md 표시 중
| name | tts |
| description | Text-to-speech and speech-to-text on Apple Silicon using Vox CLI (Qwen3-TTS + MLX) |
Local text-to-speech, speech-to-text, and voice cloning powered by Qwen3-TTS/ASR + MLX on Apple Silicon.
A powerful local TTS/STT skill based on Vox CLI that runs entirely on Apple Silicon Macs. Features speech synthesis with preset and custom voices, voice cloning from audio samples, speech recognition with subtitle generation, and batch processing. All models run locally via MLX - your data never leaves your machine.
Use this skill when users:
python scripts/vox_tts.py speak "Hello, this is a test" [OPTIONS]
python scripts/vox_tts.py transcribe audio.wav [OPTIONS]
python scripts/vox_tts.py clone "Text to speak" --ref voice_sample.wav [OPTIONS]
Speak text with default voice:
python scripts/vox_tts.py speak "Hello world" -o ./output
Speak text and play immediately:
python scripts/vox_tts.py speak "Hello world" --play
Speak with a specific voice:
python scripts/vox_tts.py speak "Hello world" --voice Chelsie -o ./output
Speak with emotion/style instruction:
python scripts/vox_tts.py speak "I can't believe it!" --instruct "excited and surprised" -o ./output
Speak Chinese text:
python scripts/vox_tts.py speak "你好,世界" --voice Chelsie -o ./output
Clone a voice and speak:
python scripts/vox_tts.py clone "Text in the cloned voice" --ref sample.wav -o ./output
Register a cloned voice for reuse:
python scripts/vox_tts.py clone --ref sample.wav --register my-voice
python scripts/vox_tts.py speak "Now using my custom voice" --voice my-voice
Design a voice from description:
python scripts/vox_tts.py design "Hello everyone" --desc "A warm, friendly female voice with a slight British accent" -o ./output
Transcribe audio to text:
python scripts/vox_tts.py transcribe recording.wav
Transcribe with subtitles:
python scripts/vox_tts.py transcribe recording.wav --subtitle srt -o ./output
Batch TTS from file (one line per utterance):
python scripts/vox_tts.py batch texts.txt --voice Chelsie -o ./output
Use large model for higher quality:
python scripts/vox_tts.py speak "High quality speech" --model large -o ./output
List available voices:
python scripts/vox_tts.py voices
speak - Text-to-Speech| Argument | Description | Default |
|---|---|---|
text | Text to synthesize (required) | - |
-o, --output | Output directory | Current directory |
-v, --voice | Voice name (see voices command) | Chelsie |
-m, --model | Model size: small, large, large-hq | small |
-s, --speed | Speech speed multiplier | 1.0 |
-i, --instruct | Emotion/style instruction | None |
--play | Play audio after generation | False |
--subtitle | Generate subtitle: srt or vtt | None |
transcribe - Speech-to-Text| Argument | Description | Default |
|---|---|---|
audio | Audio file path (required) | - |
-o, --output | Output directory | Current directory |
--subtitle | Output format: srt, vtt, json | plain text |
--language | Source language | auto-detect |
clone - Voice Cloning| Argument | Description | Default |
|---|---|---|
text | Text to synthesize | None |
--ref | Reference audio file (3+ seconds, required) | - |
--register | Register as reusable voice name | None |
-o, --output | Output directory | Current directory |
design - Voice Design| Argument | Description | Default |
|---|---|---|
text | Text to synthesize (required) | - |
--desc | Voice description in natural language (required) | - |
-o, --output | Output directory | Current directory |
batch - Batch Processing| Argument | Description | Default |
|---|---|---|
file | Text file, one utterance per line (required) | - |
-v, --voice | Voice name | Chelsie |
-o, --output | Output directory | Current directory |
# Requirements: Apple Silicon Mac (M1/M2/M3/M4), Python 3.10+, macOS 13+
# Install via pipx (recommended, global CLI)
brew install pipx
pipx ensurepath
git clone https://github.com/3Craft/tts.git /tmp/vox-tts
cd /tmp/vox-tts && pipx install .
# Or install via pip in a venv
pip install -e /path/to/tts
# Optional: Chinese text support
pipx inject vox-cli 'misaki[zh]'
# Optional: Japanese text support
pipx inject vox-cli 'misaki[ja]'
OutputDir/
├── output.wav # Generated audio
└── output.srt # Subtitle (if --subtitle srt)
OutputDir/
├── 001_first_line.wav
├── 002_second_line.wav
└── ...
OutputDir/
├── recording.txt # Plain text transcript
└── recording.srt # Subtitle (if --subtitle srt)
When user requests TTS, STT, or voice cloning:
Read skill documentation:
view("/mnt/skills/user/tts/SKILL.md")
Check if vox is installed:
python /mnt/skills/user/tts/scripts/vox_tts.py check
Install if needed:
brew install pipx && pipx ensurepath
git clone https://github.com/3Craft/tts.git /tmp/vox-tts
cd /tmp/vox-tts && pipx install .
Execute command:
# TTS
python /mnt/skills/user/tts/scripts/vox_tts.py speak "TEXT" \
--voice Chelsie -o /mnt/user-data/outputs
# STT
python /mnt/skills/user/tts/scripts/vox_tts.py transcribe audio.wav \
-o /mnt/user-data/outputs
# Voice cloning
python /mnt/skills/user/tts/scripts/vox_tts.py clone "TEXT" \
--ref sample.wav -o /mnt/user-data/outputs
Present files to user:
present_files(["/mnt/user-data/outputs/..."])
Run vox voices to see all preset voices. Some commonly used:
| Voice | Description |
|---|---|
| Chelsie | Default female voice |
| Ethan | Male voice |
Custom voices can be registered via clone --register.
large-hq for highest qualityvox serve) keeps models in memory for sub-second responseQ: First run is slow?
A: Models are downloaded on first use (1-3 GB). Subsequent runs are much faster. Use daemon mode (vox serve) for instant response.
Q: "Not Apple Silicon" error? A: Vox CLI requires Apple Silicon (M1/M2/M3/M4). It uses MLX which only runs on Apple's Neural Engine.
Q: Chinese/Japanese text sounds wrong?
A: Install language support: pipx inject vox-cli 'misaki[zh]' or pipx inject vox-cli 'misaki[ja]'
Q: Out of memory?
A: Use the small model (default) instead of large. Close other apps to free memory.
Q: How to use a custom voice persistently?
A: Register it: vox clone --ref sample.wav --register my-voice, then use --voice my-voice.
User: "帮我把这段文字转成语音:今天天气真好"
Claude:
python /mnt/skills/user/tts/scripts/vox_tts.py speak "今天天气真好" \
--voice Chelsie -o /mnt/user-data/outputs --play
User: "Transcribe this audio file to SRT subtitles"
Claude:
python /mnt/skills/user/tts/scripts/vox_tts.py transcribe recording.wav \
--subtitle srt -o /mnt/user-data/outputs
User: "用这段录音克隆一个声音,然后用它朗读一段话"
Claude:
# Register the cloned voice
python /mnt/skills/user/tts/scripts/vox_tts.py clone \
--ref sample.wav --register custom-voice
# Speak with the cloned voice
python /mnt/skills/user/tts/scripts/vox_tts.py speak "这是用克隆声音朗读的文字" \
--voice custom-voice -o /mnt/user-data/outputs --play
v1.0 (Current)