| name | text-to-voice |
| description | Convert text to speech using Kyutai's Pocket TTS. Use when the user asks to "generate speech", "text to speech", "TTS", "convert text to audio", "voice synthesis", "generate voice", "read aloud", or "create audio from text". Supports voice cloning from audio samples and multiple pre-made voices (alba, marius, javert, jean, fantine, cosette, eponine, azelma). |
| license | MIT |
| metadata | {"contributor":"Aaron Adetunmbi","thanks":"kyutai-labs"} |
Text-to-Voice with Kyutai Pocket TTS
Convert text to natural speech using Kyutai's Pocket TTS - a lightweight 100M parameter model that runs efficiently on CPU.
Installation
pip install pocket-tts
uvx pocket-tts generate
Requires Python 3.10+ and PyTorch 2.5+. GPU not required.
CLI Usage
Basic Generation
uvx pocket-tts generate
pocket-tts generate --text "Hello, this is my message."
pocket-tts generate --text "Hello" --output-path ./audio/greeting.wav
pocket-tts generate \
--text "Welcome to the demo." \
--voice alba \
--output-path ./output/welcome.wav
CLI Options
| Option | Default | Description |
|---|
--text | "Hello world..." | Text to convert to speech |
--voice | alba | Voice name, local file path, or HuggingFace URL |
--output-path | ./tts_output.wav | Where to save the generated audio file |
--temperature | 0.7 | Generation temperature (higher = more expressive) |
--lsd-decode-steps | 1 | Quality steps (higher = better quality, slower) |
--eos-threshold | -4.0 | End detection threshold (lower = finish earlier) |
--frames-after-eos | auto | Extra frames after end (each frame = 80ms) |
--device | cpu | Device to use (cpu/cuda) |
-q, --quiet | false | Disable logging output |
Voice Selection (CLI)
pocket-tts generate --voice alba --text "Hello"
pocket-tts generate --voice javert --text "Hello"
pocket-tts generate --voice ./my_voice.wav --text "Hello"
pocket-tts generate --voice "hf://kyutai/tts-voices/alba-mackenna/merchant.wav" --text "Hello"
Quality Tuning (CLI)
pocket-tts generate --lsd-decode-steps 5 --temperature 0.5 --output-path high_quality.wav
pocket-tts generate --temperature 1.0 --output-path expressive.wav
pocket-tts generate --eos-threshold -3.0 --output-path shorter.wav
Local Web Server
For quick iteration with multiple voices/texts:
uvx pocket-tts serve
Available Voices
Pre-made voices (use name directly with --voice):
| Voice | Gender | License | Description |
|---|
alba | Female | CC BY 4.0 | Casual voice |
marius | Male | CC0 | Voice donation |
javert | Male | CC0 | Voice donation |
jean | Male | CC-NC | EARS dataset |
fantine | Female | CC BY 4.0 | VCTK dataset |
cosette | Female | CC-NC | Expresso dataset |
eponine | Female | CC BY 4.0 | VCTK dataset |
azelma | Female | CC BY 4.0 | VCTK dataset |
Full voice catalog: https://huggingface.co/kyutai/tts-voices
For detailed voice information, see references/voices.md.
Voice Cloning
Clone any voice from an audio sample. For best results:
- Use clean audio (minimal background noise)
- 10+ seconds recommended
- Consider Adobe Podcast Enhance to clean samples
pocket-tts generate --voice ./my_recording.wav --text "Hello" --output-path cloned.wav
Output Format
- Sample Rate: 24kHz
- Channels: Mono
- Format: 16-bit PCM WAV
- Default location:
./tts_output.wav
Python API
For programmatic use:
from pocket_tts import TTSModel
import scipy.io.wavfile
tts_model = TTSModel.load_model()
voice_state = tts_model.get_state_for_audio_prompt("alba")
audio = tts_model.generate_audio(voice_state, "Hello world!")
scipy.io.wavfile.write("./audio/output.wav", tts_model.sample_rate, audio.numpy())
TTSModel.load_model()
model = TTSModel.load_model(
variant="b6369a24",
temp=0.7,
lsd_decode_steps=1,
noise_clamp=None,
eos_threshold=-4.0
)
Voice State
voice_state = model.get_state_for_audio_prompt("alba")
voice_state = model.get_state_for_audio_prompt("./my_voice.wav")
voice_state = model.get_state_for_audio_prompt("hf://kyutai/tts-voices/alba-mackenna/casual.wav")
Generate Audio
audio = model.generate_audio(voice_state, "Text to speak")
Streaming
for chunk in model.generate_audio_stream(voice_state, "Long text..."):
pass
Properties
model.sample_rate - 24000 Hz
model.device - "cpu" or "cuda"
Performance
- ~200ms latency to first audio chunk
- ~6x real-time on MacBook Air M4 CPU
- Uses only 2 CPU cores
Limitations
- English only
- No built-in pause/silence control