Speak text aloud via ElevenLabs TTS. Voice is the primary communication channel. Audio queues sequentially across all agents.
Generate a 3-frame lip-sync portrait set for a speak voice and upload it to the daemon. Use when adding a new voice or refreshing an existing portrait.