Transcribe user-sent voice messages via local OpenAI Whisper (no API key, runs offline)
Generate character-consistent selfie images and videos, send via the active Nako messaging channel. Character reference image and description come from per-agent env.
Resolve Feishu image_key to a local file path so the agent can Read the image natively
Send voice/song audio through the active Nako messaging channel — voice.sh does plain TTS (MiniMax or Volcengine); sing.sh generates a full song with melody via MiniMax music-2.6
Control interactive BLE devices (scan/connect/playback/timeline) from terminal.