Skip to main content

tts-integration

Add or modify text-to-speech providers in assistant-api with transport-aware behavior (WS/SSE/SDK/HTTP), correct packet lifecycle, and UI/provider wiring.

설치로 이동

소스 정보

저장소
rapidaai/voice-ai
최근 소스 활동
2026년 8월 23일 07:23
감지된 SKILL.md 언어
영어
스타
729
포크
120

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
4 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
tts-integration
description
Add or modify text-to-speech providers in assistant-api with transport-aware behavior (WS/SSE/SDK/HTTP), correct packet lifecycle, and UI/provider wiring.
# TTS Integration Skill ## Mission Integrate TTS providers with correct streaming/flush/interruption semantics and stable `tts_latency_ms` metrics. ## Inputs expected from user 1. Transport model: bidirectional WebSocket, SSE/chunked HTTP, SDK callback, or request-response flush. 2. Voice/model parameter requirements. 3. Output audio format from provider. If user does not answer: - Select nearest baseline transport and preserve existing packet semantics. - Keep internal output format compatibility with existing streamer path. ## Hard boundaries In scope: - `api/assistant-api/internal/transformer/<provider>/tts.go` (+ provider option/normalizer helpers) - `api/assistant-api/internal/transformer/transformer.go` (factory case) - optional contract updates: `api/assistant-api/internal/type/tts_transformer.go`, `packet.go` - provider config JSON + UI component files for TTS Out of scope: - STT-only logic changes unrelated to shared provider setup - telephony transport internals - EOS/VAD internals ## Transport mapping - WS streaming baseline: `deepgram`, `rime`, `sarvam` - SSE/chunked baseline: `minimax` - SDK/API callback baseline: `azure`, `google` - HTTP flush baseline: `aws`, `resembleai` ## Packet contract Input packets: - `LLMResponseDeltaPacket` - `LLMResponseDonePacket` - `InterruptionPacket` Required outputs: - `TextToSpeechAudioPacket` - `TextToSpeechEndPacket` - `ConversationEventPacket{Name:"tts", ...}` - `MessageMetricPacket{Name:"tts_latency_ms"}` once per utterance ## Implementation workflow 1. Pick transport baseline and clone lifecycle shape. 2. Implement provider under `transformer/<provider>/`. 3. Add factory registration in `transformer.go`. 4. Wire UI metadata (`voices`, `languages`, `text-to-speech-models`) and config form. 5. Verify interruption behavior clears output state without deadlocks. 6. Add provider tests (init, streaming, completion, interruption, metric-on-first-audio). ## Governed lifecycle - Classify work as Fast, Standard, or Governed using `DEVELOPMENT_PROCESS.md`; use the full gated lifecycle only for Governed work. - This skill operates only in its assigned phase and path ownership; it may not approve its own plan or code review. - Governed implementation starts only from a coordinator-attested approved plan; Fast and Standard work follows its lighter documented lifecycle. - Return changed-file and verification evidence to the coordinator, then route the complete verified diff to the read-only `code-reviewer`. ## Validation commands - `go test ./api/assistant-api/internal/transformer/... -run TestTTS` - `go test ./api/assistant-api/internal/transformer/<provider>/...` - `cd ui && yarn test providers` - `./.claude/skills/tts-integration/scripts/validate.sh --check-diff --provider <provider>`
GitHub에서 보기