一键导入
streaming-stt-deepgram
Real-time streaming speech-to-text via Deepgram WebSocket API
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Real-time streaming speech-to-text via Deepgram WebSocket API
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Edit, transform, and enhance images using AI models
Extract text from images using OCR and vision AI
Neural text-to-speech via Amazon Polly
Speaker diarization — identifies and tracks who is speaking at each moment in an audio stream
Semantic endpoint detection — uses an LLM to classify whether the user's utterance is a complete thought, reducing false turn boundaries on mid-sentence pauses
Batch speech-to-text via Google Cloud Speech-to-Text API
| name | streaming-stt-deepgram |
| description | Real-time streaming speech-to-text via Deepgram WebSocket API |
| category | voice |
Real-time streaming speech-to-text using Deepgram's Nova-2 model via WebSocket.
Set DEEPGRAM_API_KEY in your environment or agent secrets store.
speaker labels)providerOptions.keywordsIn agent.config.json:
{
"voice": {
"stt": "deepgram"
}
}
Provider-specific options via providerOptions:
{
"voice": {
"stt": "deepgram",
"providerOptions": {
"model": "nova-2",
"diarize": true,
"keywords": ["AgentOS:2", "Deepgram:1.5"],
"endpointing": 300
}
}
}
| Event | Payload | Description |
|---|---|---|
transcript | TranscriptEvent | Every hypothesis (interim + final) |
interim_transcript | TranscriptEvent | Non-final hypothesis |
final_transcript | TranscriptEvent | Stable, final hypothesis |
speech_start | — | First non-empty word in an utterance |
speech_end | — | Deepgram speech_final flag raised |
error | Error | Unrecoverable provider error |
close | — | Session fully terminated |