- name
- foundry-voice-live
- description
- Build real-time voice agents on Azure AI Foundry Voice Live (GA 2026-04-10). Four-rung migration ladder: Azure OpenAI Realtime → Voice Live → Voice Live + Foundry Agent → native azure-ai-voicelive SDK. Covers semantic VAD, AEC, Neural HD voices, agent routing triplet, TTFA/TTFT benchmark, Gradio + FastRTC UI, plus 2026-04-10 GA deltas: proactive turn control, MCP tools mid-turn, OpenTelemetry via diagnostic settings, auto-truncate. USE FOR: voice live, realtime voice, voice agent, speech to speech, semantic VAD, Neural HD voices, FastRTC, Gradio voice, TTFA, gpt-realtime, byo wss, voice avatar, azure-ai-voicelive, voice live sdk, voice live mcp, mcp mid-turn, voice live proactive turn, voice live auto-truncate, voice live otel. DO NOT USE FOR: batch STT/TTS (use foundry-doc-vision-speech), non-voice agents (use foundry-hosted-agents / foundry-prompt-agents), App Insights ingestion (use foundry-observability), authoring an MCP server (use foundry-mcp-aca or ui-widget-developer).
- metadata
- {"version":"1.4.0"}
# Foundry Voice Live
Build **real-time voice agents** on Azure AI Foundry using Voice Live —
the GA (2026-04-10) server-side voice pipeline that adds semantic VAD,
echo cancellation, noise reduction, and Azure Neural HD voices on top of
the standard Azure OpenAI Realtime API.
The migration from Realtime to Voice Live is **three small code changes**;
the migration from `openai`-shim to the native `azure-ai-voicelive` SDK
is one additional step (Rung 4). This skill walks through both ladders,
the session config that unlocks Voice Live features, and the four
2026-04-10 GA deltas — proactive turn control, MCP tools mid-turn,
OpenTelemetry via diagnostic settings, and auto-truncate governance.
## When to Use
Use this skill when the task involves ANY of:
- Real-time speech-to-speech interaction with a model
- Migrating from Azure OpenAI Realtime to Voice Live
- Adding voice to a Foundry hosted or prompt agent
- Building a voice demo with Gradio / FastRTC / WebRTC
- Benchmarking realtime voice latency (TTFA, TTFT)
- Comparing Azure Neural HD voices vs OpenAI voices
Do NOT use for batch STT/TTS (`foundry-doc-vision-speech`), non-voice
agents (`foundry-hosted-agents`, `foundry-prompt-agents`), or document
extraction.
---
## 1 · The Four Rungs
The entire migration from "plain Realtime" to "native Voice Live SDK
on a Foundry Agent" is a **diff ladder** — each rung changes only the
connection-setup block.
```
Rung 1: Azure OpenAI Realtime ← the "before"
│
▼ diff = 3 small lines (api_version, websocket_base_url, extra_query)
Rung 2: Azure Voice Live ← the punchline
│
▼ diff = 1 line (extra_query gains agent-id / -project-name / -access-token)
Rung 3: Voice Live + Foundry Agent ← the endgame on `openai` SDK
│
▼ swap `openai.AsyncAzureOpenAI` → `azure.ai.voicelive.aio.connect`
Rung 4: native `azure-ai-voicelive` SDK ← the GA path
```
Everything else — the audio pipe, transcript fan-out, status events,
voice picker, UI — is **identical across all four rungs**.
---
## 2 · Connection Code (Copy-Paste Ready)
### Rung 1 — Azure OpenAI Realtime
```python
from contextlib import asynccontextmanager
from openai import AsyncAzureOpenAI
@asynccontextmanager
async def connect_realtime(*, settings, token_provider):
client = AsyncAzureOpenAI(
azure_endpoint=settings.azure_endpoint, # https://<resource>.openai.azure.com
api_version="2025-04-01-preview", # Realtime preview
azure_ad_token_provider=token_provider,
)
try:
async with client.realtime.connect(
model=settings.azure_deployment_name,
) as conn:
yield conn
finally:
await client.close()
```
### Rung 2 — Azure Voice Live (3 changed lines)
```python
@asynccontextmanager
async def connect_voicelive(*, settings, token_provider):
actual_model = settings.azure_deployment_name
client = AsyncAzureOpenAI(
azure_endpoint=settings.azure_endpoint,
api_version="2026-04-10", # ← GA
azure_ad_token_provider=token_provider,
websocket_base_url=settings.azure_voice_live_endpoint, # ← wss://.../voice-live
)
try:
async with client.realtime.connect(
model=actual_model,
extra_query={"model": actual_model}, # ← &model= not &deployment=
) as conn:
yield conn
finally:
await client.close()
```
### Rung 3 — Voice Live + Foundry Agent (extra_query extends)
```python
@asynccontextmanager
async def connect_agent(*, settings, token_provider, agent_token_provider):
client = AsyncAzureOpenAI(
azure_endpoint=settings.azure_endpoint,
api_version="2026-04-10",
azure_ad_token_provider=token_provider,
websocket_base_url=settings.azure_voice_live_endpoint,
)
try:
async with client.realtime.connect(
model=settings.azure_deployment_name,
extra_query={
"agent-id": settings.agent_id,
"agent-project-name": settings.agent_project_name,
"agent-access-token": await agent_token_provider(), # ai.azure.com scope
},
) as conn:
yield conn
finally:
await client.close()
```
### Why 3 lines, not 1
1. **`websocket_base_url`** — the headline change; redirects the WSS
connection to `/voice-live` on `services.ai.azure.com`.
2. **`api_version`** — Realtime is still on `2025-04-01-preview`
(the `openai 2.x` SDK emits `/openai/realtime`; when it adopts the
GA `/openai/v1/realtime` URL this difference collapses). Voice Live
is GA — the latest stable version is `2026-04-10` (was `2025-10-01`
pre-//build 2026).
3. **`extra_query={"model": ...}`** — the SDK adds `&deployment=…` to
the WSS URL by default; Voice Live keys off `&model=…`, so we add
it explicitly.
### Rung 4 — native `azure-ai-voicelive` SDK
The validated native stack is `azure-ai-voicelive[aiohttp]~=1.3.0`.
SDK 1.3 defaults `connect()` to `2026-07-15`; this skill deliberately
passes `api_version="2026-04-10"` to preserve the live-proven GA
contract; do not remove until a separate `2026-07-15` migration is
tested end-to-end. The SDK speaks the same wire protocol as Rungs 2–3
(so the event-handling code in §9 still works verbatim), but replaces
the `openai`-shim plumbing with a typed, Voice-Live-native client.
```python
from contextlib import asynccontextmanager
from azure.ai.voicelive.aio import connect # async client
from azure.ai.voicelive.models import (
AzureSemanticVad,
AzureStandardVoice,
InputAudioFormat,
Modality,
OutputAudioFormat,
RequestSession,
)
from azure.identity.aio import DefaultAzureCredential
@asynccontextmanager
async def connect_voicelive_sdk(*, settings):
credential = DefaultAzureCredential()
try:
# SDK reshapes https://<resource>.services.ai.azure.com/ →
# wss://<resource>.services.ai.azure.com/voice-live/realtime
# ?api-version=2026-04-10&model=<model>
async with connect(
credential=credential,
endpoint=settings.azure_voice_live_endpoint, # https://, NOT wss://
api_version="2026-04-10", # live-proven GA contract
model=settings.azure_deployment_name, # e.g. "gpt-realtime"
# credential_scopes default = ["https://ai.azure.com/.default"]
) as conn:
await conn.session.update(session=RequestSession(
modalities=[Modality.TEXT, Modality.AUDIO],
instructions="You are a friendly assistant.",
voice=AzureStandardVoice(name="en-US-Ava:DragonHDLatestNeural"),
input_audio_format=InputAudioFormat.PCM16,
output_audio_format=OutputAudioFormat.PCM16,
turn_detection=AzureSemanticVad(
create_response=True, # §12.1 proactive
auto_truncate=True, # §12.4 token governance
),
))
yield conn
finally:
await credential.close()
```
### Why move to Rung 4
| Concern | Rungs 1–3 (`openai` shim) | Rung 4 (`azure-ai-voicelive`) |
|---------|---------------------------|-------------------------------|
| Typed session config | dict literals | `RequestSession` + typed models |
| Endpoint shape | `wss://…/voice-live` + base override | `https://…services.ai.azure.com` (SDK derives) |
| Auth scope default | manual `https://ai.azure.com/.default` | SDK default `https://ai.azure.com/.default` |
| MCP tools mid-turn | manual JSON | `MCPServer` + `MCPTool` typed |
| Avatar / custom voice | manual JSON | `AvatarConfig`, `AzureCustomVoice` |
| Interim response | not exposed | `LlmInterimResponseConfig` |
| API version pin | env var | SDK constant (override via kwarg) |
The native SDK is the **recommended path for new code** as of GA
`2026-04-10`. Migrate Rungs 1–3 incrementally — the wire protocol is
identical, so the audio pipe and event handler can stay as-is.
### Endpoint hostname (services.ai vs cognitiveservices)
Voice Live lives on the `services.ai.azure.com` subdomain of your
Foundry resource — the SAME resource that serves chat models on
`cognitiveservices.azure.com`. Map your CI/prod env var like:
```python
endpoint = os.environ["AZURE_AI_ENDPOINT"].replace(
"cognitiveservices.azure.com", "services.ai.azure.com"
)
# Or set AZURE_VOICELIVE_ENDPOINT directly to the services.ai host.
```
> **Install:** `pip install "azure-ai-voicelive[aiohttp]~=1.3.0"`.
> The `[aiohttp]` extra is **required** for the async `connect`
> path — without it the import raises `ImportError: aiohttp is
> required for azure-ai-voicelive`.
---
## 3 · Session Configuration
Voice Live sessions expose capabilities that plain Realtime doesn't.
Send these in `conn.session.update(session={...})` after connection.
### Realtime session (Rung 1)
```python
{
"turn_detection": {"type": "server_vad"},
"input_audio_format": "pcm16",
"output_audio_format": "pcm16",
"voice": "alloy", # OpenAI voice set only
"instructions": "You are a friendly assistant.",
"modalities": ["text", "audio"],
"input_audio_transcription": {
"model": "whisper-1",
"language": "en",
},
}
```
### Voice Live session (Rung 2)
```python
{
"turn_detection": {"type": "azure_semantic_vad", "remove_filler_words": False},
"input_audio_format": "pcm16",
"output_audio_format": "pcm16",
"voice": {"name": "en-US-Ava:DragonHDLatestNeural", "type": "azure-standard"},
"instructions": "You are a friendly assistant.",
"modalities": ["text", "audio"],
"input_audio_echo_cancellation": {"type": "server_echo_cancellation"},
"input_audio_noise_reduction": {"type": "azure_deep_noise_suppression"},
"input_audio_transcription": {
"model": "azure-fast-transcription", # faster than whisper-1
"language": "en",
},
}
```
### Voice Live + Agent session (Rung 3)
The agent owns instructions and tools — omit `instructions` from the
session config:
```python
{
"turn_detection": {"type": "azure_semantic_vad", "remove_filler_words": False},
"input_audio_format": "pcm16",
"output_audio_format": "pcm16",
"voice": {"name": "en-US-Ava:DragonHDLatestNeural", "type": "azure-standard"},
"modalities": ["text", "audio"],
"input_audio_echo_cancellation": {"type": "server_echo_cancellation"},
"input_audio_noise_reduction": {"type": "azure_deep_noise_suppression"},
"input_audio_transcription": {
"model": "azure-fast-transcription",
"language": "en",
},
}
```
### Non-English semantic VAD
`azure_semantic_vad` is English-tuned. For other languages, the endpointer can
fire **mid-utterance** on natural hesitations ("uhm…", a pause before an ID),
clipping the user. Two fixes:
- Use the **multilingual** VAD variant and a **multilingual end-of-utterance**
detection model for non-English locales.
- `silence_duration_ms` is the lever that **bridges mid-sentence pauses** —
`threshold` and `timeout_ms` alone don't. Raising it tolerates longer pauses at
the cost of a little latency per turn (≈ the extra silence you wait for), so tune
it to the locale's natural pausing, not lower.
### Feature comparison
| Feature | Realtime | Voice Live |
|---------|----------|------------|
| VAD | `server_vad` | `azure_semantic_vad` (understands pauses vs hesitation) |
| Echo cancellation | ❌ | `server_echo_cancellation` (built-in AEC) |
| Noise reduction | ❌ | `azure_deep_noise_suppression` |
| Transcription | `whisper-1` | `azure-fast-transcription` (lower latency) |
| Voice set | OpenAI only (10 voices) | Azure Neural HD + OpenAI (per locale) |
| Voice format | bare string `"alloy"` | `{"name": "...", "type": "azure-standard"}` |
| Filler word removal | ❌ | Optional (`remove_filler_words: true`) |
### Cascade vs native realtime — latency facts
A **native realtime** model (`gpt-realtime*`) speaks directly. A **text model on
Voice Live** (`gpt-4o`, `gpt-5*`, …) runs as a **cascade**: the model emits text,
then a managed Azure TTS overlay speaks it. Two consequences worth designing for:
- **First-audio floor.** Audio can't start before the model's first token, so the
cascade has an inherent first-audio floor that no client tuning removes. Native
realtime has no overlay and starts sooner. Pick native realtime when first-audio
latency is the priority; pick the cascade when you need a specific text model's
reasoning or a managed model with no deployment.
在 GitHub 查看