Skip to main content

foundry-doc-vision-speech

Add vision (gpt-5.4 family), Document Intelligence v4, and Azure Speech (STT/TTS) capabilities to a Foundry hosted agent. Covers tool contracts, Bicep modules, runtime client code, and model selection for vision workloads. USE FOR: vision tool, image analysis, damage photo analysis, blueprint annotation, document extraction, structured doc parsing, OCR, voice intake, FNOL voice claim, STT, TTS, gpt-5.4-mini vision, gpt-5.4 vision, Document Intelligence prebuilt, custom doc model, Azure Speech, transcription. DO NOT USE FOR: deploying the agent itself (use threadlight-deploy), MCP server deployment (use foundry-mcp-aca), Foundry IQ knowledge retrieval (use foundry-iq), real-time voice / Voice Live (use foundry-voice-live).

Ir a la instalación

Datos de origen

Repositorio
aiappsgbb/awesome-gbb
Última actividad en el origen
25 de septiembre de 2026 a las 14:00
Idioma detectado de SKILL.md
inglés
Estrellas
6
Forks
3

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Explorador de archivos
3 archivos

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
foundry-doc-vision-speech
description
Add vision (gpt-5.4 family), Document Intelligence v4, and Azure Speech (STT/TTS) capabilities to a Foundry hosted agent. Covers tool contracts, Bicep modules, runtime client code, and model selection for vision workloads. USE FOR: vision tool, image analysis, damage photo analysis, blueprint annotation, document extraction, structured doc parsing, OCR, voice intake, FNOL voice claim, STT, TTS, gpt-5.4-mini vision, gpt-5.4 vision, Document Intelligence prebuilt, custom doc model, Azure Speech, transcription. DO NOT USE FOR: deploying the agent itself (use threadlight-deploy), MCP server deployment (use foundry-mcp-aca), Foundry IQ knowledge retrieval (use foundry-iq), real-time voice / Voice Live (use foundry-voice-live).
metadata
{"version":"1.2.10"}
# Foundry Doc / Vision / Speech This skill wires the **non-chat AI services** every threadlight process eventually needs — image understanding, document parsing, voice — into the agent runtime. It reads SPEC § 7b (AI Services & Model Selection) from `threadlight-design`, selects the right Azure resource per modality, generates the tool contracts, and emits the Bicep module selectors that `threadlight-deploy` Phase 6 + `azd-patterns` consume. ## When to Use Use this skill when the SPEC declares ANY of: - Image / photo / video frame analysis (returns triage damage photos, KYC selfie, blueprint annotation, supplier facility photos) - Structured-document extraction (invoices, claim forms, KYC IDs, quotes, BOMs) - Voice / audio (FNOL voice intake, call recording transcription, IVR responses) If the SPEC only needs chat (text-in, text-out), you do NOT need this skill — the default `gpt-5.4-mini` model from `threadlight-deploy` is sufficient. ## When NOT to Use - **Pure chat agents** — `threadlight-deploy` handles the chat model - **Knowledge retrieval over documents** — that's `foundry-iq` (semantic search over already-indexed content). Use this skill for the **extraction step that produces searchable text from unstructured docs**, then hand off to `foundry-iq` for retrieval. - **MCP server deployment** — that's `foundry-mcp-aca` --- ## ⚠️ GPT-4o is LEGACY (May 2026) GPT-4o and GPT-4o Vision are explicitly **forbidden** as defaults. SPEC § 7b hardcodes the modern decision tree: | Use case | Model | Notes | |----------|-------|-------| | Default chat reasoning | `gpt-5.4-mini` (2026-03-17) | 400K context, vision-capable, fastest, cheapest | | Vision (default) | `gpt-5.4-mini` | Multimodal in same model — usually no separate vision deployment | | Vision (high-stakes, large frames) | `gpt-5.4` (2026-03-05) | 1M context, more accurate on dense visual content | | Vision + reasoning premium | `gpt-5.4-pro` | When the answer must reason about the image (e.g., "is this a structural defect?") | | Bulk/cheap vision | `gpt-5.4-nano` | Background batch jobs, throwaway analysis | | Code multimodal | `gpt-5.3-codex` | Screenshot → code, error UI → diagnosis | | Structured docs | Document Intelligence v4 prebuilt | Field extraction with confidence scores; faster than vision LLM for known forms | | Custom forms | Document Intelligence v4 custom | Train on 5+ samples of customer-specific forms | | Voice → text | Azure Speech-to-Text (Whisper option) | Fast, batch, real-time options | | Text → voice | Azure Speech-to-Text (neural voices) | For IVR, audio briefings | If anyone proposes `gpt-4o`, push back: it doesn't support encrypted content, has a 128K context vs 400K-1M, and is on a deprecation glide path. There is no business case for new processes on legacy models. --- ## Decision Tree: Vision vs Document Intelligence vs Speech ``` What kind of unstructured input? ├── IMAGE │ ├── Known structured form (invoice, ID, claim form, BOM page) │ │ → Document Intelligence (prebuilt or custom) │ ├── Free-form photo (damage, facility, blueprint, screenshot, selfie) │ │ → gpt-5.4-mini vision (or gpt-5.4 if dense / high-stakes) │ └── Bulk batch (1000s of images, low-stakes) │ → gpt-5.4-nano vision │ ├── DOCUMENT (PDF / DOCX / scanned) │ ├── Structured (forms, tables, key-value) │ │ → Document Intelligence v4 prebuilt │ ├── Customer-specific (their proprietary form) │ │ → Document Intelligence v4 custom (train on 5+ samples) │ └── Free-form / mixed → render page-by-page + gpt-5.4 vision │ └── AUDIO ├── Real-time stream (live call, IVR turn) → Azure Speech-to-Text streaming ├── Recorded file (FNOL voicemail, recording) │ → Azure Speech-to-Text batch (or Whisper-equivalent endpoint) └── TTS for response → Azure Speech neural voices ``` --- ## Pattern 0 — Foundry Toolbox (recommended starting point, May 2026) Before wiring SDK calls per modality, **check whether a Foundry Toolbox can do the job**. The Foundry Toolbox is a server-side, versioned bundle of hosted tool configurations curated in the Foundry portal and exposed as a single MCP-compatible endpoint. As of May 2026, the catalog includes **Speech, Document Intelligence, Vision, Content Understanding, Translator, Language, Content Safety, Custom Vision, Azure AI Search**, plus generic hosted tools (web search, code interpreter, file search, image generation). **Why prefer Toolbox over per-modality SDK code:** - **Centralized config**: connection strings, SAS containers, model IDs all live in the portal — no env-var fan-out across containers - **Versioning**: change toolbox config, test with `version="v3"`, promote to default — no agent redeploy - **Auth handled server-side** for the upstream services (Speech / DocIntel account credentials are configured **once** at toolbox creation; the consumer only needs an Entra token to the Toolbox endpoint itself) - **MCP-compatible consumption**: works with any agent runtime (MAF, GHCP SDK via bridge, LangGraph, custom code) - **Same code across processes**: a healthcare KYC and an FSI claim agent can both consume the same `vision_doc_speech_toolbox` **When NOT to use Toolbox** (fall back to direct SDK in Pattern A/B): - You need fine-grained control over a SDK feature not exposed by the toolbox tool wrapper (e.g., DocIntel custom-trained model, Speech batch transcription with diarization, real-time streaming audio frames) - Network-secured Foundry: **Azure Speech MCP doesn't support network-secured Foundry projects** as of May 2026 — you must use direct SDK in Pattern A/B - Toolbox is preview (`azure-ai-projects>=2.1.0`); customer policy may forbid preview-tier features in prod - High-volume batch (e.g., 100k invoices/night) — toolbox introduces a network hop the SDK doesn't ### Pattern 0a — DEFAULT: MCP consumption (works with ALL runtimes) **Always start here.** `MCPStreamableHTTPTool` (or your runtime's equivalent MCP client) talks to the Toolbox endpoint over HTTPS. This is the path the Foundry team officially documents for MAF, GHCP (via bridge), LangGraph, and custom code. **Working sample (MAF — direct, with mandatory workarounds):** ```python import os, asyncio from typing import Any # Use the SYNC azure.identity here — get_bearer_token_provider in the sync # module returns a callable that returns a string when invoked. Using the # .aio variant here returns a coroutine that gets f-string-formatted as # "Bearer <coroutine object _provider at 0x...>" and the server returns 401. from azure.identity import DefaultAzureCredential, get_bearer_token_provider from agent_framework import Agent, MCPStreamableHTTPTool from agent_framework.foundry import FoundryChatClient # 1. Build a refreshing Entra token provider for the Toolbox MCP endpoint. # Scope MUST be https://ai.azure.com/.default — wrong scope = 401. credential = DefaultAzureCredential() # reads AZURE_CLIENT_ID for UAMI token_provider = get_bearer_token_provider(credential, "https://ai.azure.com/.default") # 2. Build the MCP tool. # Endpoint format (note the "toolboxes/.../versions/.../mcp" shape — pin a version): # https://<account>.services.ai.azure.com/api/projects/<project>/toolboxes/<name>/versions/<version>/mcp?api-version=v1 mcp_tool = MCPStreamableHTTPTool( name="vision_doc_speech_mcp", url=os.environ["TOOLBOX_MCP_ENDPOINT"], # GOTCHA: MAF calls header_provider(kwargs) — it MUST accept one positional # arg, not zero. Verified against agent_framework/_mcp.py in 1.x. # Use header_provider (NOT static `headers=`) so tokens refresh on expiry. header_provider=lambda kwargs: {"Authorization": f"Bearer {token_provider()}"}, load_prompts=False, # GOTCHA: Foundry MCP returns 500 on prompts/list ) # 3. GOTCHA: MAF's MCPStreamableHTTPTool._ensure_connected() calls send_ping(), # which the Foundry MCP server rejects with 500. Override to a no-op. async def _no_ping(*args: Any, **kwargs: Any) -> None: return None mcp_tool._ensure_connected = _no_ping # type: ignore[assignment] async def main() -> None: async with mcp_tool, Agent( client=FoundryChatClient(credential=credential, model="gpt-5.4-mini"), name="ClaimsIntake", tools=[mcp_tool], ) as agent: # GOTCHA: Toolbox MCP requires stream=True on tools/call. # MAF Agent.run streams by default — explicitly avoid stream=False overrides. result = await agent.run("Transcribe the call and extract claim details.") print(result.text) asyncio.run(main()) ``` **Working sample (GHCP — via the official MCP bridge):** GHCP rejects dots in tool names; Foundry MCP returns names as `{server_label}.{tool_name}`. The official sample at `https://aka.ms/foundry-toolbox-copilotsdk` provides an MCP bridge that replaces `.` → `_` automatically. Use that bridge — do NOT roll your own. > **GHCP cannot use Pattern 0b (native).** Native toolbox attachment > (`tools=toolbox` directly) is a MAF-specific surface that has no > equivalent in the GitHub Copilot SDK. For GHCP, **Pattern 0a is the > only option** — and you must use the published bridge. **MCP gotcha cheat-sheet** (from official troubleshoot, all confirmed May 2026 — these WILL bite if ignored): | Symptom | Cause | Fix | |---------|-------|-----| | `401` on MCP calls | Wrong/expired token | Scope MUST be `https://ai.azure.com/.default`; use `header_provider` (refreshes), not static `headers` | | `500` on `send_ping()` | Foundry MCP doesn't implement `ping` | Override `MCPStreamableHTTPTool._ensure_connected` to a no-op (see sample) | | `500` on `prompts/list` | Foundry MCP doesn't implement prompts | Pass `load_prompts=False` | | `500` on `tools/call` | Non-streaming not supported | Use `stream=True` (MAF default; explicitly verify if you override) | | `400 Multiple tools without identifiers` | Two unnamed tools of same type in toolbox | Toolbox creator must add `server_label`/`name` to each MCP tool | | Tool name not found / GHCP error | Foundry returns `{server_label}.{tool_name}`; GHCP rejects dots | GHCP: use the bridge that swaps `.` → `_`. MAF: use the dotted name as-is | | Custom env var silently overwritten | Platform reserves `FOUNDRY_*` prefix | Rename your env vars (e.g. `TOOLBOX_MCP_ENDPOINT`, NOT `FOUNDRY_TOOLBOX_ENDPOINT`) | ### Pattern 0b — Native MAF (`MCPStreamableHTTPTool`) Use `MCPStreamableHTTPTool` to connect a MAF agent directly to the toolbox MCP endpoint. This replaces the removed `client.get_toolbox()` convenience helper (removed in MAF 1.3.0). - **GHCP SDK does NOT support this** — use Pattern 0a with the bridge instead - The `allowed_tools` parameter filters to specific tools from the toolbox ```python # MAF 1.3.0+ — MCPStreamableHTTPTool replaces the removed get_toolbox() import os from agent_framework import Agent, MCPStreamableHTTPTool from agent_framework.foundry import FoundryChatClient from azure.identity.aio import AzureCliCredential async with AzureCliCredential() as credential: client = FoundryChatClient(credential=credential) toolbox_tool = MCPStreamableHTTPTool( url=os.environ["TOOLBOX_MCP_ENDPOINT"], load_prompts=False, headers={"Foundry-Features": "Toolboxes=V1Preview"}, allowed_tools=["azure_speech", "document_intelligence"], ) async with Agent(client=client, name="ClaimsIntake", tools=[toolbox_tool]) as agent: result = await agent.run("Transcribe and extract claim details.") ``` If something silently misbehaves (tools listed but never invoked, partial schemas, schema-inference errors), **fall back to Pattern 0a (MCP)** — it exercises the same upstream Toolbox endpoint with strictly fewer SDK abstractions in between. ### Keyless RBAC for Toolbox consumption Everything below assumes **keyless** auth — no API keys are passed anywhere in the threadlight chain (we set `disableLocalAuth: true` on every Cognitive Services resource at provisioning time per `azd-patterns`). | Identity | Role | Scope | Why | |----------|------|-------|-----| | **Consuming agent's UAMI** | `Azure AI User` | Foundry project | Required to mint tokens for `https://ai.azure.com/.default` and read the toolbox MCP endpoint | | **Toolbox creator (one-time, not runtime)** | `Azure AI Project Manager` | Foundry project | Needed to create/update toolbox versions and assign `Azure AI User` to consumers | | **Foundry project's own MI** | `Azure AI User` | Foundry account | Project proxies inference + reads upstream tool credentials configured in toolbox | | **Upstream service identities** (Speech / DocIntel / Vision MI) | per-service role (see Pattern A/B below) | each Cognitive Services resource | Configured ONCE at toolbox creation — consumers never see these | > **Why `Azure AI User` and not `Azure AI Developer`?** `Azure AI Developer` > is scoped to the legacy AML / Foundry-hub world; hosted-agent + toolbox > resources need `Azure AI User` (or higher: `Azure AI Project Manager`). **Toolbox catalog reference**: [Foundry tool catalog](https://learn.microsoft.com/azure/foundry/agents/concepts/tool-catalog) · [Toolbox how-to (Python)](https://learn.microsoft.com/azure/foundry/agents/how-to/tools/toolbox) · [Toolbox troubleshoot](https://learn.microsoft.com/azure/foundry/agents/how-to/tools/toolbox#troubleshoot) · [Hosted-agent permissions](https://learn.microsoft.com/azure/foundry/agents/concepts/hosted-agent-permissions) · [Azure Speech MCP](https://learn.microsoft.com/azure/foundry/agents/how-to/tools/azure-ai-speech) · [GHCP toolbox bridge sample](https://aka.ms/foundry-toolbox-copilotsdk) · [MAF toolbox MCP sample](https://aka.ms/foundry-toolbox-maf) If Toolbox doesn't fit (network-secured project, custom model, high-volume batch), continue with Pattern A/B per modality below. --- ## ⚠️ Keyless / RBAC Matrix (every modality, every pattern) Threadlight pilots are **keyless by default**. Every Cognitive Services resource (`AIServices`, `DocumentIntelligence`, `SpeechServices`) is provisioned with `properties.disableLocalAuth: true` per `azd-patterns`, and runtime auth is via UAMI + Entra token. **No `KEY1` / `subscription_key` appears anywhere in the chain.** When a customer needs to roll back to keyed auth (rare — typically air-gapped lab), they explicitly opt in. ### The matrix | Modality | Identity | Role assignment | Token scope | Notes | |----------|----------|-----------------|-------------|-------| | **Foundry chat (Responses endpoint)** | Agent's UAMI | `Azure AI User` on the **Foundry project** | `https://ai.azure.com/.default` | Project proxies inference using its own MI; this is the threadlight default | | **Direct AOAI account endpoint** (bypassing project) | Agent's UAMI | `Cognitive Services OpenAI User` on the **Foundry account** | `https://cognitiveservices.azure.com/.default` | Only when you must hit the account endpoint — rare | | **Foundry Toolbox via MCP** | Agent's UAMI | `Azure AI User` on Foundry project | `https://ai.azure.com/.default` | Toolbox creator (one-time) needs `Azure AI Project Manager` | | **Vision via Foundry Responses** | Same as chat | Same as chat | Same as chat | Vision is a content-part on the chat call — no extra role |
Ver en GitHub
Este SKILL.md es muy grande, por eso SkillsMP muestra aqui solo la primera seccion. Ver en GitHub