Skip to main content

foundry-hosted-agents

Deploy + manage Foundry hosted agents — container deploy via unified azure.yaml is GA; source-code --deploy-mode code is still preview. MAF 1.14 current, azd ext install microsoft.foundry, implicit agent access (no default role grant). USE FOR: deploy foundry agent, hosted agent, container agent, azure.yaml, azd ai agent, microsoft.foundry, MAF, FoundryChatClient, ResponsesHostServer, ACR push, batch eval, agent identity, Foundry User role, entra-agent-id, Responses/Invocations protocols, Activity protocol, blue-green deploy, canary rollout, version rollback, traffic routing, version_selector, agent_endpoint, update_details. DO NOT USE FOR: Harness runtime composition (use agent-framework-harness), prompt agents (use foundry-prompt-agents), ACA MCP (use foundry-mcp-aca), GHCP coding agent (use ghcp-hosted-agents), Citadel hub/spoke (use citadel-hub-deploy), pilot pipeline (use threadlight-deploy), continuous eval (use foundry-evals), Routines (use foundry-routines), A2A wiring (use foundry-toolbox).

Jump to install

Source facts

Repository
aiappsgbb/awesome-gbb
Last source activity
September 25, 2026 at 14:00
Detected SKILL.md language
English
Stars
6
Forks
3

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

File Explorer
30 files

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
foundry-hosted-agents
description
Deploy + manage Foundry hosted agents — container deploy via unified azure.yaml is GA; source-code --deploy-mode code is still preview. MAF 1.14 current, azd ext install microsoft.foundry, implicit agent access (no default role grant). USE FOR: deploy foundry agent, hosted agent, container agent, azure.yaml, azd ai agent, microsoft.foundry, MAF, FoundryChatClient, ResponsesHostServer, ACR push, batch eval, agent identity, Foundry User role, entra-agent-id, Responses/Invocations protocols, Activity protocol, blue-green deploy, canary rollout, version rollback, traffic routing, version_selector, agent_endpoint, update_details. DO NOT USE FOR: Harness runtime composition (use agent-framework-harness), prompt agents (use foundry-prompt-agents), ACA MCP (use foundry-mcp-aca), GHCP coding agent (use ghcp-hosted-agents), Citadel hub/spoke (use citadel-hub-deploy), pilot pipeline (use threadlight-deploy), continuous eval (use foundry-evals), Routines (use foundry-routines), A2A wiring (use foundry-toolbox).
metadata
{"version":"3.0.0"}
# Microsoft Foundry Hosted Agents — Reference Guide Production-tested patterns for deploying hosted agents on Microsoft Foundry. **Container-based deployment (Dockerfile + unified `azure.yaml` + `azd`) is GA.** Source-code (`--deploy-mode code`) deployment is a separate, still-preview surface — see [§ Preview appendix: source-code deploy](#preview-appendix-source-code-deploy-still-preview) for that path in isolation. Covers the `Agent` + `FoundryChatClient` + `ResponsesHostServer` (MAF) variant exclusively. > **✅ Core 1.14.0 is the current dependency stack.** Use the canonical pins in > this skill (`agent-framework-core~=1.14.0`, `agent-framework-foundry~=1.11.0`, > `agent-framework-foundry-hosting==1.0.0b260813`) for new hosted-agent work and > refreshes. > > `MAF 1.8.0` matters here only as the historical feature boundary where the > caller-side `agent_framework.foundry.FoundryAgent(timeout=...)` knob landed > (overrides the OpenAI-SDK 5s connect / 600s total defaults) — see > [§ MAF 1.8.0 update](#maf-180-update-june-2026). > > If you are still on 1.3.x, upgrade directly to current 1.14.0 and absorb the > 1.4.0 breaking changes below plus the 1.6.0 telemetry bundling and the 1.8.0 > timeout knob in one pass. ## When to Use - Deploying a custom container agent to Foundry (GA path) - Debugging hosted agent failures (401, 500, import errors) - Understanding agent identity + when (rarely) explicit RBAC is needed - Authoring the unified `azure.yaml` (`azure.ai.project` + `azure.ai.agent` services) - Blue-green / canary / rollback traffic routing across agent versions - **Migrating from MAF 1.3.x → 1.4.0** ([§ below](#maf-140-breaking-changes-may-2026)) - **Migrating from the legacy two-file `agent.yaml` + `agent.manifest.yaml` contract** — select and verify the versioned consumer before migration; the supported unified profile uses a single `azure.yaml` ([§ azure.yaml](#azureyaml-unified-hosted-agent-configuration)) | Need | Use | |---|---| | Build an already-hostable `Agent` with `create_harness_agent` | [`agent-framework-harness`](../agent-framework-harness/SKILL.md) | > **Choosing a model?** See [`references/model-selection.md`](references/model-selection.md) for the model / region / capacity / data-residency decision before you `azd provision`. ## Deployment preflight **Select a versioned contract before authoring.** The authoritative profile is [`references/hosted-contract.json`](references/hosted-contract.json), consumed by [`hosted_contract.py`](references/python/hosted_contract.py) and the stager. It binds the tested consumer, manifest/environment shape, protocol, runtime dependencies, identities and lifecycle; the existing template and pyproject remain the source of their actual content. The recorded private BASIC profile is azd **1.34.1 + azure.ai.agents 1.0.0-beta.14**, not “this version or newer”. Its historical no-tools evidence does not certify other modes or business tools. Legacy two-file manifests are **migration-only**. Preserve them until the owner selects and verifies a migration; do not delete them merely because this skill documents a unified manifest. Current Learn pages use an `env` map while the recorded consumer template uses an `environmentVariables` list. Do not mix the shapes or rewrite a working cohort without checking actual consumer emission. An unknown toolchain is `UNSUPPORTED_CAPABILITY`, not permission to upgrade. Run the new **capabilities** phase of the existing preflight before building or publishing an image; then the setup and execution phases at their own boundaries. See [capability evidence](references/deployment-preflight.md#early-capability-evidence). The capability check is not a new deployment framework and never changes Azure. **Platform-managed compute does not mean that private project setup is optional.** Before hosted registration, select the account/project/network mode and follow the single [deployment preflight and evidence contract](references/deployment-preflight.md). Run its read-only `references/python/deploy_preflight.py` gate on fresh, operator-collected observations. It performs no Azure calls or repairs. | Selected route | Capability-host prerequisite before registration | |---|---| | Public, platform-managed azd setup | Let the selected template/platform own setup; do not add manual capability hosts from legacy instructions. Raw/legacy public provisioning needs its own explicit contract review. | | Private Basic, including raw SDK or brownfield azd | Read the **project** capability-host inventory and exact host GET. Require `Agents` / `Succeeded`, no BYO stores. The canonical Basic template has a project host only; no customer-created account host is required by that template. | | Private Standard/BYO | Verify the account host/subnet and project host with the selected existing BYO connection names. Preserve hosts, connections and stores. | For a missing Basic project host, stop and request explicit authorization to use only the [existing project-host module](../foundry-vnet-deploy/templates/basic-vnet/modules-network-secured/add-project-capability-host.bicep). An unreadable inventory is not absence. A failed, incompatible or BYO host is not permission to recreate it; use [`foundry-caphost-lifecycle`](../foundry-caphost-lifecycle/SKILL.md) to inspect. `azd deploy` against a pre-existing project does not certify these prerequisites. The gate also checks project/model scope, actual operator/runtime-path evidence, the credential-free **project ContainerRegistry connection** and retained native identity mapping, mode-appropriate project-MI ACR pull access/policy, immutable image and the selected runtime cohort. Preserve the default container identity and verify writable `/home/session/.sessions`; a fixed UID can break the native mounted session home. Keep management and runtime environments separate. Do not change pins merely because another working deployment uses a different tested cohort. **Registration, readiness, invocation and business proof are separate.** LIST versions is inventory, not readiness: require two consecutive **direct version GET** observations, then exact endpoint routing, authenticated invocation and the real requested tool result with independent audit. Account `Succeeded`, image build, LIST `active`, health, or a noop are insufficient. Lost create ACK means reconcile existing versions against the frozen definition before any retry. Retain old versions and signed bindings; changed image/version/identity requires a new exact observation and association. A prompt agent is not an automatic substitute. ## Private BASIC consumer For an existing private project/ACR, execute the bounded [private BASIC procedure](references/private-basic.md) and its [opt-in private fixture](test-fixture/private_consumer_prompt.md). This no-tools bootstrap is independent of tool consent and governed business execution; the public CI consumer and private prompt-agent E2E are not substitutes. Use the canonical stager, runtime and native SDK oracle below. Source builds retain lock/base/context provenance and need a network-reachable builder. Prebuilt images use a service-level immutable `image` with native `docker.imagePassthrough: true` / `remoteBuild: false`, plus explicit `azd deploy "$AGENT" --from-package "$IMAGE" --no-prompt` on the live-validated **azd 1.34.1 + azure.ai.agents 1.0.0-beta.14** compatibility baseline (not a minimum-version claim). **azd 1.27.0 is blocked for this private prebuilt route:** it can still republish despite that flag. This remote-image argument is distinct from `azd publish --from-package`, which consumes a local package and pushes it. A matching connection is reused; missing configuration routes to separately authorized native connection-only provisioning, while unreadable or conflicting state blocks without repair. No public-network opening, new grants, global tool update or OCI-format ban is part of this route. **Validated 2026-09-17 ([#500](https://github.com/aiappsgbb/awesome-gbb/pull/500)):** the no-tools private BASIC consumer activated from the existing immutable image, completed one model request and passed independent response/session/version readbacks. Exact owned native objects were subsequently verified absent; retained image/cache and service history have explicit bounded owner custody, not a purge claim. This proves the tested BASIC path, not governed business execution or other runtime/tooling combinations. --- ## MAF 1.4.0 breaking changes (May 2026) > **Historical boundary only.** The pins below are preserved for the > original 1.4.0 cutover. If you are still on 1.3.x today, do **not** copy > these historical 1.4 pins — use the current canonical 1.14 stack in > [`references/python/pyproject.toml`](references/python/pyproject.toml), > which already incorporates the 1.4 breaking changes, the 1.6 telemetry > update, and the 1.8 timeout correction. Azure renamed the Foundry data-plane role from **"Azure AI User"** to **"Foundry User"** and changed the AAD token audience the SDK requests from `https://cognitiveservices.azure.com/.default` to `https://ai.azure.com/.default`. **`agent-framework-core` 1.4.0** (2026-05-14, alongside `agent-framework-foundry-hosting` 1.0.0a260514) is the first SDK that requests the new scope; everything pinned to 1.3.x or earlier requests the old scope and gets **`401 Unauthorized`** from the post-rename data plane. ### The three changes you have to absorb | # | Change | Symptom on pinned 1.3.x | Fix | |---|--------|------------------------|-----| | 1 | **AAD token scope:** `cognitiveservices.azure.com` → `ai.azure.com` | Every Responses request → `401 Unauthorized` (with valid `Foundry User` RBAC) | Upgrade to `agent-framework-core~=1.4.0` + `agent-framework-foundry~=1.4.0` + `agent-framework-foundry-hosting==1.0.0a260514` | | 2 | **Role display name** "Azure AI User" → "Foundry User" | `az role assignment create --role "Azure AI User"` fails with `RoleDefinitionNotFound` | Use `--role "Foundry User"` OR pin by GUID `53ca6127-db72-4b80-b1b0-d745d6d5456d` (unchanged across the rename — the safest call-site form is the GUID) | | 3 | **`AzureOpenAIChatClient` removed** from `agent_framework.azure` | `ImportError: cannot import name 'AzureOpenAIChatClient'` after `pip install -U` | Use `OpenAIChatClient(azure_endpoint=..., model=..., credential=...)` from `agent_framework.openai` — see snippet below | ### `AzureOpenAIChatClient` → `OpenAIChatClient` migration `FoundryChatClient` is the right choice for hosted Foundry agents and is unaffected. The removal hits **adjacent code paths** that talked directly to Azure OpenAI (eval judges, batch scoring, sidecar services, agents that route through APIM): ```python # OLD (MAF 1.3.x — REMOVED in 1.4.0) from agent_framework.azure import AzureOpenAIChatClient from azure.identity import get_bearer_token_provider, DefaultAzureCredential client = AzureOpenAIChatClient( endpoint=AZURE_OPENAI_ENDPOINT, deployment_name=DEPLOYMENT, ad_token_provider=get_bearer_token_provider( DefaultAzureCredential(), "https://cognitiveservices.azure.com/.default", ), ) # NEW (MAF 1.4.0) from agent_framework.openai import OpenAIChatClient from azure.identity import DefaultAzureCredential client = OpenAIChatClient( azure_endpoint=AZURE_OPENAI_ENDPOINT, # NB: kwarg renamed from `endpoint` model=DEPLOYMENT, # NB: kwarg renamed from `deployment_name` credential=DefaultAzureCredential(), # SDK derives the right token scope internally ) ``` Drop the explicit `get_bearer_token_provider` / `ad_token_provider` — `OpenAIChatClient` now derives the scope itself. The unused `get_bearer_token_provider` import can be removed. ### Image-tag staleness trap (mandatory read for any pilot that pins by digest) Hosted agent versions reference the orchestrator container by ACR image **reference**. There are two common shapes: - `<acr>.azurecr.io/tl-maf-orchestrator:latest` — re-resolves to the newest pushed image on every container start. Auto-picks up MAF rebuilds the next time the agent provisions. - `<acr>.azurecr.io/tl-maf-orchestrator@sha256:abc…` — pinned to a specific registry manifest/index digest (not a local image configuration ID or an individual layer digest). Reproducible, but **frozen at the MAF version that was in the image when the digest was computed**. After a MAF 1.4.0 rebuild, agents using `:latest` work; agents using `@sha256:…` digests from the 1.3.x era keep hitting the old token scope and fail. **Re-import every hosted agent version** (or pin by digest to the freshly-built 1.4.0 image) as part of the upgrade. `azd ai agent deploy` does this automatically if the YAML references `:latest`. ### Rebuild recipe (copy-paste) ```bash # 1. Upgrade orchestrator pyproject.toml + regenerate uv.lock sed -i.bak 's/"agent-framework-core[^"]*"/"agent-framework-core~=1.4.0"/' src/orchestrator/pyproject.toml sed -i.bak 's/"agent-framework-foundry[^"]*"/"agent-framework-foundry~=1.4.0"/' src/orchestrator/pyproject.toml sed -i.bak 's/"agent-framework-foundry-hosting[^"]*"/"agent-framework-foundry-hosting==1.0.0a260514"/' src/orchestrator/pyproject.toml (cd src/orchestrator && uv lock) # 2. ACR remote build — tag with :latest AND a date-pinned tag for forensics ACR=tl<your-acr-suffix> az acr build --registry "$ACR" \ --image "tl-maf-orchestrator:maf14-$(date +%Y%m%d%H%M)" \ --image "tl-maf-orchestrator:latest" \ --file src/orchestrator/dockerfile src/orchestrator/ # 3. Re-deploy + re-import each agent so it picks up the fresh image azd deploy agents (cd infra/scripts && uv run deploy_job.py) # Then: azd ai agent show — confirm new image digest under each version ``` > **Why date-pinned tags matter.** `:latest` is fine for the agent > reference but leaves zero forensic trail in ACR's image history. > The `maf14-YYYYMMDDHHMM` tag lets you correlate "which agent > version is running which MAF build?" months later when a regression > bisect needs it. ## MAF 1.6.0 update (May 2026) `agent-framework-core` 1.6.0 and `agent-framework-foundry-hosting` 1.0.0a260521 ship with **instrumentation enabled by default**. The hosting package transitively pulls `azure-ai-agentserver-core>=2.0.0b3` which depends on `microsoft-opentelemetry>=1.0.0` (resolves to 1.1.0), bundling all OTel instrumentors: | Bundled instrumentor | What it captures | |---|---| | `opentelemetry-instrumentation-openai-v2==2.3b0` | gen_ai spans: model name, token usage, latency for every OpenAI/Responses call | | `opentelemetry-instrumentation-openai-agents-v2==0.1.0` | Agent-level invocation spans | | `opentelemetry-instrumentation-httpx` | HTTP dependency spans | | `azure-core-tracing-opentelemetry` | Azure SDK call tracing | ### What this means for `container.py` **Remove all manual OTel code.** No `configure_azure_monitor()`, no `OpenAIInstrumentor().instrument()`, no custom `TracerProvider`. The platform handles everything. Your container code only needs: ```python # Ensure env var is set (deploy.py passthrough for O-012 workaround) cs = os.getenv("APPLICATION_INSIGHTS_CONNECTION_STRING") or \ os.getenv("APPLICATIONINSIGHTS_CONNECTION_STRING") if cs: os.environ["APPLICATIONINSIGHTS_CONNECTION_STRING"] = cs # Optionally: SDK-level telemetry setup (after FoundryChatClient creation) try: await client.configure_azure_monitor(enable_sensitive_data=True) except Exception: pass # Agent works without telemetry ``` ### gen_ai spans appear under `cloud_RoleName = "agent_framework"` **Important for KQL queries:** gen_ai dependency spans use `cloud_RoleName = "agent_framework"`, NOT `agent-{jobId}-maf`. Queries filtering `cloud_RoleName startswith "agent-"` will miss them. ```kql // gen_ai spans — model, tokens, latency dependencies | where timestamp > ago(1h) | where cloud_RoleName == "agent_framework" | project timestamp, name, duration, model=tostring(customDimensions['gen_ai.response.model']), input_tokens=tostring(customDimensions['gen_ai.usage.input_tokens']), output_tokens=tostring(customDimensions['gen_ai.usage.output_tokens']), op=tostring(customDimensions['gen_ai.operation.name']) | order by timestamp desc ```
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub