Skip to main content

foundry-hosted-agents

Deploy + manage Foundry hosted agents — container deploy via unified azure.yaml is GA; source-code --deploy-mode code is still preview. MAF 1.14 current, azd ext install microsoft.foundry, implicit agent access (no default role grant). USE FOR: deploy foundry agent, hosted agent, container agent, azure.yaml, azd ai agent, microsoft.foundry, MAF, FoundryChatClient, ResponsesHostServer, ACR push, batch eval, agent identity, Foundry User role, entra-agent-id, Responses/Invocations protocols, Activity protocol, blue-green deploy, canary rollout, version rollback, traffic routing, version_selector, agent_endpoint, update_details. DO NOT USE FOR: Harness runtime composition (use agent-framework-harness), prompt agents (use foundry-prompt-agents), ACA MCP (use foundry-mcp-aca), GHCP coding agent (use ghcp-hosted-agents), Citadel hub/spoke (use citadel-hub-deploy), pilot pipeline (use threadlight-deploy), continuous eval (use foundry-evals), Routines (use foundry-routines), A2A wiring (use foundry-toolbox).

Zur Installation springen

Quellinformationen

Repository
aiappsgbb/awesome-gbb
Letzte Quellaktivität
25. September 2026 um 14:00
Erkannte Sprache von SKILL.md
Englisch
Sterne
6
Forks
3

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

Datei-Explorer
30 Dateien

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
foundry-hosted-agents
description
Deploy + manage Foundry hosted agents — container deploy via unified azure.yaml is GA; source-code --deploy-mode code is still preview. MAF 1.14 current, azd ext install microsoft.foundry, implicit agent access (no default role grant). USE FOR: deploy foundry agent, hosted agent, container agent, azure.yaml, azd ai agent, microsoft.foundry, MAF, FoundryChatClient, ResponsesHostServer, ACR push, batch eval, agent identity, Foundry User role, entra-agent-id, Responses/Invocations protocols, Activity protocol, blue-green deploy, canary rollout, version rollback, traffic routing, version_selector, agent_endpoint, update_details. DO NOT USE FOR: Harness runtime composition (use agent-framework-harness), prompt agents (use foundry-prompt-agents), ACA MCP (use foundry-mcp-aca), GHCP coding agent (use ghcp-hosted-agents), Citadel hub/spoke (use citadel-hub-deploy), pilot pipeline (use threadlight-deploy), continuous eval (use foundry-evals), Routines (use foundry-routines), A2A wiring (use foundry-toolbox).
metadata
{"version":"3.0.0"}
# Microsoft Foundry Hosted Agents — Reference Guide Production-tested patterns for deploying hosted agents on Microsoft Foundry. **Container-based deployment (Dockerfile + unified `azure.yaml` + `azd`) is GA.** Source-code (`--deploy-mode code`) deployment is a separate, still-preview surface — see [§ Preview appendix: source-code deploy](#preview-appendix-source-code-deploy-still-preview) for that path in isolation. Covers the `Agent` + `FoundryChatClient` + `ResponsesHostServer` (MAF) variant exclusively. > **✅ Core 1.14.0 is the current dependency stack.** Use the canonical pins in > this skill (`agent-framework-core~=1.14.0`, `agent-framework-foundry~=1.11.0`, > `agent-framework-foundry-hosting==1.0.0b260813`) for new hosted-agent work and > refreshes. > > `MAF 1.8.0` matters here only as the historical feature boundary where the > caller-side `agent_framework.foundry.FoundryAgent(timeout=...)` knob landed > (overrides the OpenAI-SDK 5s connect / 600s total defaults) — see > [§ MAF 1.8.0 update](#maf-180-update-june-2026). > > If you are still on 1.3.x, upgrade directly to current 1.14.0 and absorb the > 1.4.0 breaking changes below plus the 1.6.0 telemetry bundling and the 1.8.0 > timeout knob in one pass. ## When to Use - Deploying a custom container agent to Foundry (GA path) - Debugging hosted agent failures (401, 500, import errors) - Understanding agent identity + when (rarely) explicit RBAC is needed - Authoring the unified `azure.yaml` (`azure.ai.project` + `azure.ai.agent` services) - Blue-green / canary / rollback traffic routing across agent versions - **Migrating from MAF 1.3.x → 1.4.0** ([§ below](#maf-140-breaking-changes-may-2026)) - **Migrating from the legacy two-file `agent.yaml` + `agent.manifest.yaml` contract** — select and verify the versioned consumer before migration; the supported unified profile uses a single `azure.yaml` ([§ azure.yaml](#azureyaml-unified-hosted-agent-configuration)) | Need | Use | |---|---| | Build an already-hostable `Agent` with `create_harness_agent` | [`agent-framework-harness`](../agent-framework-harness/SKILL.md) | > **Choosing a model?** See [`references/model-selection.md`](references/model-selection.md) for the model / region / capacity / data-residency decision before you `azd provision`. ## Deployment preflight **Select a versioned contract before authoring.** The authoritative profile is [`references/hosted-contract.json`](references/hosted-contract.json), consumed by [`hosted_contract.py`](references/python/hosted_contract.py) and the stager. It binds the tested consumer, manifest/environment shape, protocol, runtime dependencies, identities and lifecycle; the existing template and pyproject remain the source of their actual content. The recorded private BASIC profile is azd **1.34.1 + azure.ai.agents 1.0.0-beta.14**, not “this version or newer”. Its historical no-tools evidence does not certify other modes or business tools. Legacy two-file manifests are **migration-only**. Preserve them until the owner selects and verifies a migration; do not delete them merely because this skill documents a unified manifest. Current Learn pages use an `env` map while the recorded consumer template uses an `environmentVariables` list. Do not mix the shapes or rewrite a working cohort without checking actual consumer emission. An unknown toolchain is `UNSUPPORTED_CAPABILITY`, not permission to upgrade. Run the new **capabilities** phase of the existing preflight before building or publishing an image; then the setup and execution phases at their own boundaries. See [capability evidence](references/deployment-preflight.md#early-capability-evidence). The capability check is not a new deployment framework and never changes Azure. **Platform-managed compute does not mean that private project setup is optional.** Before hosted registration, select the account/project/network mode and follow the single [deployment preflight and evidence contract](references/deployment-preflight.md). Run its read-only `references/python/deploy_preflight.py` gate on fresh, operator-collected observations. It performs no Azure calls or repairs. | Selected route | Capability-host prerequisite before registration | |---|---| | Public, platform-managed azd setup | Let the selected template/platform own setup; do not add manual capability hosts from legacy instructions. Raw/legacy public provisioning needs its own explicit contract review. | | Private Basic, including raw SDK or brownfield azd | Read the **project** capability-host inventory and exact host GET. Require `Agents` / `Succeeded`, no BYO stores. The canonical Basic template has a project host only; no customer-created account host is required by that template. | | Private Standard/BYO | Verify the account host/subnet and project host with the selected existing BYO connection names. Preserve hosts, connections and stores. | For a missing Basic project host, stop and request explicit authorization to use only the [existing project-host module](../foundry-vnet-deploy/templates/basic-vnet/modules-network-secured/add-project-capability-host.bicep). An unreadable inventory is not absence. A failed, incompatible or BYO host is not permission to recreate it; use [`foundry-caphost-lifecycle`](../foundry-caphost-lifecycle/SKILL.md) to inspect. `azd deploy` against a pre-existing project does not certify these prerequisites. The gate also checks project/model scope, actual operator/runtime-path evidence, the credential-free **project ContainerRegistry connection** and retained native identity mapping, mode-appropriate project-MI ACR pull access/policy, immutable image and the selected runtime cohort. Preserve the default container identity and verify writable `/home/session/.sessions`; a fixed UID can break the native mounted session home. Keep management and runtime environments separate. Do not change pins merely because another working deployment uses a different tested cohort. **Registration, readiness, invocation and business proof are separate.** LIST versions is inventory, not readiness: require two consecutive **direct version GET** observations, then exact endpoint routing, authenticated invocation and the real requested tool result with independent audit. Account `Succeeded`, image build, LIST `active`, health, or a noop are insufficient. Lost create ACK means reconcile existing versions against the frozen definition before any retry. Retain old versions and signed bindings; changed image/version/identity requires a new exact observation and association. A prompt agent is not an automatic substitute. ## Private BASIC consumer For an existing private project/ACR, execute the bounded [private BASIC procedure](references/private-basic.md) and its [opt-in private fixture](test-fixture/private_consumer_prompt.md). This no-tools bootstrap is independent of tool consent and governed business execution; the public CI consumer and private prompt-agent E2E are not substitutes. Use the canonical stager, runtime and native SDK oracle below. Source builds retain lock/base/context provenance and need a network-reachable builder. Prebuilt images use a service-level immutable `image` with native `docker.imagePassthrough: true` / `remoteBuild: false`, plus explicit `azd deploy "$AGENT" --from-package "$IMAGE" --no-prompt` on the live-validated **azd 1.34.1 + azure.ai.agents 1.0.0-beta.14** compatibility baseline (not a minimum-version claim). **azd 1.27.0 is blocked for this private prebuilt route:** it can still republish despite that flag. This remote-image argument is distinct from `azd publish --from-package`, which consumes a local package and pushes it. A matching connection is reused; missing configuration routes to separately authorized native connection-only provisioning, while unreadable or conflicting state blocks without repair. No public-network opening, new grants, global tool update or OCI-format ban is part of this route. **Validated 2026-09-17 ([#500](https://github.com/aiappsgbb/awesome-gbb/pull/500)):** the no-tools private BASIC consumer activated from the existing immutable image, completed one model request and passed independent response/session/version readbacks. Exact owned native objects were subsequently verified absent; retained image/cache and service history have explicit bounded owner custody, not a purge claim. This proves the tested BASIC path, not governed business execution or other runtime/tooling combinations. --- ## MAF 1.4.0 breaking changes (May 2026) > **Historical boundary only.** The pins below are preserved for the > original 1.4.0 cutover. If you are still on 1.3.x today, do **not** copy > these historical 1.4 pins — use the current canonical 1.14 stack in > [`references/python/pyproject.toml`](references/python/pyproject.toml), > which already incorporates the 1.4 breaking changes, the 1.6 telemetry > update, and the 1.8 timeout correction. Azure renamed the Foundry data-plane role from **"Azure AI User"** to **"Foundry User"** and changed the AAD token audience the SDK requests from `https://cognitiveservices.azure.com/.default` to `https://ai.azure.com/.default`. **`agent-framework-core` 1.4.0** (2026-05-14, alongside `agent-framework-foundry-hosting` 1.0.0a260514) is the first SDK that requests the new scope; everything pinned to 1.3.x or earlier requests the old scope and gets **`401 Unauthorized`** from the post-rename data plane. ### The three changes you have to absorb | # | Change | Symptom on pinned 1.3.x | Fix | |---|--------|------------------------|-----| | 1 | **AAD token scope:** `cognitiveservices.azure.com` → `ai.azure.com` | Every Responses request → `401 Unauthorized` (with valid `Foundry User` RBAC) | Upgrade to `agent-framework-core~=1.4.0` + `agent-framework-foundry~=1.4.0` + `agent-framework-foundry-hosting==1.0.0a260514` | | 2 | **Role display name** "Azure AI User" → "Foundry User" | `az role assignment create --role "Azure AI User"` fails with `RoleDefinitionNotFound` | Use `--role "Foundry User"` OR pin by GUID `53ca6127-db72-4b80-b1b0-d745d6d5456d` (unchanged across the rename — the safest call-site form is the GUID) | | 3 | **`AzureOpenAIChatClient` removed** from `agent_framework.azure` | `ImportError: cannot import name 'AzureOpenAIChatClient'` after `pip install -U` | Use `OpenAIChatClient(azure_endpoint=..., model=..., credential=...)` from `agent_framework.openai` — see snippet below | ### `AzureOpenAIChatClient` → `OpenAIChatClient` migration `FoundryChatClient` is the right choice for hosted Foundry agents and is unaffected. The removal hits **adjacent code paths** that talked directly to Azure OpenAI (eval judges, batch scoring, sidecar services, agents that route through APIM): ```python # OLD (MAF 1.3.x — REMOVED in 1.4.0) from agent_framework.azure import AzureOpenAIChatClient from azure.identity import get_bearer_token_provider, DefaultAzureCredential client = AzureOpenAIChatClient( endpoint=AZURE_OPENAI_ENDPOINT, deployment_name=DEPLOYMENT, ad_token_provider=get_bearer_token_provider( DefaultAzureCredential(), "https://cognitiveservices.azure.com/.default", ), ) # NEW (MAF 1.4.0) from agent_framework.openai import OpenAIChatClient from azure.identity import DefaultAzureCredential client = OpenAIChatClient( azure_endpoint=AZURE_OPENAI_ENDPOINT, # NB: kwarg renamed from `endpoint` model=DEPLOYMENT, # NB: kwarg renamed from `deployment_name` credential=DefaultAzureCredential(), # SDK derives the right token scope internally ) ``` Drop the explicit `get_bearer_token_provider` / `ad_token_provider` — `OpenAIChatClient` now derives the scope itself. The unused `get_bearer_token_provider` import can be removed. ### Image-tag staleness trap (mandatory read for any pilot that pins by digest) Hosted agent versions reference the orchestrator container by ACR image **reference**. There are two common shapes: - `<acr>.azurecr.io/tl-maf-orchestrator:latest` — re-resolves to the newest pushed image on every container start. Auto-picks up MAF rebuilds the next time the agent provisions. - `<acr>.azurecr.io/tl-maf-orchestrator@sha256:abc…` — pinned to a specific registry manifest/index digest (not a local image configuration ID or an individual layer digest). Reproducible, but **frozen at the MAF version that was in the image when the digest was computed**. After a MAF 1.4.0 rebuild, agents using `:latest` work; agents using `@sha256:…` digests from the 1.3.x era keep hitting the old token scope and fail. **Re-import every hosted agent version** (or pin by digest to the freshly-built 1.4.0 image) as part of the upgrade. `azd ai agent deploy` does this automatically if the YAML references `:latest`. ### Rebuild recipe (copy-paste) ```bash # 1. Upgrade orchestrator pyproject.toml + regenerate uv.lock sed -i.bak 's/"agent-framework-core[^"]*"/"agent-framework-core~=1.4.0"/' src/orchestrator/pyproject.toml sed -i.bak 's/"agent-framework-foundry[^"]*"/"agent-framework-foundry~=1.4.0"/' src/orchestrator/pyproject.toml sed -i.bak 's/"agent-framework-foundry-hosting[^"]*"/"agent-framework-foundry-hosting==1.0.0a260514"/' src/orchestrator/pyproject.toml (cd src/orchestrator && uv lock) # 2. ACR remote build — tag with :latest AND a date-pinned tag for forensics ACR=tl<your-acr-suffix> az acr build --registry "$ACR" \ --image "tl-maf-orchestrator:maf14-$(date +%Y%m%d%H%M)" \ --image "tl-maf-orchestrator:latest" \ --file src/orchestrator/dockerfile src/orchestrator/ # 3. Re-deploy + re-import each agent so it picks up the fresh image azd deploy agents (cd infra/scripts && uv run deploy_job.py) # Then: azd ai agent show — confirm new image digest under each version ``` > **Why date-pinned tags matter.** `:latest` is fine for the agent > reference but leaves zero forensic trail in ACR's image history. > The `maf14-YYYYMMDDHHMM` tag lets you correlate "which agent > version is running which MAF build?" months later when a regression > bisect needs it. ## MAF 1.6.0 update (May 2026) `agent-framework-core` 1.6.0 and `agent-framework-foundry-hosting` 1.0.0a260521 ship with **instrumentation enabled by default**. The hosting package transitively pulls `azure-ai-agentserver-core>=2.0.0b3` which depends on `microsoft-opentelemetry>=1.0.0` (resolves to 1.1.0), bundling all OTel instrumentors: | Bundled instrumentor | What it captures | |---|---| | `opentelemetry-instrumentation-openai-v2==2.3b0` | gen_ai spans: model name, token usage, latency for every OpenAI/Responses call | | `opentelemetry-instrumentation-openai-agents-v2==0.1.0` | Agent-level invocation spans | | `opentelemetry-instrumentation-httpx` | HTTP dependency spans | | `azure-core-tracing-opentelemetry` | Azure SDK call tracing | ### What this means for `container.py` **Remove all manual OTel code.** No `configure_azure_monitor()`, no `OpenAIInstrumentor().instrument()`, no custom `TracerProvider`. The platform handles everything. Your container code only needs: ```python # Ensure env var is set (deploy.py passthrough for O-012 workaround) cs = os.getenv("APPLICATION_INSIGHTS_CONNECTION_STRING") or \ os.getenv("APPLICATIONINSIGHTS_CONNECTION_STRING") if cs: os.environ["APPLICATIONINSIGHTS_CONNECTION_STRING"] = cs # Optionally: SDK-level telemetry setup (after FoundryChatClient creation) try: await client.configure_azure_monitor(enable_sensitive_data=True) except Exception: pass # Agent works without telemetry ``` ### gen_ai spans appear under `cloud_RoleName = "agent_framework"` **Important for KQL queries:** gen_ai dependency spans use `cloud_RoleName = "agent_framework"`, NOT `agent-{jobId}-maf`. Queries filtering `cloud_RoleName startswith "agent-"` will miss them. ```kql // gen_ai spans — model, tokens, latency dependencies | where timestamp > ago(1h) | where cloud_RoleName == "agent_framework" | project timestamp, name, duration, model=tostring(customDimensions['gen_ai.response.model']), input_tokens=tostring(customDimensions['gen_ai.usage.input_tokens']), output_tokens=tostring(customDimensions['gen_ai.usage.output_tokens']), op=tostring(customDimensions['gen_ai.operation.name']) | order by timestamp desc ```
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen