- name
- foundry-mcp-aca
- description
- Deploy custom MCP servers as Azure Container Apps or Azure Functions for use with Foundry hosted agents. Covers Cosmos DB MCPToolKit, Playwright MCP, custom MCP servers, protocol requirements, ACA configuration, and authentication + hardening (ACA built-in auth / Easy Auth, OAuth, managed identity). USE FOR: deploy MCP server, MCP on ACA, Cosmos MCP, Playwright MCP in Foundry, custom MCP server, Azure Functions MCP, MCP ACA deployment, remote MCP endpoint, MCP for hosted agent, connect hosted agent to MCP, secure MCP server, harden MCP server, MCP authentication, MCP OAuth, ACA Easy Auth for MCP. DO NOT USE FOR: deploying the hosted agent itself (use threadlight-deploy), long-running job-backed MCP orchestration or external ACA Job handoff (use foundry-mcp-aca-jobs), local MCP development (use mcp-config.json directly), general Azure deploy.
- metadata
- {"version":"2.0.0"}
> **📦 This skill is for MCP server PRODUCERS (deploying servers to ACA).** If you want to CONSUME an existing MCP server from a Foundry hosted agent, see [foundry-hosted-agents](../foundry-hosted-agents/SKILL.md) § MCP Tools or [foundry-toolbox](../foundry-toolbox/SKILL.md) § Learn MCP. If the MCP server should hand work to an ACA Job, use [foundry-mcp-aca-jobs](../foundry-mcp-aca-jobs/SKILL.md) instead.
# Foundry MCP ACA Deployment
> 🎯 **Scope: PRODUCER-side only.** This skill is for **hosting** MCP servers on
> Azure Container Apps or Azure Functions, NOT for **consuming** a remote MCP from
> a Foundry hosted agent (e.g. calling `https://learn.microsoft.com/api/mcp` or
> `https://api.github.com/mcp` from your agent's `container.py`).
>
> - **Producer (this skill):** you're WRITING + DEPLOYING an MCP server.
> Cosmos MCP, Playwright MCP, custom MCP for an internal API, etc.
> - **Consumer (different skill):** you're WRITING a Foundry hosted agent that
> CALLS a remote MCP server. Use [`foundry-hosted-agents`](../foundry-hosted-agents/SKILL.md)
> § MCP Tools via FoundryChatClient + § MCP Tools — recommended pattern.
>
> Confusing the two costs hours. The consumer-side pattern is `MCPStreamableHTTPTool`
> wired into `Agent(tools=[…])`; the producer-side is everything below.
> ⚠️ **Azure Tenant Isolation (mandatory).** Before any `azd` or `az`
> operation, verify tenant isolation per
> [`azure-tenant-isolation`](../azure-tenant-isolation/): set config dirs,
> check token, assert subscription. See that skill's Agent preflight.
Deploy custom MCP servers as **Azure Container Apps** or **Azure Functions** for
use with Foundry hosted agents. The hosted agent container connects to these MCP
servers via HTTP at runtime using `client.get_mcp_tool()`.
## When to Use
- Deploying a Cosmos DB MCP server for agent data access
- Running Playwright/browser automation as a remote MCP server
- Creating a custom MCP server for an API or data store not covered by Foundry built-ins
- Deploying an MCP server as an Azure Function (consumption billing)
## Capability and identity preflight
Before preparing an image, record the selected ACA environment type, region,
supported ingress/auth features, registry authorization mode and exact
API/SDK cohort from this skill's upstream pin. An unsupported feature is
`UNSUPPORTED_CAPABILITY`; a missing owner choice is a blocker, not permission
to change the environment. Do not use the hosted agent's azd profile as proof
of compatibility for an ACA service.
Reuse the [hosted capability/evidence vocabulary](../foundry-hosted-agents/references/deployment-preflight.md#early-capability-evidence)
for the producer's read-only review. The producer has distinct subjects:
deployer, ACA image-pull identity, inbound caller, server MI accessing the
source, and independent result reader. For each, record the actual target,
required data actions, effective roles/conditions and API version.
An RG-wide elevated-role audit is not proof of these permissions.
**Registry feasibility precedes build.** Repository-only access requires an
ABAC registry, effective repository conditions and exclusion of broader pull
grants. Legacy `AcrPull` is registry-wide. Do not grant a broad role, change a
shared registry's mode or treat a successful deployer push as runtime pull.
The existing ACA module declares an identity; it does not prove that grant.
**Ingress trust is deployment-specific.** Easy Auth must enforce authentication
on the actual route, replace caller-supplied principal headers, and leave no
direct backend bypass. Socket peer, forwarded client IP and authenticated
principal are different facts. Ignore arbitrary forwarded headers; trust only
the explicitly verified proxy chain. No wildcard trusted-proxy shortcut.
Private routing, authConfig support and effective config must be verified on
the selected environment before depending on them.
Follow the [operation recovery contract](../foundry-hosted-agents/references/operation-recovery.md)
for stateful tools and delivery adapters. A tool error, generic404 or missing
browser assertion is not permission to repeat a write. Capture the original
operation identity/status before parsing and use the declared result reader.
## Architecture
```
┌────────────────────────┐ HTTPS ┌─────────────────────┐
│ Hosted Agent Container │ ─────────────► │ MCP ACA │
│ client.get_mcp_tool() │ │ (e.g. Cosmos MCP) │
│ Agent + ResponsesHost │ │ Port 8080 /mcp │
└────────────────────────┘ └─────────────────────┘
```
The hosted agent container:
1. Loads `mcp-config.json` at startup (or `MCP_SERVER_URL` env var)
2. Creates `client.get_mcp_tool(name=..., url=..., approval_mode="never_require")` per server
3. Passes tools to `Agent(tools=[...])` alongside skill-loaded instructions
---
## MCP Protocol Requirements
**Successful JSON-RPC requests return HTTP 200; accepted notifications return HTTP 202** (per [MCP 2025-06-18 transport spec](https://modelcontextprotocol.io/specification/2025-06-18/basic/transports)).
Authentication, transport and invalid-session failures retain their meaningful
HTTP status; never turn them into a success-shaped response.
Failing to handle any of these causes `FoundryChatClient.get_mcp_tool()` to silently fail.
| Method | Purpose | Notes |
|--------|---------|-------|
| `initialize` | Protocol handshake | Must return server capabilities |
| `notifications/initialized` | Client notification | HTTP 202 Accepted (no body) |
| `tools/list` | Discover available tools | Must return tool definitions |
| `prompts/list` | List prompts | Required by agent-framework (return empty list) |
| `resources/list` | List resources | Required by agent-framework (return empty list) |
| `logging/setLevel` | Set log level | Per [MCP spec § Logging](https://modelcontextprotocol.io/specification/2025-06-18/server/utilities/logging) — camelCase `setLevel` (capital L). Lowercase `setlevel` returns `-32601 Method not found` from spec-compliant clients. |
**Transport requirements:**
- Foundry only accepts **remote HTTP** MCP endpoints (no stdio, no local)
- Use Streamable HTTP transport (HTTP POST with JSON-RPC at `/mcp`)
- Non-streaming tool call timeout: **100 seconds**
- Private MCP (VNet) requires Standard Agent Setup
- Port 8080 is convention for ACA MCP servers
- Health endpoint at `/health` (separate from MCP protocol)
---
## Option A: Cosmos DB MCPToolKit (.NET)
A pre-built .NET Cosmos DB MCPToolKit image provides 10 tools out of the box:
| Tool | Type | Purpose |
|------|------|---------|
| `list_databases` | Read | List all Cosmos databases |
| `list_collections` | Read | List containers in a database |
| `find_document_by_id` | Read | Get single document by id |
| `text_search` | Read | Text search across documents |
| `query_documents` | Read | SQL query against a container |
| `get_approximate_schema` | Read | Infer schema from sample docs |
| `get_recent_documents` | Read | Get N most recent documents |
| `vector_search` | Read | Semantic vector search |
| `upsert_document` | Write | Create or update a document |
| `delete_document` | Write | Delete a document by id |
### Deployment
Deploy as a per-project ACA. **Prefer keyless (managed identity) over Cosmos keys** —
disable account keys at the Cosmos resource (`disableLocalAuth: true`).
| Variable | Required | Purpose |
|----------|----------|---------|
| `COSMOS_ENDPOINT` | ✅ | Cosmos DB account endpoint |
| `COSMOS_DATABASE` | ✅ | Default database name |
| `AZURE_CLIENT_ID` | ✅ (keyless) | UAMI client ID — the MCPToolKit's Cosmos SDK uses `DefaultAzureCredential` which reads this |
| `COSMOS_AUTH_KEY` | ❌ avoid | Cosmos master key. Only for local dev; for ACA, **disable account keys** at the Cosmos resource (`disableLocalAuth: true`) and grant the UAMI `Cosmos DB Built-in Data Contributor` (data-plane RBAC, NOT control-plane Contributor) on the database scope |
| `DEV_BYPASS_AUTH` | No | Set `true` only for local dev; never in prod |
> **RBAC pin (verified May 2026).** Cosmos DB SQL API data-plane access is
> NOT granted by control-plane roles like `Contributor` or
> `DocumentDB Account Contributor`. You MUST assign the data-plane role
> `Cosmos DB Built-in Data Contributor`
> (`00000000-0000-0000-0000-000000000002`) via `az cosmosdb sql role
> assignment create`. This is the same gotcha called out in
> `threadlight-hitl-patterns` — keep both wirings consistent.
> **⚠️ aiohttp dep is mandatory for the Python Cosmos MCP server.** The Python Cosmos MCPToolKit uses `azure-cosmos` async client which
> silently requires `aiohttp` as the HTTP transport. Without it the container starts
> fine, `tools/list` returns the 11 tools, but every `upsert_item` / `query_items`
> call fails server-side with what looks like a Cosmos error but is actually an
> ImportError swallowed by FastMCP. Pin in `src/mcp/requirements.txt`:
>
> ```
> fastmcp>=2.0.0,<3.0.0 # MUST upper-bound — see callout below
> azure-cosmos>=4.15.0 # see kwarg callout below
> azure-identity>=1.19.0
> mcp>=1.10.0
> aiohttp>=3.9.0 # REQUIRED — async HTTP transport for azure-cosmos
> ```
> **⚠️ Pin `fastmcp<3.0.0` — the unbounded `>=2.0.0` pin is a re-deploy time bomb.**
> FastMCP 3.x changed the streamable-http mount path. If `requirements.txt` says
> `fastmcp>=2.0.0` (no upper bound), the next container rebuild will pull
> **fastmcp 3.x silently** the moment it ships on PyPI. Symptoms — agent says
> *"case read failed"* / *"audit-log screening read failed"* on every tool call,
> bot/MCP logs show every single request as `POST /mcp HTTP/1.1" 404 Not Found`,
> the MCP container itself is `Healthy` and `Running`. **FastMCP itself prints
> the warning at boot:**
>
> ```
> FastMCP 3.0 is coming!
> Pin `fastmcp < 3` in production, then upgrade when you're ready.
> ```
>
> If you skim past that and ship, every Cosmos tool call will 404. A real
> deployment burned an hour on this when an unrelated `azd deploy` rebuild
> jumped fastmcp 2.14.7 → 3.2.4 and broke point reads + queries simultaneously.
> Always upper-bound: `fastmcp>=2.0.0,<3.0.0`. Same rule applies to **any
> client** that imports `fastmcp` (e.g. an ACA Job that drives the MCP) — pin
> client + server to the same major.
>
> **🛑 DO NOT bump `fastmcp` major version without re-running the demo scenarios** — the streamable-http mount path changed between 2.x and 3.x. Every Cosmos tool call will fail silently if the path moves. **DO test locally first:** `pip install 'fastmcp>=3' && python -m pytest tests/ -k cosmos_mcp`.
> **⚠️ `enable_cross_partition_query` was DROPPED in `azure-cosmos>=4.15` async.**
> If your `query_items` tool implementation passes `enable_cross_partition_query=True`,
> the kwarg leaks down to `aiohttp.ClientSession._request()` and raises
> `TypeError: ClientSession._request() got an unexpected keyword argument
> 'enable_cross_partition_query'`. FastMCP swallows the traceback and surfaces
> only `Error calling tool 'query_items'` — the agent then says "case lookup
> failed" on every related read while point reads (`get_item`) keep working.
>
> The new async signature is partition-key-aware by inference:
>
> ```python
> # ❌ Old (works on azure-cosmos<4.15, breaks on >=4.15):
> async for item in container.query_items(
> query=q, parameters=p, enable_cross_partition_query=True,
> ):
> ...
>
> # ✅ New: omit partition_key for cross-partition; pass it for single-partition:
> kwargs = {"query": q, "parameters": p}
> if partition_key is not None:
> kwargs["partition_key"] = partition_key
> async for item in container.query_items(**kwargs):
> ...
> ```
>
> This is a **silent migration trap** — the SDK dependency floats forward, the
> kwarg used to be valid, and the runtime error message blames the tool name
> not the SDK call. Catches every Cosmos MCP that pinned `azure-cosmos>=4.7`
> instead of `>=4.15`.
>
> **🛑 DO NOT rely on cross-partition queries for MCP tool calls without explicit user consent** — partition-scoped queries are cheaper and more predictable. **DO scope per-partition or accept the cost & latency impact.** If you must cross-partition, document it as a tool contract in SPEC § 6 so the agent knows to prefer partition-scoped alternatives when available.
> **Historical stale-session observation after MCP redeploy — verify the selected client.**
> FastMCP's streamable-http maintains per-client session state in-memory on the MCP
> container. When you redeploy the MCP server (`azd deploy cosmos-mcp` /
> `azd deploy <mock-mcp-service>` / any new container revision), every session is wiped.
> In the observed older client, the cached `mcp-session-id` from the previous
> initialize handshake kept being sent with `tools/call` without a successful
> re-handshake. In that incident, tool calls returned
> 404 silently, agent self-reports `case read failed` / `audit-log query failed`
> on EVERY tool, MCP container is `Healthy` and `Running`, MCP logs show
> `POST /mcp HTTP/1.1" 404 Not Found` **without** the preceding `new transport
> with session ID: ...` log line that a fresh handshake would produce. External
> probes to `/mcp` with proper Accept headers return `200 OK` — the path is fine,
> the SESSION is gone.
>
> **Distinguishing this from the FastMCP 3.x mount-path 404:**
>
> | Symptom | FastMCP 3.x mount-path | Stale session-id |
> |---|---|---|
> | MCP log line | `POST /mcp HTTP/1.1" 404` (no transport log either way) | `POST /mcp HTTP/1.1" 404` (no `new transport with session ID` log preceding) |
> | External probe with `Accept: application/json, text/event-stream` | `404 Not Found` (path moved) | `200 OK` (path fine) |
> | External probe with stale `mcp-session-id` header | `404 Not Found` (path moved) | `404 {"error":{"code":-32600,"message":"Session not found"}}` |
> | Discriminator before remediation | Verify the actual mounted route and pinned package | Verify the original session error and the client's reinitialization contract |
>
> **Historical recovery sequence, not an automatic repair:**
>
> ```bash
> # After: azd deploy cosmos-mcp (or any MCP service)
> # Only after identifying the stale session and obtaining explicit authorization:
> azd deploy <agent-service-name> # creates a new agent version, fresh compute
>
> # And restart the bot ACA replica so its connection pool is dropped:
> az containerapp revision restart \
> -g <rg> -n <bot-aca-name> \
> --revision $(az containerapp revision list -g <rg> -n <bot-aca-name> \
> --query "[?properties.active] | [0].name" -o tsv)
> ```
>
> **Alternatively** — wait ~15 min idle and the refreshed-preview hosted agent
> auto-deprovisions; the next user message will spin up fresh compute with a
> fresh MCP session. But "wait 15 min" isn't a fix you can put in a runbook.
>
> Client reinitialization support is version-specific. This incident does not
> establish a universal coupled-redeploy requirement or prove that a failed
> business call had no effect.
>
> Before any recovery, distinguish a wrong route, auth failure and an actual
> stale session with a bounded handshake/read-only probe using the same client
> cohort. `azd ai agent show` is a read, not a cache refresh. Do not automatically
> redeploy the agent or restart the bot, and never replay the failed business
> tool to diagnose its session. If reinitialization cannot be proven safe for the
> original operation, preserve UNKNOWN and return the specific blocked decision.
### Cosmos firewall + ACA egress (the trap that wastes 45 min on every fresh deployment)
The single biggest "first-deploy doesn't work" gotcha for ACA→Cosmos:
| Default | What happens | Fix |
|---|---|---|
| Cosmos `publicNetworkAccess: Disabled` (the Azure default) | All ACA→Cosmos traffic returns `Forbidden — public access disabled` | Set `publicNetworkAccess: Enabled` for development (production: use private endpoint) |
| Cosmos `networkAclBypass: None` (default) AND `ipRules: []` | All ACA→Cosmos traffic returns `Forbidden — Request originated from IP <egress-ip> through public internet. This is blocked` | Either: (a) set `networkAclBypass: AzureServices` ⚠️ (see caveat) OR (b) add the ACA environment's egress IP to `ipRules` (proven path for non-production environments) |
Ver no GitHub