- name
- foundry-observability
- description
- End-to-end observability for Azure AI pilots — App Insights + Log Analytics + OpenTelemetry across hosted agents, ACA MCP servers, ACA jobs, bot service, workspace UIs. Closes the silent telemetry gap where `azd up` returns 0 but **zero traces ever reach App Insights**. Covers Bicep modules, Foundry account-level telemetry connection, ACA-side instrumentation, `Monitoring Metrics Publisher` RBAC, and KQL diagnostic queries. Read the full skill body for the 3-layer wiring sequence — do not instrument from this summary alone. USE FOR: app insights, application insights, OpenTelemetry, OTel, configure_azure_monitor, agent traces missing, no telemetry, blank appin, log analytics, KQL, observability, trace MCP, silent cron, Monitoring Metrics Publisher RBAC, AppInsights connection foundry, account-level appin, AppIn PUT 400, credentials null, silent injection, server_error telemetry. DO NOT USE FOR: continuous eval (foundry-evals), pre-deploy gates (threadlight-safe-check), Foundry IQ monitoring (foundry-iq).
- metadata
- {"version":"1.2.5"}
# Foundry Observability
End-to-end telemetry across every component of a Threadlight pilot:
Foundry hosted agent, MCP servers on ACA, ACA jobs (cron triggers),
bot service, workspace UI. **Default discipline**, not optional.
For the complete per-agent adoption and release-evidence workflow, see
[`foundry-agentops`](../foundry-agentops/SKILL.md). This skill remains
authoritative for OpenTelemetry and App Insights wiring; AgentOps aggregates
evidence and never replaces instrumentation or verification of the telemetry path.
> **Downstream FinOps consumer.** [`foundry-cost-monitoring`](../foundry-cost-monitoring/SKILL.md)
> joins the `gen_ai.usage.*` spans this skill emits with the Azure
> Retail Prices API to compute per-agent / per-project / per-tenant
> cost projection — wire it whenever a FinOps stakeholder needs to
> answer "what is this agent costing us right now?"
> **Why this skill exists.** Recent pilots deployed cleanly
> (`azd up` returned 0, all resources provisioned) but App Insights
> stayed **completely empty** — no agent traces, no MCP tool calls,
> no cron logs. Root cause: no one wired the connection at any layer.
> The intel for *each layer* lives scattered across `threadlight-deploy`,
> `foundry-hosted-agents`, `foundry-mcp-aca`, `threadlight-event-triggers` —
> but no single skill walks an operator through the full chain. That's
> what this skill does. Pair with `threadlight-safe-check` Step 5.6
> (App Insights existence + first-trace probe) to gate it shut.
---
## Mental model — three layers, one signal
```
┌─────────────────────────────────────────────────────────────────────┐
│ Layer 3: ACA workloads (MCP / bot / workspace / cron jobs) │
│ • configure_azure_monitor() reads APPLICATIONINSIGHTS_CONNECTION_STRING │
│ • Env var set by Bicep from app-insights.outputs.connectionString │
│ • OTel exporter ships spans + logs + metrics over HTTPS │
└─────────────────────────────────────────────────────────────────────┘
▲
│ direct push from container code
│
┌─────────────────────────────────┼───────────────────────────────────┐
│ Layer 2: Foundry hosted agent (the runtime) │
│ • Account-level AppInsights connection (category: AppInsights) │
│ • Platform AUTO-INJECTS APPLICATIONINSIGHTS_CONNECTION_STRING │
│ • RBAC: Monitoring Metrics Publisher on agent identities │
│ • Tracing emitted by the runtime — no app code change │
└─────────────────────────────────┼───────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────┐
│ Layer 1: Bicep substrate │
│ • app-insights.bicep — workspace-based (LAW-bound) │
│ • log-analytics.bicep — single LAW for ALL workloads in the RG │
│ • ACA env wiring: dapr.appInsightsConnectionString OR direct env │
│ • Output `connectionString` consumed by every workload │
└─────────────────────────────────────────────────────────────────────┘
```
**Single connection string, three fan-outs.** All telemetry lands in
the same App Insights resource. No per-workload AppIn — that fragments
the trace graph and makes correlation impossible.
---
## What ships from this skill
```
foundry-observability/
├── SKILL.md
└── references/
├── bicep/
│ ├── log-analytics.bicep # LAW (workspace) — required by both AppIn and ACA env
│ ├── app-insights.bicep # AppIn workspace-based + UAMI Monitoring Metrics Publisher RBAC
│ └── aca-env-monitoring.bicep # ACA env wired to LAW + AppIn
├── python/
│ └── otel_init.py # configure_azure_monitor() for ACA workloads
├── postprovision/
│ └── connect_foundry_appinsights.py # creates the account-level AppInsights connection
└── queries/
├── agent-traces.kql # hosted-agent traces, last 1h
├── mcp-tool-calls.kql # MCP tool invocation breakdown
├── silent-cron-debug.kql # ACA Job exec failures with no console logs
└── first-trace-probe.kql # smoke query — "did ANY trace land in last 5 min?"
```
Drop these into a PoC's `infra/modules/`, `infra/scripts/`, `src/<svc>/`,
and `docs/queries/` respectively. The patterns work as-shipped — replace
parameter values, re-deploy, traces flow.
---
## Layer 1 — Bicep substrate
### Step 1.1 — Single LAW for the whole pilot
```bicep
// infra/modules/log-analytics.bicep — drop-in
@description('Log Analytics workspace. ONE per pilot. AppIn binds to it; ACA env binds to it; cron jobs ship console+system logs to it.')
param location string = resourceGroup().location
param name string
resource law 'Microsoft.OperationalInsights/workspaces@2023-09-01' = {
name: name
location: location
properties: {
sku: { name: 'PerGB2018' }
retentionInDays: 30
features: { enableLogAccessUsingOnlyResourcePermissions: true }
}
}
output workspaceId string = law.id // ARM ID — used by AppIn + ACA env
output customerId string = law.properties.customerId // GUID — for KQL queries
output workspaceName string = law.name
```
### Step 1.2 — App Insights bound to that LAW
```bicep
// infra/modules/app-insights.bicep — drop-in
@description('App Insights component. workspace-based (legacy classic mode is deprecated). Auto-grants Monitoring Metrics Publisher to the workload UAMI so OTel traces, logs, and metrics can be ingested keylessly.')
param location string = resourceGroup().location
param name string
param workspaceId string
param uamiPrincipalId string // Foundry agent UAMI principal — for RBAC
resource appin 'Microsoft.Insights/components@2020-02-02' = {
name: name
location: location
kind: 'web'
properties: {
Application_Type: 'web'
WorkspaceResourceId: workspaceId
DisableLocalAuth: true // RBAC-only ingestion — keyless mandate
publicNetworkAccessForIngestion: 'Enabled'
publicNetworkAccessForQuery: 'Enabled'
}
}
// Monitoring Metrics Publisher — required for OTel trace/log/metric ingestion
// when DisableLocalAuth is true. Has both Microsoft.Insights/Metrics/Write and
// Microsoft.Insights/Telemetry/Write dataActions.
// Role GUID: 3913510d-42f4-4e42-8a64-420c390055eb (well-known)
resource dataIngestor 'Microsoft.Authorization/roleAssignments@2022-04-01' = {
scope: appin
name: guid(appin.id, uamiPrincipalId, '3913510d-42f4-4e42-8a64-420c390055eb')
properties: {
principalId: uamiPrincipalId
principalType: 'ServicePrincipal'
roleDefinitionId: subscriptionResourceId(
'Microsoft.Authorization/roleDefinitions',
'3913510d-42f4-4e42-8a64-420c390055eb'
)
}
}
output id string = appin.id
output name string = appin.name
output connectionString string = appin.properties.ConnectionString
output instrumentationKey string = appin.properties.InstrumentationKey
```
> **Why `DisableLocalAuth: true`.** Threadlight pilots are keyless by
> mandate (see `azure-tenant-isolation` and `citadel-spoke-onboarding`).
> AppIn ingestion keys are a back-door around RBAC — disable them and
> rely on `Monitoring Metrics Publisher` (GUID
> `3913510d-42f4-4e42-8a64-420c390055eb`) to gate writes.
> That built-in role is sufficient for OTel trace/log/metric ingestion
> because it includes `Microsoft.Insights/Telemetry/Write`.
> `Application Insights Data Ingestor` is a common misconception, not a
> real built-in role. Using the wrong role with `DisableLocalAuth: true`
> causes HTTP 400 "Bad Request" from the
> `azure-monitor-opentelemetry-exporter`.
### Step 1.3 — main.bicep wiring
```bicep
// infra/main.bicep — observability is ALWAYS-ON
module law 'modules/log-analytics.bicep' = {
name: 'law-${envName}'
params: { name: 'log-${envName}' }
}
module appInsights 'modules/app-insights.bicep' = {
name: 'appin-${envName}'
params: {
name: 'appin-${envName}'
workspaceId: law.outputs.workspaceId
uamiPrincipalId: uami.outputs.principalId
}
}
// ACA env binds to the same LAW so console+system logs land alongside traces
module acaEnv 'modules/aca-env-monitoring.bicep' = {
name: 'env-${envName}'
params: {
name: 'env-${envName}'
workspaceCustomerId: law.outputs.customerId
workspaceSharedKey: listKeys(law.outputs.workspaceId, '2023-09-01').primarySharedKey // ACA env still requires shared-key today
appInsightsConnectionString: appInsights.outputs.connectionString
}
}
// Every ACA app + ACA job container reads APPLICATIONINSIGHTS_CONNECTION_STRING
// from env. Pass it through `containers[].env`:
// - { name: 'APPLICATIONINSIGHTS_CONNECTION_STRING', value: appInsights.outputs.connectionString }
//
// MCP server, bot, workspace, deadline-watcher cron — all four get the same env var.
// The Foundry hosted agent does NOT — Foundry auto-injects it (Layer 2).
```
> **The shared-key gotcha.** Azure Container Apps environment binding
> to LAW still uses `customerId + sharedKey` (not RBAC) as of late 2025.
> This is the one remaining keyed surface in an otherwise-keyless stack.
> Document the exception in your README; don't fight it.
---
## Layer 2 — Foundry hosted agent (the runtime)
> **⚡ START HERE if `az rest --method PUT` returns 400/AAD or `GET` returns `credentials: null`:**
> Jump directly to the O-012 workaround below. The auto-injection path does NOT work for all account types — if PUT fails, skip to O-012 immediately.
> 🚨 **START HERE if your AppInsights connection PUT fails or your agent
> reports `server_error` after a fresh `azd provision`.**
>
> The "normal" Layer 2 path below assumes the platform auto-injects
> `APPLICATIONINSIGHTS_CONNECTION_STRING` after you create an account-level
> connection. **On O-012-affected accounts (recurring in 2026-05-28), the
> auto-injection silently drops** — the connection looks fine but the env
> var never lands in the container, so traces never reach AppInsights.
>
> **Decision tree:**
> 1. Try the Step 2.1 PUT below (`authType: ApiKey + credentials.key`).
> 2. **If PUT returns HTTP 400 `ValidationError "AuthType for AppInsights
> Connection can only be ApiKey"`** → you sent `authType: AAD`; switch
> to `ApiKey` and retry. See O-012 row in § Common silent-failure modes.
> 3. **If PUT returns HTTP 200 but GET returns `credentials: null`** → you
> hit the silent-drop variant of O-012; use the underscored env-var
> passthrough workaround (`APPLICATION_INSIGHTS_CONNECTION_STRING`, with
> underscore — NOT the reserved no-underscore name) from
> `HostedAgentDefinition.environment_variables` via `create_version()`,
> and call `configure_azure_monitor(connection_string=...)` explicitly
> in `container.py`. See O-012 row for the full forensic.
> 4. **If you see no traces in AppInsights despite a healthy connection
> record** → check `agent.yaml` did NOT set
> `APPLICATIONINSIGHTS_CONNECTION_STRING` (it's a reserved name; setting
> it blocks auto-injection). Remove and redeploy.
The hosted agent runtime emits traces automatically — but ONLY if the
account has an AppInsights connection registered. **Without that
connection, the runtime silently drops every span.**
> **Architecture (MAF 1.6.0+).** The platform (`azure.ai.agentserver`)
> manages its OWN OTel pipeline via `_tracing.py:_setup_log_export()`.
> This pipeline captures platform log records (HTTP requests, agent
> lifecycle, message routing) but does NOT export dependency spans.
> With 1.6.0, the hosting package bundles `microsoft-opentelemetry`
> which adds the SpanExporter + all instrumentors (openai-v2, httpx,
> etc.) — gen_ai dependency spans flow automatically.
>
> **Do NOT call standalone `configure_azure_monitor()` in
> `container.py`** — it conflicts with the platform's TracerProvider
> setup, causing duplicate log records and/or lost spans. The only
> telemetry code you need is the env var passthrough (for O-012
> workaround) and optionally `client.configure_azure_monitor()`.
> See `foundry-hosted-agents` § MAF 1.6.0 update.
### Step 2.1 — Create the account-level connection (postprovision)
> 🛣️ **Path A (recommended — try this first):** PUT with `authType: ApiKey`
> + `credentials.key` from your AppInsights instrumentation key. Succeeds
> on most accounts; data lands within 1-2 min after first invocation.
>
> 🛣️ **Path B (O-012 fallback — only if Path A returns 400/AAD or GET
> returns `credentials: null`):** skip the PUT entirely, pass
> `APPLICATION_INSIGHTS_CONNECTION_STRING` (with underscore between
> APPLICATION and INSIGHTS) directly in the agent's
> `HostedAgentDefinition.environment_variables`, and call
> `configure_azure_monitor(connection_string=...)` explicitly in
> `container.py`. The reserved no-underscore name is platform-managed; the
> underscored variant is accepted as a user override. Verified: 88 traces +
> gen_ai dependency spans landed in AppInsights from a hosted agent on
> `northcentralus`. See O-012 row for full forensic.
```bash
# infra/scripts/connect_foundry_appinsights.py — see references/postprovision/
uv run infra/scripts/connect_foundry_appinsights.py
```
The script:
1. Reads `${AZURE_FOUNDRY_ACCOUNT_NAME}` and `${AZURE_APPINSIGHTS_RESOURCE_ID}` from azd env
2. PUTs to `https://management.azure.com/subscriptions/.../accounts/{name}/connections/AppInsights?api-version=2025-04-01-preview`
3. Body: `{ "properties": { "category": "AppInsights", "target": "<armResourceId>", "metadata": { "ApiType": "Azure" } } }`
4. The connection is on the **account**, not the project — this is the
trap that bites everyone
After this runs, the platform begins injecting
`APPLICATIONINSIGHTS_CONNECTION_STRING` into every hosted-agent
container revision automatically.
### Step 2.2 — RBAC for the agent identities
The hosted-agent platform creates **two** managed identities per agent
(`AgentService-<agent-name>` + `Foundry-<workspace>`). Both need
`Monitoring Metrics Publisher` (role GUID
`3913510d-42f4-4e42-8a64-420c390055eb`) on the AppInsights resource,
OR they fail to ingest telemetry with HTTP 400
"Bad Request". The Bicep in Step 1.2 grants it to your workload UAMI; the
postprovision script extends the same grant to the platform-managed
identities.
### Step 2.3 — agent.yaml — what NOT to set
This historical heading covers the same reserved-name rule in the selected
[versioned hosted manifest](../foundry-hosted-agents/references/hosted-contract.json).
Do not regenerate a legacy `agent.yaml` for a unified-profile consumer. Keep
management environment values distinct from user-declared container variables;
management PATCH success is not evidence that the serving process loaded the
new telemetry configuration. Use a safe actual-runtime discriminator, not an
automatic restart or a diagnostic invoke that replays business work.
```yaml
# agent.yaml
environment_variables:
- name: COSMOS_ENDPOINT
value: ${AZURE_COSMOS_ENDPOINT}
- name: SEARCH_ENDPOINT
value: ${AZURE_SEARCH_ENDPOINT}
# ❌ DO NOT add APPLICATIONINSIGHTS_CONNECTION_STRING here
# The platform injects it from the account-level connection.
# Setting it manually causes telemetry collisions.
```
Add it to `agent.yaml` and the agent runtime errors with
`APPLICATIONINSIGHTS_CONNECTION_STRING is reserved`.
> **🔑 Env-var naming distinction (O-012 workaround):**
> - **Reserved (platform-managed, NO underscore):** `APPLICATIONINSIGHTS_CONNECTION_STRING` — set by platform from account-level connection; CANNOT be set in `agent.yaml` or via normal env vars.
> - **Override name (user-settable, WITH underscore):** `APPLICATION_INSIGHTS_CONNECTION_STRING` — passed ONLY via `HostedAgentDefinition.environment_variables` in `create_version()` when using O-012 workaround; platform reserves the no-underscore name but accepts this underscored variant.
> Make this distinction explicit when configuring O-012 fallback — use the underscored name only in HostedAgentDefinition, never in agent.yaml.
> **⚠️ Layer 2 caveat — hosted-agent containers MUST guard the init too.**
> The "platform auto-injects `APPLICATIONINSIGHTS_CONNECTION_STRING`"
> promise is **best-effort, not contractual** — we have field evidence
> that it can silently fail. In some regions, the AppInsights
> account-level connection can persist as
> `credentials: null` (silent-drop on AAD-rejected → ApiKey-fallback PUT
> — see "Auth-type platform forensic" below). The platform did NOT
> inject the env var. The hosted-agent container called raw
> `configure_azure_monitor()` as the first line of `main()`. The SDK
> raised `ValueError`. The container crashed before `ResponsesHostServer`
> bound. Foundry returned `server_error`/`model:""` on every smoke —
> with ZERO telemetry to debug it (telemetry init was what crashed).
> **The agent itself was fine.**
>
> **🚨 CRITICAL.** If PUT returns 400/AAD or GET returns `credentials:null`, you HAVE hit O-012 — use the underscored env var passthrough (`APPLICATION_INSIGHTS_CONNECTION_STRING`, with underscore between APPLICATION and INSIGHTS) + explicit `configure_azure_monitor(connection_string=...)`. DO NOT silently fall through to assuming auto-injection works.
>
> **Discipline.** Hosted-agent `container.py` MUST use the same
> guarded-init pattern as ACA workloads (Layer 3 helper below), NOT
> raw `configure_azure_monitor()`. Treat the platform's auto-injection
> guarantee as best-effort — guard it like Layer 3 does. See gap rows
> O-011 / O-012 for the full forensic and the canonical inline helper
> shape.
---
## Layer 3 — ACA workloads + hosted-agent containers (MCP / bot / workspace / cron / agent)
The Foundry runtime auto-instruments hosted agents, but **everything
else** (your MCP server, your bot, your workspace, your cron jobs) is
plain Python or Node — you wire OTel yourself. **And per the Layer 2
Ver en GitHub