Skip to main content

foundry-observability

End-to-end observability for Azure AI pilots — App Insights + Log Analytics + OpenTelemetry across hosted agents, ACA MCP servers, ACA jobs, bot service, workspace UIs. Closes the silent telemetry gap where `azd up` returns 0 but **zero traces ever reach App Insights**. Covers Bicep modules, Foundry account-level telemetry connection, ACA-side instrumentation, `Monitoring Metrics Publisher` RBAC, and KQL diagnostic queries. Read the full skill body for the 3-layer wiring sequence — do not instrument from this summary alone. USE FOR: app insights, application insights, OpenTelemetry, OTel, configure_azure_monitor, agent traces missing, no telemetry, blank appin, log analytics, KQL, observability, trace MCP, silent cron, Monitoring Metrics Publisher RBAC, AppInsights connection foundry, account-level appin, AppIn PUT 400, credentials null, silent injection, server_error telemetry. DO NOT USE FOR: continuous eval (foundry-evals), pre-deploy gates (threadlight-safe-check), Foundry IQ monitoring (foundry-iq).

설치로 이동

소스 정보

저장소
aiappsgbb/awesome-gbb
최근 소스 활동
2026년 9월 25일 14:00
감지된 SKILL.md 언어
영어
스타
6
포크
3

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
15 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
foundry-observability
description
End-to-end observability for Azure AI pilots — App Insights + Log Analytics + OpenTelemetry across hosted agents, ACA MCP servers, ACA jobs, bot service, workspace UIs. Closes the silent telemetry gap where `azd up` returns 0 but **zero traces ever reach App Insights**. Covers Bicep modules, Foundry account-level telemetry connection, ACA-side instrumentation, `Monitoring Metrics Publisher` RBAC, and KQL diagnostic queries. Read the full skill body for the 3-layer wiring sequence — do not instrument from this summary alone. USE FOR: app insights, application insights, OpenTelemetry, OTel, configure_azure_monitor, agent traces missing, no telemetry, blank appin, log analytics, KQL, observability, trace MCP, silent cron, Monitoring Metrics Publisher RBAC, AppInsights connection foundry, account-level appin, AppIn PUT 400, credentials null, silent injection, server_error telemetry. DO NOT USE FOR: continuous eval (foundry-evals), pre-deploy gates (threadlight-safe-check), Foundry IQ monitoring (foundry-iq).
metadata
{"version":"1.2.5"}
# Foundry Observability End-to-end telemetry across every component of a Threadlight pilot: Foundry hosted agent, MCP servers on ACA, ACA jobs (cron triggers), bot service, workspace UI. **Default discipline**, not optional. For the complete per-agent adoption and release-evidence workflow, see [`foundry-agentops`](../foundry-agentops/SKILL.md). This skill remains authoritative for OpenTelemetry and App Insights wiring; AgentOps aggregates evidence and never replaces instrumentation or verification of the telemetry path. > **Downstream FinOps consumer.** [`foundry-cost-monitoring`](../foundry-cost-monitoring/SKILL.md) > joins the `gen_ai.usage.*` spans this skill emits with the Azure > Retail Prices API to compute per-agent / per-project / per-tenant > cost projection — wire it whenever a FinOps stakeholder needs to > answer "what is this agent costing us right now?" > **Why this skill exists.** Recent pilots deployed cleanly > (`azd up` returned 0, all resources provisioned) but App Insights > stayed **completely empty** — no agent traces, no MCP tool calls, > no cron logs. Root cause: no one wired the connection at any layer. > The intel for *each layer* lives scattered across `threadlight-deploy`, > `foundry-hosted-agents`, `foundry-mcp-aca`, `threadlight-event-triggers` — > but no single skill walks an operator through the full chain. That's > what this skill does. Pair with `threadlight-safe-check` Step 5.6 > (App Insights existence + first-trace probe) to gate it shut. --- ## Mental model — three layers, one signal ``` ┌─────────────────────────────────────────────────────────────────────┐ │ Layer 3: ACA workloads (MCP / bot / workspace / cron jobs) │ │ • configure_azure_monitor() reads APPLICATIONINSIGHTS_CONNECTION_STRING │ │ • Env var set by Bicep from app-insights.outputs.connectionString │ │ • OTel exporter ships spans + logs + metrics over HTTPS │ └─────────────────────────────────────────────────────────────────────┘ ▲ │ direct push from container code │ ┌─────────────────────────────────┼───────────────────────────────────┐ │ Layer 2: Foundry hosted agent (the runtime) │ │ • Account-level AppInsights connection (category: AppInsights) │ │ • Platform AUTO-INJECTS APPLICATIONINSIGHTS_CONNECTION_STRING │ │ • RBAC: Monitoring Metrics Publisher on agent identities │ │ • Tracing emitted by the runtime — no app code change │ └─────────────────────────────────┼───────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────┐ │ Layer 1: Bicep substrate │ │ • app-insights.bicep — workspace-based (LAW-bound) │ │ • log-analytics.bicep — single LAW for ALL workloads in the RG │ │ • ACA env wiring: dapr.appInsightsConnectionString OR direct env │ │ • Output `connectionString` consumed by every workload │ └─────────────────────────────────────────────────────────────────────┘ ``` **Single connection string, three fan-outs.** All telemetry lands in the same App Insights resource. No per-workload AppIn — that fragments the trace graph and makes correlation impossible. --- ## What ships from this skill ``` foundry-observability/ ├── SKILL.md └── references/ ├── bicep/ │ ├── log-analytics.bicep # LAW (workspace) — required by both AppIn and ACA env │ ├── app-insights.bicep # AppIn workspace-based + UAMI Monitoring Metrics Publisher RBAC │ └── aca-env-monitoring.bicep # ACA env wired to LAW + AppIn ├── python/ │ └── otel_init.py # configure_azure_monitor() for ACA workloads ├── postprovision/ │ └── connect_foundry_appinsights.py # creates the account-level AppInsights connection └── queries/ ├── agent-traces.kql # hosted-agent traces, last 1h ├── mcp-tool-calls.kql # MCP tool invocation breakdown ├── silent-cron-debug.kql # ACA Job exec failures with no console logs └── first-trace-probe.kql # smoke query — "did ANY trace land in last 5 min?" ``` Drop these into a PoC's `infra/modules/`, `infra/scripts/`, `src/<svc>/`, and `docs/queries/` respectively. The patterns work as-shipped — replace parameter values, re-deploy, traces flow. --- ## Layer 1 — Bicep substrate ### Step 1.1 — Single LAW for the whole pilot ```bicep // infra/modules/log-analytics.bicep — drop-in @description('Log Analytics workspace. ONE per pilot. AppIn binds to it; ACA env binds to it; cron jobs ship console+system logs to it.') param location string = resourceGroup().location param name string resource law 'Microsoft.OperationalInsights/workspaces@2023-09-01' = { name: name location: location properties: { sku: { name: 'PerGB2018' } retentionInDays: 30 features: { enableLogAccessUsingOnlyResourcePermissions: true } } } output workspaceId string = law.id // ARM ID — used by AppIn + ACA env output customerId string = law.properties.customerId // GUID — for KQL queries output workspaceName string = law.name ``` ### Step 1.2 — App Insights bound to that LAW ```bicep // infra/modules/app-insights.bicep — drop-in @description('App Insights component. workspace-based (legacy classic mode is deprecated). Auto-grants Monitoring Metrics Publisher to the workload UAMI so OTel traces, logs, and metrics can be ingested keylessly.') param location string = resourceGroup().location param name string param workspaceId string param uamiPrincipalId string // Foundry agent UAMI principal — for RBAC resource appin 'Microsoft.Insights/components@2020-02-02' = { name: name location: location kind: 'web' properties: { Application_Type: 'web' WorkspaceResourceId: workspaceId DisableLocalAuth: true // RBAC-only ingestion — keyless mandate publicNetworkAccessForIngestion: 'Enabled' publicNetworkAccessForQuery: 'Enabled' } } // Monitoring Metrics Publisher — required for OTel trace/log/metric ingestion // when DisableLocalAuth is true. Has both Microsoft.Insights/Metrics/Write and // Microsoft.Insights/Telemetry/Write dataActions. // Role GUID: 3913510d-42f4-4e42-8a64-420c390055eb (well-known) resource dataIngestor 'Microsoft.Authorization/roleAssignments@2022-04-01' = { scope: appin name: guid(appin.id, uamiPrincipalId, '3913510d-42f4-4e42-8a64-420c390055eb') properties: { principalId: uamiPrincipalId principalType: 'ServicePrincipal' roleDefinitionId: subscriptionResourceId( 'Microsoft.Authorization/roleDefinitions', '3913510d-42f4-4e42-8a64-420c390055eb' ) } } output id string = appin.id output name string = appin.name output connectionString string = appin.properties.ConnectionString output instrumentationKey string = appin.properties.InstrumentationKey ``` > **Why `DisableLocalAuth: true`.** Threadlight pilots are keyless by > mandate (see `azure-tenant-isolation` and `citadel-spoke-onboarding`). > AppIn ingestion keys are a back-door around RBAC — disable them and > rely on `Monitoring Metrics Publisher` (GUID > `3913510d-42f4-4e42-8a64-420c390055eb`) to gate writes. > That built-in role is sufficient for OTel trace/log/metric ingestion > because it includes `Microsoft.Insights/Telemetry/Write`. > `Application Insights Data Ingestor` is a common misconception, not a > real built-in role. Using the wrong role with `DisableLocalAuth: true` > causes HTTP 400 "Bad Request" from the > `azure-monitor-opentelemetry-exporter`. ### Step 1.3 — main.bicep wiring ```bicep // infra/main.bicep — observability is ALWAYS-ON module law 'modules/log-analytics.bicep' = { name: 'law-${envName}' params: { name: 'log-${envName}' } } module appInsights 'modules/app-insights.bicep' = { name: 'appin-${envName}' params: { name: 'appin-${envName}' workspaceId: law.outputs.workspaceId uamiPrincipalId: uami.outputs.principalId } } // ACA env binds to the same LAW so console+system logs land alongside traces module acaEnv 'modules/aca-env-monitoring.bicep' = { name: 'env-${envName}' params: { name: 'env-${envName}' workspaceCustomerId: law.outputs.customerId workspaceSharedKey: listKeys(law.outputs.workspaceId, '2023-09-01').primarySharedKey // ACA env still requires shared-key today appInsightsConnectionString: appInsights.outputs.connectionString } } // Every ACA app + ACA job container reads APPLICATIONINSIGHTS_CONNECTION_STRING // from env. Pass it through `containers[].env`: // - { name: 'APPLICATIONINSIGHTS_CONNECTION_STRING', value: appInsights.outputs.connectionString } // // MCP server, bot, workspace, deadline-watcher cron — all four get the same env var. // The Foundry hosted agent does NOT — Foundry auto-injects it (Layer 2). ``` > **The shared-key gotcha.** Azure Container Apps environment binding > to LAW still uses `customerId + sharedKey` (not RBAC) as of late 2025. > This is the one remaining keyed surface in an otherwise-keyless stack. > Document the exception in your README; don't fight it. --- ## Layer 2 — Foundry hosted agent (the runtime) > **⚡ START HERE if `az rest --method PUT` returns 400/AAD or `GET` returns `credentials: null`:** > Jump directly to the O-012 workaround below. The auto-injection path does NOT work for all account types — if PUT fails, skip to O-012 immediately. > 🚨 **START HERE if your AppInsights connection PUT fails or your agent > reports `server_error` after a fresh `azd provision`.** > > The "normal" Layer 2 path below assumes the platform auto-injects > `APPLICATIONINSIGHTS_CONNECTION_STRING` after you create an account-level > connection. **On O-012-affected accounts (recurring in 2026-05-28), the > auto-injection silently drops** — the connection looks fine but the env > var never lands in the container, so traces never reach AppInsights. > > **Decision tree:** > 1. Try the Step 2.1 PUT below (`authType: ApiKey + credentials.key`). > 2. **If PUT returns HTTP 400 `ValidationError "AuthType for AppInsights > Connection can only be ApiKey"`** → you sent `authType: AAD`; switch > to `ApiKey` and retry. See O-012 row in § Common silent-failure modes. > 3. **If PUT returns HTTP 200 but GET returns `credentials: null`** → you > hit the silent-drop variant of O-012; use the underscored env-var > passthrough workaround (`APPLICATION_INSIGHTS_CONNECTION_STRING`, with > underscore — NOT the reserved no-underscore name) from > `HostedAgentDefinition.environment_variables` via `create_version()`, > and call `configure_azure_monitor(connection_string=...)` explicitly > in `container.py`. See O-012 row for the full forensic. > 4. **If you see no traces in AppInsights despite a healthy connection > record** → check `agent.yaml` did NOT set > `APPLICATIONINSIGHTS_CONNECTION_STRING` (it's a reserved name; setting > it blocks auto-injection). Remove and redeploy. The hosted agent runtime emits traces automatically — but ONLY if the account has an AppInsights connection registered. **Without that connection, the runtime silently drops every span.** > **Architecture (MAF 1.6.0+).** The platform (`azure.ai.agentserver`) > manages its OWN OTel pipeline via `_tracing.py:_setup_log_export()`. > This pipeline captures platform log records (HTTP requests, agent > lifecycle, message routing) but does NOT export dependency spans. > With 1.6.0, the hosting package bundles `microsoft-opentelemetry` > which adds the SpanExporter + all instrumentors (openai-v2, httpx, > etc.) — gen_ai dependency spans flow automatically. > > **Do NOT call standalone `configure_azure_monitor()` in > `container.py`** — it conflicts with the platform's TracerProvider > setup, causing duplicate log records and/or lost spans. The only > telemetry code you need is the env var passthrough (for O-012 > workaround) and optionally `client.configure_azure_monitor()`. > See `foundry-hosted-agents` § MAF 1.6.0 update. ### Step 2.1 — Create the account-level connection (postprovision) > 🛣️ **Path A (recommended — try this first):** PUT with `authType: ApiKey` > + `credentials.key` from your AppInsights instrumentation key. Succeeds > on most accounts; data lands within 1-2 min after first invocation. > > 🛣️ **Path B (O-012 fallback — only if Path A returns 400/AAD or GET > returns `credentials: null`):** skip the PUT entirely, pass > `APPLICATION_INSIGHTS_CONNECTION_STRING` (with underscore between > APPLICATION and INSIGHTS) directly in the agent's > `HostedAgentDefinition.environment_variables`, and call > `configure_azure_monitor(connection_string=...)` explicitly in > `container.py`. The reserved no-underscore name is platform-managed; the > underscored variant is accepted as a user override. Verified: 88 traces + > gen_ai dependency spans landed in AppInsights from a hosted agent on > `northcentralus`. See O-012 row for full forensic. ```bash # infra/scripts/connect_foundry_appinsights.py — see references/postprovision/ uv run infra/scripts/connect_foundry_appinsights.py ``` The script: 1. Reads `${AZURE_FOUNDRY_ACCOUNT_NAME}` and `${AZURE_APPINSIGHTS_RESOURCE_ID}` from azd env 2. PUTs to `https://management.azure.com/subscriptions/.../accounts/{name}/connections/AppInsights?api-version=2025-04-01-preview` 3. Body: `{ "properties": { "category": "AppInsights", "target": "<armResourceId>", "metadata": { "ApiType": "Azure" } } }` 4. The connection is on the **account**, not the project — this is the trap that bites everyone After this runs, the platform begins injecting `APPLICATIONINSIGHTS_CONNECTION_STRING` into every hosted-agent container revision automatically. ### Step 2.2 — RBAC for the agent identities The hosted-agent platform creates **two** managed identities per agent (`AgentService-<agent-name>` + `Foundry-<workspace>`). Both need `Monitoring Metrics Publisher` (role GUID `3913510d-42f4-4e42-8a64-420c390055eb`) on the AppInsights resource, OR they fail to ingest telemetry with HTTP 400 "Bad Request". The Bicep in Step 1.2 grants it to your workload UAMI; the postprovision script extends the same grant to the platform-managed identities. ### Step 2.3 — agent.yaml — what NOT to set This historical heading covers the same reserved-name rule in the selected [versioned hosted manifest](../foundry-hosted-agents/references/hosted-contract.json). Do not regenerate a legacy `agent.yaml` for a unified-profile consumer. Keep management environment values distinct from user-declared container variables; management PATCH success is not evidence that the serving process loaded the new telemetry configuration. Use a safe actual-runtime discriminator, not an automatic restart or a diagnostic invoke that replays business work. ```yaml # agent.yaml environment_variables: - name: COSMOS_ENDPOINT value: ${AZURE_COSMOS_ENDPOINT} - name: SEARCH_ENDPOINT value: ${AZURE_SEARCH_ENDPOINT} # ❌ DO NOT add APPLICATIONINSIGHTS_CONNECTION_STRING here # The platform injects it from the account-level connection. # Setting it manually causes telemetry collisions. ``` Add it to `agent.yaml` and the agent runtime errors with `APPLICATIONINSIGHTS_CONNECTION_STRING is reserved`. > **🔑 Env-var naming distinction (O-012 workaround):** > - **Reserved (platform-managed, NO underscore):** `APPLICATIONINSIGHTS_CONNECTION_STRING` — set by platform from account-level connection; CANNOT be set in `agent.yaml` or via normal env vars. > - **Override name (user-settable, WITH underscore):** `APPLICATION_INSIGHTS_CONNECTION_STRING` — passed ONLY via `HostedAgentDefinition.environment_variables` in `create_version()` when using O-012 workaround; platform reserves the no-underscore name but accepts this underscored variant. > Make this distinction explicit when configuring O-012 fallback — use the underscored name only in HostedAgentDefinition, never in agent.yaml. > **⚠️ Layer 2 caveat — hosted-agent containers MUST guard the init too.** > The "platform auto-injects `APPLICATIONINSIGHTS_CONNECTION_STRING`" > promise is **best-effort, not contractual** — we have field evidence > that it can silently fail. In some regions, the AppInsights > account-level connection can persist as > `credentials: null` (silent-drop on AAD-rejected → ApiKey-fallback PUT > — see "Auth-type platform forensic" below). The platform did NOT > inject the env var. The hosted-agent container called raw > `configure_azure_monitor()` as the first line of `main()`. The SDK > raised `ValueError`. The container crashed before `ResponsesHostServer` > bound. Foundry returned `server_error`/`model:""` on every smoke — > with ZERO telemetry to debug it (telemetry init was what crashed). > **The agent itself was fine.** > > **🚨 CRITICAL.** If PUT returns 400/AAD or GET returns `credentials:null`, you HAVE hit O-012 — use the underscored env var passthrough (`APPLICATION_INSIGHTS_CONNECTION_STRING`, with underscore between APPLICATION and INSIGHTS) + explicit `configure_azure_monitor(connection_string=...)`. DO NOT silently fall through to assuming auto-injection works. > > **Discipline.** Hosted-agent `container.py` MUST use the same > guarded-init pattern as ACA workloads (Layer 3 helper below), NOT > raw `configure_azure_monitor()`. Treat the platform's auto-injection > guarantee as best-effort — guard it like Layer 3 does. See gap rows > O-011 / O-012 for the full forensic and the canonical inline helper > shape. --- ## Layer 3 — ACA workloads + hosted-agent containers (MCP / bot / workspace / cron / agent) The Foundry runtime auto-instruments hosted agents, but **everything else** (your MCP server, your bot, your workspace, your cron jobs) is plain Python or Node — you wire OTel yourself. **And per the Layer 2
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기