Skip to main content

foundry-cost-monitoring

Continuous Foundry token-cost monitoring, chargeback, budget alerts, and FinOps automation. Joins gen_ai.usage spans emitted by foundry-observability with the Azure Retail Prices API to compute per-agent / per-project / per-tenant cost projection. Wires Cost Management budgets + Action Groups, and uses Foundry's project-tag attribution (preview) for chargeback. USE FOR: foundry cost monitoring, token cost dashboard, FinOps Foundry, project-level chargeback, project tag cost attribution, Azure retail prices API, gen_ai.usage cost KQL, budget alert Foundry, Action Group cost webhook, token cost anomaly detection, OTel vs Cost Management reconciliation, per-agent / per-tenant token spend. DO NOT USE FOR: one-shot PTU vs PAYGO sizing (use paygo-ptu-cost-analyzer); emitting cost telemetry from agent code (use foundry-observability); gateway-only chargeback via x-app-id (use citadel-spoke-onboarding); fine-tuning cost estimates.

Zur Installation springen

Quellinformationen

Repository
aiappsgbb/awesome-gbb
Letzte Quellaktivität
18. August 2026 um 20:57
Erkannte Sprache von SKILL.md
Englisch
Sterne
6
Forks
3

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

Datei-Explorer
3 Dateien

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
foundry-cost-monitoring
description
Continuous Foundry token-cost monitoring, chargeback, budget alerts, and FinOps automation. Joins gen_ai.usage spans emitted by foundry-observability with the Azure Retail Prices API to compute per-agent / per-project / per-tenant cost projection. Wires Cost Management budgets + Action Groups, and uses Foundry's project-tag attribution (preview) for chargeback. USE FOR: foundry cost monitoring, token cost dashboard, FinOps Foundry, project-level chargeback, project tag cost attribution, Azure retail prices API, gen_ai.usage cost KQL, budget alert Foundry, Action Group cost webhook, token cost anomaly detection, OTel vs Cost Management reconciliation, per-agent / per-tenant token spend. DO NOT USE FOR: one-shot PTU vs PAYGO sizing (use paygo-ptu-cost-analyzer); emitting cost telemetry from agent code (use foundry-observability); gateway-only chargeback via x-app-id (use citadel-spoke-onboarding); fine-tuning cost estimates.
metadata
{"version":"1.0.5"}
# Foundry Cost Monitoring Continuous, FinOps-grade cost monitoring for Microsoft Foundry. Where [`foundry-observability`](../foundry-observability/SKILL.md) emits the `gen_ai.usage.*` spans, this skill **consumes** them — joining them with the Azure Retail Prices API to project per-agent, per-project, and per-tenant cost in near-real time, then wiring Cost Management budgets and Action Groups so the FinOps team sees overruns within minutes instead of next-month's invoice. > **Sibling skills.** Pair with `foundry-observability` for the telemetry > emit path, [`paygo-ptu-cost-analyzer`](../paygo-ptu-cost-analyzer/SKILL.md) > for one-shot PTU sizing decisions, and > [`citadel-spoke-onboarding`](../citadel-spoke-onboarding/SKILL.md) when > chargeback is gateway-mediated via APIM `x-app-id` headers. --- ## §1 — Three layers of cost truth Foundry cost lands in three places, with three different latencies and three different fidelities: | Layer | Source | Latency | Fidelity | Use for | |-------|--------|---------|----------|---------| | **1. Foundry portal estimate** | `ai.azure.com` → Operate → Overview / Build → Agent / Model Monitor | ≈ 2–5 min | **Estimate** — not invoiceable | Near-real-time engineering insight (per-agent rollup) | | **2. OTel real-time projection** | App Insights `customDimensions['gen_ai.usage.*']` × Retail Prices rate card | ≈ 2–5 min ingest + your KQL window | **Engineering projection** — not invoiceable | Per-tenant / per-agent / per-conversation projection, anomaly detection, dashboards | | **3. Cost Mgmt actuals** | `Microsoft.CostManagement/query` REST + invoice CSV exports | ≈ 8–24 h | **Source of truth** for billing | Finance reconciliation, chargeback, budgets | **Reconciliation rule** (and this is the rule that prevents the awkward quarterly call with finance): treat **layer 3** as authoritative for anything that ends up on the bill. Layers 1 and 2 are engineering tools to **see overruns before** layer 3 catches up. Per the MS Learn cost guidance: *"treat your invoice and meter records as the source of truth."* --- ## §2 — Pricing source: Azure Retail Prices REST API The [Retail Prices API](https://learn.microsoft.com/rest/api/cost-management/retail-prices/azure-retail-prices) is anonymous, public, and rate-card complete: ``` GET https://prices.azure.com/api/retail/prices ?$filter=<OData filter> [&$top=N] [&$skip=N] ``` No auth header. No subscription. Just `curl` (or `urllib.request`). The filter shape that matters for Foundry models in Sweden Central: ```bash # Foundry models (post-rename — current GA service classification) curl -s "https://prices.azure.com/api/retail/prices?\$filter=serviceName eq 'Foundry Models' and armRegionName eq 'swedencentral' and contains(meterName, 'Tokens')" | jq '.Items[0:3]' # Embeddings — Sweden Central requires GlobalStandard SKU; filter accordingly curl -s "https://prices.azure.com/api/retail/prices?\$filter=serviceName eq 'Foundry Models' and armRegionName eq 'swedencentral' and contains(productName, 'embedding')" ``` Each `Items[]` row carries the fields you need to price a token row: | Field | Example | What it means | |-------|---------|---------------| | `serviceName` | `Foundry Models` | Top-level service classification | | `productName` | `Azure OpenAI`, `Azure Llama Models`, `Azure OpenAI GPT5` | Model family bucket | | `skuName` | `gpt 4.1 Inp regnl` | Deployment-shape SKU | | `meterName` | `gpt 4.1 Inp regnl Tokens` | Billing meter | | `retailPrice` | `0.00242` | Per `unitOfMeasure` | | `unitOfMeasure` | `1K`, `1M` | Token unit (mix of `1K` and `1M` — **normalize before multiplying**) | | `currencyCode` | `USD` | Pricing currency | | `armRegionName` | `swedencentral` | Region | | `type` | `Consumption` | `Consumption` for PAYGO, `Reservation` for PTU | **Pagination**: response includes `NextPageLink`; iterate until null when caching the full catalog (~thousands of rows for `AI + Machine Learning`). > **Service-name drift.** The Retail Prices API used to surface OpenAI > meters under `serviceName eq 'Cognitive Services'`. Current GA > classification is **`Foundry Models`**. The drift detector for this > skill polls the Retail Prices endpoint weekly — if a new classification > appears, refresh § 11 of this skill and bump PATCH. **Cache TTL recommendation**: 24 h. Rates change weekly at most; daily refresh comfortably beats invoice cadence. --- ## §3 — Foundry `project` tag chargeback (Preview) Foundry automatically tags Models-sold-by-Azure usage records with a `project` tag whose value is the Foundry project name. The chargeback flow is then: 1. In Cost Management → Cost Analysis, scope to your Foundry resource. 2. **Add filter** → **Tag** → `project`. 3. Pick one or more projects → cost split by project, time-series. CLI equivalent: ```bash az consumption usage list \ --subscription "$SUB" \ --start-date "$(date -u -v-30d +%Y-%m-%d)" \ --end-date "$(date -u +%Y-%m-%d)" \ --query "[?tags.project=='proj-fsi-claims'].{date:usageStart, meter:meterDetails.meterName, qty:quantity, cost:pretaxCost}" \ -o table ``` > **Verbatim preview limit (MS Learn, [manage-costs#chargeback-with-project-level-cost-attribution-preview](https://learn.microsoft.com/azure/foundry/concepts/manage-costs#chargeback-with-project-level-cost-attribution-preview)):** > > *"Project-level cost attribution is currently supported for Models > sold by Azure (Azure Direct models, including Azure OpenAI). It isn't > yet supported for models served through Azure Marketplace."* Concretely: Azure OpenAI / Foundry-native models → tagged automatically. Llama / Mistral / DeepSeek / Cohere via Azure Marketplace → **no project tag**; fall back to layer-2 OTel projection (§ 4) or per-resource grouping (one project ↔ one resource group). The tag is **not manually applied** — do not try to add it via `az tag create-or-update`; Foundry control plane writes it on the underlying meter records. --- ## §4 — In-flight cost projection: KQL on `gen_ai.usage.*` This is the engineering-insight layer. Requires `foundry-observability` already wired (account-level App Insights connection + ACA-side `configure_azure_monitor()`). The MAF 1.6.0+ schema emits OTel spans with the following `customDimensions`: | Key | Type | Example | |-----|------|---------| | `gen_ai.system` | string | `azure_openai` | | `gen_ai.request.model` | string | `gpt-5.4-mini` | | `gen_ai.response.model` | string | `gpt-5.4-mini-2026-01-15` | | `gen_ai.usage.input_tokens` | int | `1284` | | `gen_ai.usage.output_tokens` | int | `342` | | `gen_ai.operation.name` | string | `chat`, `embeddings` | Spans land in `traces` / `dependencies` depending on instrumentation. The canonical filter for MAF-emitted spans is `cloud_RoleName == "agent_framework"` (verified at MAF 1.6.0+ — see [`foundry-observability`](../foundry-observability/SKILL.md) § "OTel cloud_RoleName" for the rationale). ### 4.1 Per-agent cost rollup (last 24 h) ```kql // Per-agent token cost (last 24h) — projection-only, not invoiceable let rateCard = externaldata(model:string, in_rate_usd_per_1k:real, out_rate_usd_per_1k:real) [@"https://raw.githubusercontent.com/<your-org>/<your-repo>/main/rate-card.csv"] with (format="csv", ignoreFirstRecord=true); dependencies | where timestamp > ago(24h) | where cloud_RoleName == "agent_framework" | where isnotempty(customDimensions["gen_ai.usage.input_tokens"]) | extend agent = tostring(customDimensions["gen_ai.agent.name"]), model = tostring(customDimensions["gen_ai.request.model"]), in_tokens = toint(customDimensions["gen_ai.usage.input_tokens"]), out_tokens = toint(customDimensions["gen_ai.usage.output_tokens"]) | join kind=leftouter rateCard on model | extend in_cost = (in_tokens / 1000.0) * in_rate_usd_per_1k, out_cost = (out_tokens / 1000.0) * out_rate_usd_per_1k | summarize calls = count(), input_tokens = sum(in_tokens), output_tokens = sum(out_tokens), projected_cost_usd = sum(in_cost + out_cost) by agent, model | order by projected_cost_usd desc ``` ### 4.2 Per-project cost rollup `gen_ai.agent.name` carries the agent name. To get per-project rollup without the preview tag, instrument the project name as a resource attribute on the OTel exporter (e.g. `OTEL_RESOURCE_ATTRIBUTES="foundry.project=proj-fsi-claims"`) and group on `customDimensions["foundry.project"]` instead of `agent`. Where the `project` tag (§ 3) IS available, prefer Cost Management for authoritative reconciliation and reserve KQL for in-flight projection. ### 4.3 Per-tenant cost rollup (Citadel APIM gateway path) If traffic goes through the AI Citadel gateway (see § 5), the APIM fragment-policy stamps the tenant identifier into the OTel span as `enduser.id` (gateway maps `x-app-id` header → `enduser.id`). The per-tenant rollup is identical to § 4.1 but groups on `tostring(customDimensions["enduser.id"])`. --- ## §5 — Citadel APIM `x-app-id` chargeback path When Foundry traffic is fronted by an APIM AI Gateway (the AI Citadel Hub topology), tenant identity arrives on the `x-app-id` request header and the gateway logs it into App Insights. This is the cleanest chargeback path for multi-tenant SaaS deployments because it does not depend on Foundry's project-tag preview. Wiring path (full detail in [`citadel-spoke-onboarding`](../citadel-spoke-onboarding/SKILL.md)): ``` Tenant request → APIM (validates JWT, stamps x-app-id) → Foundry ↓ OTel span tagged with enduser.id = x-app-id ``` KQL aggregates over `enduser.id` then rejoin to the customer-ID-to-name table (kept outside this skill — usually a Cosmos/Kusto reference table). --- ## §6 — Budget alerts via Azure Cost Management Set bottom-up: subscription budget + per-RG budgets for spike isolation. ### 6.1 Bicep (recommended) ```bicep // budget.bicep — per-RG budget with notifications at 80% / 100% / 120% @description('Budget amount in USD per month') param budgetAmount int = 5000 @description('FinOps distribution list') param finopsEmails array = [ 'finops@<example-bank>.com' ] resource budget 'Microsoft.Consumption/budgets@2024-08-01' = { name: 'budget-foundry-${resourceGroup().name}' properties: { amount: budgetAmount timeGrain: 'Monthly' timePeriod: { startDate: '2026-06-01' } category: 'Cost' notifications: { Warning_80: { enabled: true operator: 'GreaterThan' threshold: 80 contactEmails: finopsEmails thresholdType: 'Actual' } Critical_100: { enabled: true operator: 'GreaterThan' threshold: 100 contactEmails: finopsEmails contactGroups: [ actionGroup.id ] thresholdType: 'Actual' } Forecast_120: { enabled: true operator: 'GreaterThan' threshold: 120 contactEmails: finopsEmails thresholdType: 'Forecasted' } } filter: { tags: { name: 'project' operator: 'In' values: [ 'proj-fsi-claims' 'proj-fsi-onboarding' ] } } } } resource actionGroup 'Microsoft.Insights/actionGroups@2024-10-01-preview' = { name: 'ag-finops-${resourceGroup().name}' location: 'global' properties: { groupShortName: 'finops' enabled: true webhookReceivers: [ { name: 'finops-teams' serviceUri: '<teams-incoming-webhook-url>' useCommonAlertSchema: true } ] } } ``` ### 6.2 CLI equivalent ```bash # Resource-group scoped budget — fires Action Group on actual cost > $5000 az consumption budget create-with-rg \ --resource-group "$RG" \ --budget-name "budget-foundry-$RG" \ --amount 5000 \ --time-grain Monthly \ --start-date "$(date -u +%Y-%m-01)" \ --category Cost \ --notifications-properties \ "[{operator:'GreaterThan',threshold:80,contactEmails:['finops@<example-bank>.com'],thresholdType:'Actual'}]" ``` > **Hard-limit caveat (MS Learn manage-costs § Create budgets):** > *"Azure OpenAI doesn't currently provide [hard-limit] functionality. > You can start automation from action groups as part of your budget > notifications to take more advanced actions, but this functionality > requires additional custom development."* > > Budgets are **alerts**, not throttles. For hard ceilings, wire the > Action Group → Logic App / Function that disables the offending > deployment (e.g. `az cognitiveservices account deployment update > --properties.callRateLimit=0`). --- ## §7 — Action Group → Webhook / Logic App / Teams Standard Azure pattern. The Action Group from § 6.1 wires three common downstream channels: | Channel | Use case |
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen