- name
- threadlight-consumption-iq
- description
- Project Azure cost from deployed Bicep + SPEC § 12, compare viable SKUs, and emit cost-projection.md + cost-manifest.json after safe-check. Also supports no-pilot pre-sales estimates with phased rollout, EA/MCA discount, and a one-pager. Owns opt-in read-only actuals and scope-bound reconciliation from Cost Management, Monitor, and Log Analytics, publishing threadlight-cost-actuals/v1 and threadlight-cost-reconciliation/v1. threadlight-production-ready only consumes verified artifacts. Advisory. USE FOR: Azure consumption projection, post-deploy cost, SKU diff, PAYG vs PTU, load profile, cost manifest, pre-sales estimate, EA/MCA discount, cost actuals, forecast vs actual, reconciliation. DO NOT USE FOR: AOAI-only break-even without a pilot (use paygo-ptu-cost-analyzer); Bicep mutation (use threadlight-deploy).
- metadata
- {"version":"0.4.0"}
# Threadlight Consumption IQ — post-deploy cost projection + SKU diff
> The single skill in the chain that asks "**what should this pilot actually
> cost to run at the customer's real production load, and which SKUs should
> we swap to before they sign off?**" and answers with a structured,
> evidence-backed artefact instead of a guess.
>
> Naming: the `-iq` suffix matches the target Field Outcomes
> view's `ai-foundry-account-iq` / `azure-account-iq` family under
> **Build & Deliver → Drive Consumption**.
## Current evidence contract
This skill owns the cost evidence lifecycle for a pilot:
1. **Forecast** — `docs/cost-projection.md` + `specs/cost-manifest.json`
project declared load against public pricing.
2. **Read-only actuals collection** — `specs/cost-actuals-manifest.json`
(`threadlight-cost-actuals/v1`) captures Cost Management / Monitor / Log
Analytics evidence for a settled window.
3. **Scope-bound reconciliation manifest** —
`specs/cost-reconciliation-manifest.json`
(`threadlight-cost-reconciliation/v1`) proves the actuals still match the
exact forecast, SPEC § 14 policy, and collection scope/window they were
reconciled against.
4. **Human reconciliation report** — `docs/cost-reconciliation.md` explains the
variance, coverage, maturity, and unit-economics result.
Production-ready consumes verified artifacts and does not query or recompute.
`threadlight-production-ready`'s cost/readiness findings are downstream
consumers of what this skill publishes:
- `COST-102` = **mature/fresh/scope-bound reconciliation**.
- `COST-103` = **PAYG/PTU recommendation at observed token volume**.
- `KPI-003` = **measured cost per successful interaction joined with eval and
telemetry evidence**.
## Why this skill exists
The `threadlight-*` chain ships a working agent in one session (design →
local-test → deploy → safe-check). `threadlight-production-ready`'s
pillar 10 (cost) then asks *static* questions: is a Budget declared? is
an anomaly alert wired? is `docs/cost-projection.md` present? Those are
yes/no checks. They do not answer **what the pilot will actually cost**
when the customer turns on production traffic.
Meanwhile `paygo-ptu-cost-analyzer` (in `awesome-gbb`) only covers AOAI
PAYG-vs-PTU break-even. Every other deployed resource — ACA, Cosmos,
Storage, APIM, AI Search, Foundry hosted-agent — gets eyeballed.
Without this skill the cost conversation at architecture review goes:
> **Customer FinOps:** "What's this going to cost per month at our actual
> load?"
> **SE:** "Uh, roughly… a few thousand?"
That's how pilots become lab graveyards.
This skill produces the answer in one command.
## What this skill does NOT replace
| Concern | Use instead |
|---|---|
| AOAI-only PAYG-vs-PTU break-even on a notebook (no deployed pilot) | `paygo-ptu-cost-analyzer` (awesome-gbb) |
| Static cost-pillar checks (Budget declared? anomaly alert? projection present?) | `threadlight-production-ready` pillar 10 |
| Live budget presence / anomaly posture on the target subscription or RG | `threadlight-production-ready` `COST-101` + cost-pillar budget/anomaly checks |
| Assessing the verified reconciliation bundle in a readiness report | `threadlight-production-ready` `COST-102`, `COST-103`, `KPI-003` *(artifact consumers only; no query/recompute)* |
| Bicep mutation from recommendations | `threadlight-deploy` on the next run (this skill is advisory) |
| Real-time anomaly detection | Azure-native budget/anomaly automation; this skill stays evidence-only |
| Demand forecasting / usage time-series | out of scope; foundry-observability owns the trace side |
## When to invoke
| You start with… | Phase | What's produced |
|---|---|---|
| Green `threadlight-safe-check --phase post-deploy` and an upcoming customer architecture review | `run --all` | `docs/cost-projection.md` + `specs/cost-manifest.json` (+ back-filled SPEC § 12 `load_profile{}` if wizard ran) |
| Pre-deploy spec review and you want to sanity-check the SKU choices in `infra/main.bicep` before you push | `run --all --pre-deploy` | same artefacts, marked `pre_deploy: true` in manifest (no `azd env` walk; Bicep-only) |
| Re-run with `load_profile{}` already populated and recent deploy | `run --all` (wizard auto-skips) | refreshed artefacts |
| You want to check just one resource (e.g. "did we pick the right APIM tier?") | `project --only Microsoft.ApiManagement/service` | partial artefacts, scoped manifest |
> **Rule of thumb.** Run this **after** `safe-check` and **before**
> `production-ready`. `threadlight-auto` will wire it in automatically.
## The chain (where this fits)
```
threadlight-design → threadlight-demo-data-factory → threadlight-local-test →
threadlight-deploy → threadlight-safe-check →
threadlight-consumption-iq ← THIS SKILL
foundry-evals + foundry-observability →
threadlight-production-ready
```
## Inputs (contracts)
| Input | Source | Required |
|---|---|---|
| `specs/manifest.json → deployment_manifest{}` | `threadlight-design` | yes — **or** derived from the export's `infra/` + `azure.yaml` in Kratos-export mode |
| `infra/main.bicep` + modules | repo | yes |
| `azd env get-values` | live azd env | yes (skip with `--pre-deploy` to read Bicep only) |
| `specs/SPEC.md § 11c` (tech-stack selectors) | `threadlight-design` | yes — not present in a Kratos export; resources come from `infra/` instead |
| `specs/SPEC.md § 12 → load_profile{}` (NEW sub-block) | this skill's wizard OR hand-authored | yes (skill writes it back if absent; in Kratos-export mode it writes to `use-cases/<x>/load-profile.yml` since there's no SPEC) |
| Recent Application Insights / `foundry-observability` traces | live monitor | optional (fidelity boost post-launch) |
### Kratos-export mode (discover from `infra/`, not from a SPEC)
For a **Kratos-exported project** (`src/hosted-agent/` + `use-cases/<x>/`,
trimmed `infra/` — see [`docs/KRATOS-BRIDGE.md`](../../docs/KRATOS-BRIDGE.md))
there is no `specs/SPEC.md` and no `specs/manifest.json`. Adapt the `discover`
phase:
- **Walk the export's `infra/` (`az bicep build` → ARM JSON) + `azd env
get-values`** to enumerate deployed resources, instead of reading a
`deployment_manifest{}`. This is the same ARM-walker path the skill already
uses — just sourced from the export's own Bicep.
- **Tolerate Kratos resource naming.** The trimmed Kratos infra names Cosmos /
Foundry / ACA resources differently than a `threadlight-design` deployment.
Match resources by **ARM type** (e.g. `Microsoft.DocumentDB/databaseAccounts`,
`Microsoft.App/containerApps`, `Microsoft.CognitiveServices/accounts`) and by
`azd env` output keys, **not** by hard-coded `threadlight-design` names. The
trimmed infra has **no APIM** — simply omit the APIM projector rather than
reporting it missing.
- **`load_profile{}` has no SPEC to write back to.** Run the wizard as usual,
but persist the result to `use-cases/<x>/load-profile.yml` (and still emit
`specs/cost-manifest.json` + `docs/cost-projection.md`). Idempotent on re-run.
### NEW: SPEC § 12 `load_profile{}` sub-block
Documented in `references/load-profile-schema.md`. Seven required
fields; the skill **fails fast** (no math, friendly error) if any are
blank after the wizard completes — we never produce a projection on
guessed numbers.
```yaml
load_profile:
workload_class: chat-agent | batch | scheduled | hybrid
peak_concurrent_sessions: 50
avg_requests_per_session: 8
avg_tokens_per_request: 1500
peak_requests_per_second: 12
business_hours_only: true
cosmos_gb_year_one: 50
storage_gb_year_one: 100
ai_search_documents: 50000
monthly_growth_rate: 0.15
declared_constraints:
max_p95_latency_ms: 2500
min_redundancy: zone-redundant
pinned_region: eastus2
```
## Outputs (contracts)
| Output | Consumer | Format |
|---|---|---|
| `docs/cost-projection.md` | humans + `threadlight-production-ready` `COST-005` | markdown: per-resource sections, side-by-side SKU tables, top-N recommendations, mermaid cost share donut |
| `specs/cost-manifest.json` | `threadlight-production-ready` `COST-005..007` + `threadlight-auto` resumability + downstream CI | strict v1 schema (see `references/cost-manifest-schema.md`) |
| `specs/cost-actuals-manifest.json` | reconciliation + audit trail | read-only settled-window capture (`threadlight-cost-actuals/v1`) |
| `specs/cost-reconciliation-manifest.json` | `threadlight-production-ready` `COST-102`, `COST-103`, `KPI-003` | strict reconciliation verdict bundle (`threadlight-cost-reconciliation/v1`) |
| `docs/cost-reconciliation.md` | humans + handoff review | markdown summary of scope/window, variance, coverage, unallocated cost, and unit economics |
| `specs/SPEC.md § 12 load_profile{}` | re-runs of this skill, future deploys, `threadlight-design` template | back-filled if wizard ran |
## Resource coverage matrix (v1)
| Azure resource | Compared variants |
|---|---|
| AOAI model deployments | PAYG ↔ PTU (1/4/10/25/50/100 units); region (eastus2 / sweden / etc.); model swap (gpt-4o ↔ gpt-4o-mini) |
| Foundry hosted-agent | tier ↑/↓ |
| Azure Container Apps | Consumption ↔ Dedicated D4/D8/E4/E8; min/max replicas |
| Cosmos DB (NoSQL) | provisioned (1k/4k/10k RU) ↔ serverless ↔ autoscale |
| Storage account | redundancy (LRS/ZRS/GRS); access tier (hot/cool/cold/archive) |
| APIM | Consumption ↔ Basic v2 ↔ Standard v2 ↔ Premium |
| Azure AI Search | Free / Basic / S1 / S2 / S3; replica × partition combos |
**Out of scope for v1:** Reservations / Savings Plans, EA / MCA discounts,
automatic Bicep mutation, multi-region failover cost, spot ACA pricing,
cross-cloud comparison.
## Phases
| Phase | Script | What it does | Resumable from |
|---|---|---|---|
| 1. discover | `discover.py` | Walk Bicep + `azd env` → list of normalized resource selectors | always |
| 2. load-profile | `load_profile_wizard.py` | Read SPEC § 12; if missing, run interactive prompt; write back to SPEC | always (idempotent) |
| 3. price | `pricing_client.py` | Hit `Azure-pricing` MCP for each SKU + 2–3 alternatives per resource | cache to `.threadlight/cost-cache.json` (TTL 24h) |
| 4. project | `projectors/<resource>.py` | Apply per-resource consumption formulas to load_profile | always |
| 5. compare | `projectors/<resource>.py` | Build per-resource alternative comparisons | always |
| 6. recommend | `recommender.py` | Score alternatives vs `declared_constraints`; rank by $/mo savings | always |
| 7. emit | `emitter.py` | Write `docs/cost-projection.md` + `specs/cost-manifest.json` | always |
## CLI surface
```bash
# Individual phases
scripts/consumption_iq.py discover
scripts/consumption_iq.py load-profile [--non-interactive]
scripts/consumption_iq.py price
scripts/consumption_iq.py project [--only <resource_kind>]
scripts/consumption_iq.py recommend
scripts/consumption_iq.py emit
# Chained
scripts/consumption_iq.py run --all
scripts/consumption_iq.py run --all --pre-deploy # skip azd env walk; Bicep-only
# Opt-in live actuals (read-only; never reached without these commands/flags)
scripts/consumption_iq.py actuals \
--start 2026-08-01 --end 2026-08-08 \
--subscription <sub-id> --resource-group <rg> \
--monitor-resource-id /subscriptions/<sub-id>/resourceGroups/<rg>/providers/Microsoft.CognitiveServices/accounts/<account> \
--workspace-resource-id /subscriptions/<sub-id>/resourceGroups/<rg>/providers/Microsoft.OperationalInsights/workspaces/<ws>
scripts/consumption_iq.py reconcile # local documents only, no Azure calls
scripts/consumption_iq.py reconcile --expect-subscription <sub-id> --expect-resource-group <rg>
scripts/consumption_iq.py run --all --with-actuals --start 2026-08-01 --end 2026-08-08
```
### Live actuals + reconciliation (opt-in)
`run --all` is unchanged and **offline**: it reads Bicep/SPEC and the public
Retail Prices API and never contacts a customer subscription. Live evidence is
reached only by naming `actuals` / `reconcile`, or by adding `--with-actuals`.
| Command | Reads Azure? | Writes |
|---|---|---|
| `actuals` | yes (read-only) | `specs/cost-actuals-manifest.json` |
| `reconcile` | **no** — not one call | `specs/cost-reconciliation-manifest.json`, `docs/cost-reconciliation.md`, `specs/cost-history/` |
| `run --all --with-actuals` | yes (read-only) | forecast artefacts **plus** all of the above |
**What `reconcile` trusts.** `reconcile` re-projects local documents and
issues **no Azure call at all**, so it takes the scope and window recorded in
the actuals manifest you point it at as given. Standalone `reconcile` still has
no `--start`/`--end`, so it does not re-collect or re-verify the window from
Azure, and `actuals_ref.sha256` still pins the exact evidence bytes you chose.
If you want a local scope guardrail, pass `--expect-subscription` and/or
`--expect-resource-group`: the command compares the recorded manifest scope to
those expected values before reconciling and exits `2` on mismatch. It
publishes the chosen subscription, resource group and window under
**Collection scope** either way, so a reviewer can see exactly what the offline
verdict rests on. (Under `run --all --with-actuals` the collection flags do
exist, and a manifest whose scope or window disagrees with them stops the run.)
**Flags.** `--start` / `--end` are `YYYY-MM-DD` (`--end` exclusive, must be
after `--start`). `--subscription` defaults to `AZURE_SUBSCRIPTION_ID`,
`--resource-group` to `AZURE_RESOURCE_GROUP`; if neither is resolvable the
command exits `2` **before** any projection or network call.
`--monitor-resource-id` and `--workspace-resource-id` are full **ARM resource
IDs** — the AI/Cognitive account to read token metrics from, and the Log
Analytics workspace holding traces. They are never a customer/tenant GUID, and
no customer GUID is ever required or recorded. Both are optional: without them
the token and interaction rows degrade to a warning and the cost evidence is
still collected. Sidecar paths: `--actuals-manifest`,
Ver no GitHub