- name
- threadlight-production-ready
- description
- Use when a pilot needs production-readiness review, CISO or customer architecture handoff, current runtime-binding evidence, missing live proof, remediation ownership or an explicit CI readiness gate. Not for deployment, runtime implementation, model evaluation or hub provisioning.
- metadata
- {"version":"0.15.1"}
# Threadlight Production Ready — paving the path to production
## Enterprise handoff: three areas
The report opens with a presentation-only view of existing findings:
| Area | Enterprise baseline to assess | Optional / conditional modules |
|---|---|---|
| DevOps | Private network boundary, dedicated workload identity, least privilege, protected secrets, reproducible release and rollback | Citadel when selected, landing-zone integration, private-runner provisioning, multi-region |
| Runtime governance | Explicit tools, input validation, ordinary backend authorization and code execution disabled by default in the proposed configuration | SAFE/ACS/Agent Hooks for selected tools, authenticated HITL/resume, OBO where required, grounding/quality/authority assurance |
| Operations | Owner/escalation, health/telemetry, delivered alerts, runbook/recovery, cost/budget ownership, restore where state requires it | Azure SRE Agent, continuous eval, recurring red-team, advanced load tests, PTU optimization, dashboards and upgrade automation |
These are requirements to assess, not controls installed by rendering a report.
Explicit opt-in selects advanced modules; a domain, template or finding is not
consent. A module required by the selected business process remains mandatory.
Rapid prototyping stays unchanged. No new score, gate, manifest or prototype stage
is introduced, and existing selected-binding gates are not relaxed.
The view groups the current manifest's emitted findings exactly once, preserving
their statuses and experimental selection. Missing/N/A findings and a lack of
open findings never establish baseline acceptance. Detailed pillar, evidence,
waiver and KPI sections remain available; legacy scores are advisory, not proof
that the enterprise baseline is complete. Implementation and deployment still
require an approved specification and explicit action.
## Optional AgentOps operations evidence
`AOPS-001` is an additive, tier-0 `sre-handover` aggregate. Per-agent
`agentops.yaml` is the only opt-in: without it the finding is not-applicable
and non-scoring; opt-in without valid evidence is not-verified. Consume only
`specs/agentops-manifest.json` through the independent strict shared contract
validator (freshness, binding, artifact hashes, all roots, statuses and counts),
never native JSON or the producer implementation. Aggregate the worst
operational status across opted-in agents.
Canonical eval, red-team and binding-scoped governance checks retain ownership.
Exclude a domain blocker from AOPS only when its exact source is represented
by current validated canonical evidence; a claimed mapping alone is not proof.
Unknown operational blockers and worse conflicts remain. Remediate adoption
through the pinned merged `foundry-agentops` sibling, then run
`threadlight-agentops`; this is not certification or a tagged-release claim.
No paid probes, deployment, new secrets, RBAC expansion or Citadel writes occur
as a side effect of assessment.
## Current runtime-binding contract
Consume `specs/governance-manifest.json` (`threadlight-governance-manifest/v1`)
through the shared strict validator and `governance_readiness.assess`, also used
by evidence_gate and Auto. This is **not whole-agent** governance. SAFE is the
method; ACS/Rego is the PDP; Agent Hooks is the host/interceptor contract; the
native host/gateway is the PEP; AGT is the toolkit; ASSERT is assurance.
Selected bindings need current live evidence for their exact tool, point and
path. Preserve unbound reads without ACS. Consequential unbound actions require
explicit current acceptance bound to tool, risk, source commit, environment and
owner; a boolean does not suffice. Invalid selected configuration is not off.
Legacy v2 green is provenance only, never a readiness fallback.
The current collector certifies only `governance_probe_noop` allow/deny and
supported signed-policy/audit requirements. Business bindings, other lifecycle
points and unproved approval requirements stay unverified. Do not force
`gaps: []`, a whole-agent `governed` verdict or SAFE-complete from local tests.
The canonical returns business write has local native proof but no live proof.
Full signed envelope/key, bundle, configuration and observed target must still
match the verified record. Re-signing, expiry or any selected config change
invalidates that evidence. A new deploy attempt needs fresh after-deployment
collection, not file mtime. The ordinary scorecard remains advisory; explicit
`readiness-proof` CI is strict and fails on missing evidence. Route runtime
implementation to `threadlight-govern`/`threadlight-governed-actions` real
generators rather than just recommending another assessment.
> **v0.12.0 — advisory assessment + explicit remediation workflow.** The
> current contract is simple: the assessment phase is always read-only. It
> inventories evidence, scores 13 pillars, and writes the scorecard/report
> without patching the repo, querying cost actuals on its own, or deploying
> anything. Repo remediation, CI scaffold generation, and any downstream
> deployment require an explicit operator action (`--onboard`,
> `--scaffold-cicd`, or a follow-on remediation/deploy skill); none of those
> happen as a side effect of a plain assessment run.
The single skill in the chain that asks "**is this pilot ready for the
customer architecture review, or is it about to land in the lab graveyard?**"
and answers with a structured, evidence-backed artefact instead of tribal
knowledge.
> **Why this skill exists.** The `threadlight-*` chain ships a working
> agent in one session (design → local-test → deploy → safe-check).
> `safe-check` proves the pilot is **structurally complete and behaves**:
> every selector landed, every channel reaches, every cron ran, no
> placeholder image. But "green safe-check" ≠ "production-ready". The
> next conversation — CISO, SRE, FinOps, network architect, data
> protection — needs an artefact that says: **what posture is this in,
> what's missing, what would the uplift cost, who owns each gap, can we
> go live with waivers?** Without that artefact every pilot grows a
> tribal-knowledge answer that takes weeks to assemble, the customer
> defers the production phase, and the pilot quietly becomes a "lab
> graveyard" demo.
>
> This skill produces the artefact in one command.
## What this skill does NOT replace
| Concern | Use instead |
|---|---|
| Authoring SPEC / `deployment_manifest{}` | `threadlight-design` |
| Running `azd up` | `threadlight-deploy` |
| Structural / behavioural deploy gate | `threadlight-safe-check --phase post-deploy` |
| Invocation testing of the agent | `foundry-evals` |
| Wiring App Insights / OTel | `foundry-observability` |
| Provisioning Citadel hub | `citadel-hub-deploy` |
| Onboarding spoke to Citadel | `citadel-spoke-onboarding` |
| Provisioning Azure SRE Agent | `azure-sre-agent` |
| Authoring / linting the AGT policy | `foundry-agt` |
| Generating Bicep / Terraform | `azd-patterns`, `azureterraform`, `bicepschema` |
| Deploying to a VNet-injected Foundry | `foundry-vnet-deploy` |
**This skill separates assessment from follow-on actions.** The scorecard is
always advisory and read-only. The remediation workflow is explicit, reviewable,
and opt-in.
## What this skill does in v0.12.0
```mermaid
flowchart LR
A[Phase 1: Assess<br/>production_ready.py] -->|apply-plan.json| B[Phase 2: Refine+Deploy<br/>agent + Edit/Write + sibling skills]
B -->|repo commits<br/>+ pipeline deferrals| C[Phase 3: CI/CD Handoff<br/>azd-deploy-prod.yml<br/>+ UAMI readme]
C -.->|post-deploy| A
```
The skill exposes an **explicit production-onboarding workflow**:
1. **Assess (always safe, always read-only).** A Python script
(`scripts/production_ready.py`) inventories your target Azure
subscription/resource group, scores it against 192 findings spanning 13
production-readiness pillars, and emits the scorecard/report plus an
`apply-plan.json` that names every must-fix gap and the remediation recipe
that closes it.
2. **Refine + remediate (explicit action).** Only when you invoke an explicit
follow-up (`--onboard`, or ask the agent to execute the apply-plan) does the
Copilot agent read `apply-plan.json`, open the named recipes, and prepare
repo edits / sibling-skill follow-ups. Items marked `kind: manual` are
surfaced to you for acknowledgement before any change. Items marked
`kind: deferred-to-pipeline` are recorded for the CI/CD handoff.
**Stale-plan detection (the agent MUST do this before applying anything).**
Each `apply-plan.json` carries a `manifest_sha256` field — the SHA256 of
the canonical JSON of the `production-readiness-manifest.json` it was
built from. Before executing the first item, the agent recomputes
`sha256(canonical_json(<current production-readiness-manifest.json>))`
and compares. If the two hashes differ, the manifest has been
re-generated since the plan was emitted (e.g. the operator re-ran
`--onboard` in another shell, or hand-edited the manifest) and the plan
is stale: **the agent refuses to apply and tells the operator to re-run
the Assess phase to get a fresh apply-plan.** Skipping this check risks
applying remediations for findings that no longer exist or missing
findings that now do — silent drift between the plan and reality is the
single failure mode this gate exists to prevent.
3. **CI/CD handoff (explicit action).** `--scaffold-cicd` delegates to the
authoritative **`threadlight-cicd` verified-release generator** for GitHub
Actions or Azure DevOps. Pipeline-deferred findings alone only emit a hint;
they do not create a pipeline. Both entrypoints emit the same executable
tooling, non-executable policy example, separate validation/production setup
runbooks and central-platform boundary. The legacy
`docs/threadlight-cicd/central-team-uami-readme.md` path is a handoff pointer,
not an alternate deployment implementation.
Follow the [release contract](../threadlight-cicd/references/release-contract.md):
configure reviewed application adapters, isolated targets/identities and
protected CI environments before enabling execution. Validation runs the
actual required producers; promotion verifies the receipt and reuses the
immutable image. Generating YAML grants no deployment or native-operation
authority, and the scorecard does not certify a release.
4. **Deployment (explicit action outside this assessment).** Production Ready
never runs `azd up` for you. Deployment happens later through the generated
pipeline or a downstream deployment skill/session.
**The Python script is assessor-only for remediation findings.** It never
mutates your repo or subscription for findings — fixes are dispatched to the
agent as apply-plan tasks. The explicit `--scaffold-cicd` exception writes the
selected pipeline, portable release tooling
(`.threadlight/skills/threadlight-cicd/scripts/release_runner.py`), policy example
(`specs/release-policy.example.json`), private-output Git exclusions and runbooks
(`docs/threadlight-cicd/release-contract.md`) into the customer repo. It does not enable the
workflow, generate passing evidence, execute adapters or apply remediation.
## Assessment-to-plan handoff
Current `must-fix`, `should-fix`, and `not-verified` findings are preserved
in the apply-plan from native pillar output or a flat findings list.
Legacy `fail` and `warn` remain supported. `pass`, `not-applicable`, and
`waived` findings are not proposed for remediation.
The plan retains source order and the full source-manifest hash.
Restricted environments retain the existing manual-handoff behavior.
A plan is a proposal, not approval to edit, provision or deploy.
This status compatibility fix does not change the onboarding CLI's input
contract, readiness scoring, governance selection, or prototype workflow.
## How to invoke
First set and verify `CATALOG` and `PROJECT` using the [CLI setup](#cli).
Run these examples from the pilot working directory: onboarding/scaffolding
resolve relative framing, apply-plan and scaffold paths from that directory.
Live reads still require approved identity and target scope.
```bash
: "${CATALOG:?Set and verify CATALOG using CLI setup}"
: "${PROJECT:?Set PROJECT to the existing pilot root}"
cd "$PROJECT" || exit 1
```
### Quick assess (no changes)
```
python3 "$CATALOG/skills/threadlight-production-ready/scripts/production_ready.py" --root "$PROJECT" \
--target-sub <SUB> --target-rg <RG>
```
### Full onboarding (interactive framing wizard)
```
python3 "$CATALOG/skills/threadlight-production-ready/scripts/production_ready.py" --root "$PROJECT" --onboard
```
### Headless / CI-friendly
```
python3 "$CATALOG/skills/threadlight-production-ready/scripts/production_ready.py" --root "$PROJECT" --onboard \
--framing-file framing.json \
--apply-plan-out apply-plan.json \
--no-rights-probe
```
### Phase 3 scaffold only
```
python3 "$CATALOG/skills/threadlight-production-ready/scripts/production_ready.py" --root "$PROJECT" \
--framing-file framing.json \
--scaffold-cicd \
--repo-full-name owner/repo
```
## Framing wizard questions
`--onboard` asks these 8 questions; `--framing-file` uses the same IDs.
The IDs and prompts are sourced from `FRAMING_QUESTIONS` in
`scripts/production_ready.py`.
| # | ID | Description |
|---|---|---|
| 1 | `target_subscription_id` | Azure subscription ID for the production target. |
| 2 | `target_resource_group` | Resource group name for the production target. |
| 3 | `target_posture` | Posture profile: `citadel-spoke`, `standard-ai-gateway`, `agt`, or `hybrid`. |
| 4 | `provisioning_rights` | Whether the operator has Contributor-or-higher rights on the target resource group. |
| 5 | `central_platform_team` | Whether a central platform/Citadel team owns shared gateways, Key Vault, or networking. |
| 6 | `restricted_environment` | Whether direct writes are restricted and all changes must go through CI/CD. |
| 7 | `cicd_target` | CI/CD target; `github-actions` is the only supported value. |
| 8 | `azure_tenant_id` | Azure tenant ID (UUID) where the production subscription lives. |
## Remediation recipes
Every must-fix finding has a recipe at `references/remediation-recipes/{FINDING_ID}.md`.
Recipes declare `kind: repo-edit | sibling-skill | manual | deferred-to-pipeline`
in their YAML front-matter; the apply-plan inherits this `kind` field so the
agent knows whether to edit a file, invoke a sibling skill, or surface a
prompt to the operator. See `references/remediation-recipes/_template.md`
for the schema, and `references/sibling-skills-map.md` for the awesome-gbb
skills threadlight delegates to.
## When to invoke
| You are at… | Run | Get |
|---|---|---|
| `safe-check --phase post-deploy` returned green and the customer wants to talk about production | Pinned complete-catalog CLI below, live reads only after approval | Markdown report + JSON manifest |
| Customer architecture review in 3 days | Add `--target citadel-spoke` (or other posture) to that CLI | Same as above, scored against the declared target |
| Pilot has been parked for weeks; someone asks "could we ship this?" | Use `--static --no-rights-probe`; retain valid fresh safe-check prerequisites | Static scorecard from repo + safe-check manifests |
| You inherited a pilot whose SPEC has no § 12 | Skill still runs — posture falls back to `standard-ai-gateway`, an `RDY-002` warning surfaces "SPEC § 12 missing — add it from `references/spec-section-12-template.md`" | Author § 12, re-run for full scorecard |
| Historical AGT v4 compatibility inspection | Add `--pillar agent-governance --agt-profile v4_preview` | Legacy diagnostics only; never substitutes for current v1 runtime-binding readiness |
| Customer accepted some `must-fix` findings as risk | Author `tests/production-readiness-waivers.json`, re-run | Report shows `score_with_waivers` and `would_fail_hard_gate` flags |
> **Rule of thumb.** This skill runs at most twice per pilot
> lifecycle: once when the pilot is heading into the customer
> architecture review (the artefact that lives in the deck), and once
> immediately before the go-live decision (the artefact that goes to
> CISO / the change advisory board). Running it every commit is noise.
## The 13 pillars
Each pillar gets its own reference doc under `references/pillars/`; the
skill ships with prose-heavy guidance per pillar so the LLM can reason
about findings, not just emit them.
| # | Pillar | What "good" looks like | Primary remediation skill |
|---|---|---|---|
| 1 | [`network-posture`](references/pillars/01-network-posture.md) | Resolved posture target met (Citadel spoke / AGT / VNet / standard); data-residency considered (model region, APIM region, backup region, cross-border support) — declarative SPEC check, not sub-scored | `citadel-spoke-onboarding`, `foundry-vnet-deploy`, `foundry-network-runbook` |
| 2 | [`agent-governance`](references/pillars/02-agent-governance.md) | Current selected-binding v1 evidence and exact signed policy/config/deployment; module imports and legacy policy/verifier artifacts alone are not enforcement | `threadlight-govern`, `threadlight-governed-actions` |
| 3 | [`identity-access`](references/pillars/03-identity-access.md) | Workloads use managed identity; **no client secrets**; RBAC least-privilege; KV access via RBAC not access policies | `foundry-hosted-agents`, `azure-tenant-isolation`, `azd-patterns` |
| 4 | [`secrets`](references/pillars/04-secrets.md) | Key Vault with **soft-delete + purge protection**; no hardcoded secrets in repo; rotation policy declared; control-plane vs data-plane access scoped | `azd-patterns`, `foundry-hosted-agents` |
| 5 | [`observability`](references/pillars/05-observability.md) | App Insights connected at **account-level** (Foundry); OTel emit verified (recent traces); alert rules wired; workbook + retention declared | `foundry-observability` |
| 6 | [`continuous-evals`](references/pillars/06-continuous-evals.md) | SPEC § 9 scenarios scheduled (Plan A or Plan B); threshold alerts wired; last run within freshness window; eval datasets stored | `foundry-evals` |
| 7 | [`responsible-ai`](references/pillars/07-responsible-ai.md) | Content filters, jailbreak shields, grounded-language eval; AGT RAI policy; PII redaction declared; allow/deny tested | `foundry-agt`, `foundry-evals` |
| 8 | [`hitl-audit`](references/pillars/08-hitl-audit.md) | If SPEC § 8 declares gates: wired, persistent audit trail, escalation channel reachable, idempotent | `threadlight-hitl-patterns` |
عرض على GitHub