- name
- threadlight-cicd
- description
- Use when a Threadlight pilot must reach production through a CI/CD pipeline instead of the agent running `azd up` directly — prod deploys run under a federated identity with scoped RBAC from private-VNet runners and the agent has no standing deploy rights. Generates a GitHub Actions or Azure DevOps prod-deploy pipeline + env-setup runbooks, with an onboarding-path gate and a central-platform boundary vs citadel-hub-deploy. USE FOR: ci/cd pipeline, prod deploy pipeline, github actions pipeline, azure devops pipeline, OIDC, workload identity federation, WIF, UAMI, scoped RBAC, self-hosted runners, private vnet runners, restricted environment deploy, platform team handoff, onboarding path gate. DO NOT USE FOR: deploying the central Citadel hub / shared AI gateway (use citadel-hub-deploy); wiring a pilot to an existing hub (use citadel-spoke-onboarding); the first-run sandbox deploy (use threadlight-deploy).
- metadata
- {"version":"0.6.0"}
# Threadlight CI/CD — verified release + environment setup
> The skill that answers "**how does this pilot actually deploy to production
> when the agent can't run `azd up` and has no standing rights?**" — by
> validating a separate preproduction candidate, then promoting its immutable
> image through a protected production job. The federated-identity pipeline and
> runbooks define the separation. The customer's platform team configures the
> identity, permissions, and
> runners. Secret-free by construction; parallel-track-safe by design.
## When to use
- The pilot's prod environment is **locked down**: direct writes are restricted,
deploys must go through a pipeline, and the agent has no deploy rights.
- You need a **GitHub Actions or Azure DevOps** prod-deploy pipeline that logs in
with **OIDC / Workload Identity Federation** (no `AZURE_CREDENTIALS`, no client
secret, no PAT).
- The platform team needs **ready-to-run `az` runbooks** for the UAMI, federated
credentials, least-privilege RBAC, and (for private VNets) the runners.
- You must make the **central-platform boundary** explicit so the pilot pipeline
never touches the Citadel hub.
**When NOT to use:** deploying the hub itself (`citadel-hub-deploy`), onboarding a
spoke onto an existing hub (`citadel-spoke-onboarding`), the permissive first-run
deploy (`threadlight-deploy`), or the readiness scorecard
(`threadlight-production-ready`). See the description's DO NOT USE FOR list.
## Onboarding-path decision gate (runs FIRST)
Before generating anything, resolve which track the pilot is on. `resolve_onboarding_path()`
branches on two questions and is the source of truth for posture + RBAC scope:
```
Is a central platform env required? (Citadel hub / shared AI gateway /
shared networking / platform Key Vault)
│
├─ no ──────────────────► standalone (validate target sub/RG, shared-resource
│ usage, network exposure FIRST)
│ posture: standard-ai-gateway | agt | direct
│ RBAC scope: target-rg
│
└─ yes ─► already deployed?
│
├─ yes ─────────► spoke-onboard (consume hub via Access Contract →
│ citadel-spoke-onboarding)
│ posture: citadel-spoke · RBAC scope: spoke-rg
│
└─ no ─────────► hub-deploy-then-spoke
(stand up hub on the SEPARATE central
track → citadel-hub-deploy, THEN
citadel-spoke-onboarding)
posture: citadel-spoke · RBAC scope: spoke-rg
```
**Invariant:** a spoke pilot's deploy identity is scoped to the **spoke resource
group only** — never the hub, regardless of whether the hub already exists. The
resolved decision is written to `docs/threadlight-cicd/onboarding-path.json` so
the choice is auditable.
## Parallel-track boundary (the must-tell)
The pilot pipeline is a **separate repo and pipeline** from the central platform.
| Concern | Owner | Track / skill |
|---|---|---|
| Citadel hub, shared APIM AI gateway, shared networking, platform Key Vault | Central platform team | **`citadel-hub-deploy`** (awesome-gbb, separate repo) |
| Wire the pilot to consume the hub via an Access Contract | Platform / SE | **`citadel-spoke-onboarding`** |
| Deploy the pilot's **use-case** resources into the spoke/target RG | This pipeline | **`threadlight-cicd`** |
The generated `central-platform-boundary.md` states the pilot pipeline **must not**
deploy or modify the hub, and that its UAMI RBAC is **spoke-RG-scoped only**.
```mermaid
flowchart LR
subgraph central["Central platform repo (separate)"]
HUB["citadel-hub-deploy<br/>hub · shared APIM · networking · KV"]
end
subgraph pilot["Pilot repo (this pipeline)"]
PIPE["threadlight-cicd<br/>azd-deploy-prod → spoke/target RG only"]
end
HUB -. "Access Contract<br/>(citadel-spoke-onboarding)" .-> PIPE
PIPE -. "never writes" .-x HUB
```
## Quick reference
| Goal | Command |
|---|---|
| Interactive onboarding-path gate + generate | `python scripts/generate_pipeline.py --onboard` |
| GitHub Actions, standalone (public) | `python scripts/generate_pipeline.py --platform github-actions --central-env-required no --repo-full-name owner/repo --target-sub <sub> --target-rg <rg> --tenant-id <tid>` |
| Azure DevOps, spoke onto existing hub | `python scripts/generate_pipeline.py --platform azure-devops --central-env-required yes --central-env-exists yes --ado-org <org> --ado-project <proj> --ado-service-connection <sc> --target-sub <sub> --target-rg <rg> --tenant-id <tid> --hub-sub <hsub> --hub-apim-id <apim-id> --access-contract-product <product>` |
| Private-VNet target (self-hosted / managed pool) | add `--private-network` (and `--ado-pool-name <pool>` for ADO) |
| Validation target | add `--validation-env-name`, `--validation-sub`, `--validation-rg` and `--validation-client-id` or `--validation-service-connection` |
| Required release checks | Always blocking; legacy `--eval-gate hard` / `--mcp-gate hard` remain accepted, but soft or invalid modes fail |
| Approved producer/promotion contract | `--release-policy specs/release-policy.json`; see [release contract](references/release-contract.md) |
| Selected two-tool returns reference | `--reference-application returns-mcp/v1`; emits concrete adapter composition and a deliberately incomplete operator configuration |
| Optional AgentOps integration | `--agentops auto` (default; no opt-in means no pipeline change) or `--agentops off` |
| Explicit post-deploy Doctor opt-in | add `--agentops-refresh-doctor`; runtime owner approval and existing application telemetry scope are still required |
| Doctor-only schedule | add `--agentops-refresh-doctor --agentops-doctor-schedule "0 6 * * *"`; never schedules deployment |
| From a saved framing file | `--framing-file framing.json` |
| Run the test suite | `python -m pytest tests/ -v` |
> If `python` resolves to Python 2 on your machine, use `python3` (the generator
> renderer uses Python 3 stdlib; optional AgentOps authoritative service
> discovery uses the shared contract and PyYAML).
**Spoke flags:** `--hub-sub` / `--hub-apim-id` / `--access-contract-product` surface
the hub coordinates the platform team needs to wire the Access Contract; they are
echoed into `central-platform-boundary.md` and the runbooks (never used to deploy
the hub).
## What it emits
Rendered deterministically (offline, no Azure calls, no secrets) into the pilot repo:
- **Pipeline** — `.github/workflows/azd-deploy-prod.yml` or `azure-pipelines.yml`.
`validation` prepares an isolated candidate, executes configured real
producers and the supplied MCP checker, then enforces required domains.
`promotion` depends on that success and uses a different protected
environment and deployment identity. It verifies the receipt SHA-256 supplied
separately by the validation job, source/run/attempt, policy and input hashes.
- **Executable release tooling** — `release_runner.py`, the shared strict
canonical evidence consumer and actual MCP producer. The runner invokes the
reviewed application's deployment/evaluation/scanner adapters, not `echo`
instructions. Missing adapters, authority or current evidence are blockers.
The metadata receipt is not a business-write proof or readiness certificate.
- **Policy example** — `specs/release-policy.example.json`, deliberately separate
from the executable policy. Configure and commit its real target, producer,
observer, input and threshold contract before enabling the workflow.
- **Env-setup runbooks** — `docs/threadlight-cicd/env-setup/` for production and
`docs/threadlight-cicd/validation-env-setup/` for validation:
- `01-uami-federated-credentials.md` + `.sh` (UAMI + GH OIDC or ADO WIF — no secrets)
- `02-rbac-role-assignments.md` + `.sh` (target-RG-scoped: deploy role **plus**
*Role Based Access Control Administrator* so keyless `azd provision` can assign
the app identity's data-plane roles; ensures the target RG exists)
- `03-runners-private-vnet.md` + `.sh` (managed **and** self-hosted options, with
subnet/egress/private-DNS prerequisites)
- `README.md` (what to hand the dev team vs the platform team)
- **Boundary + decision record** — `central-platform-boundary.md`, `onboarding-path.json`.
Public targets default to hosted runners (`ubuntu-latest` / ADO `vmImage`); private
targets switch to `self-hosted` labels / a named ADO pool.
### Required-domain acceptance
Quality checks compare actual metrics with approved thresholds. Red teaming
requires fresh, sufficiently sized evidence for every required attack category.
MCP checks execute `mcp_sbom.py` and reject missing or contradictory counts,
unresolved findings and lock drift. Neither an empty object nor an aggregate
`partial` verdict grants release acceptance. Optional evaluation capabilities
can be scoped explicitly in reviewed policy; core execution, dataset, freshness
and threshold checks cannot be disabled.
All producer evidence is bound to the current source, observed candidate,
inputs and CI attempt. A failed gate never starts the production job. A failed
promotion is not rollback: the deployment adapter owns immutable-image use,
durable idempotency, business-routing closure and reconciliation.
### Selected returns reference
`--reference-application returns-mcp/v1` composes the existing runner with
`returns_release.py`, canonical evaluator/scanner assessors and the actual
`govern_control_plane.operator` admission helper. It does not generate another
workflow or evaluator. The observed application contract binds model deployment/
version, two tool schemas, policy/binding, connection and service revisions,
role map, native Outlook authority and backend Cosmos targets—not only the image.
Validation and production retain separate identities, resources and approvals.
Real reviewed prepare/observe/producer/traffic/promotion/recovery operators are
required; the example configuration cannot execute as supplied. The production
adapter closes and reads admission, reconciles the durable promotion record
without retrying unknown outcomes, then prepares the exact accepted image.
Only fresh postchecks, unchanged independent observation and rechecked release
authorization may admit. Explicit `release_runner.py reconcile --operation-id`
closes and reads an uncertain operation; it never retries promotion or opens
admission. See the
[selected-reference handoff](../../docs/reference-release.md) for exact
interfaces, native operator lease, human approvals and remaining live-integration
limits. Local subprocess/wire tests are not hosted, Outlook, Cosmos or business
proof. Native AgentOps opt-in and pins remain unchanged.
### `--agentops auto|off`
Discover opted-in roots from the **target output repository**, using the shared
AgentOps discovery contract. The only opt-in is each agent's `agentops.yaml`;
there is no second agent registry. `auto` without opt-in and `off` preserve
the existing workflows and do not emit AgentOps tooling.
With opt-in, retain native operations in the same generated workflow:
- PRs run the bound AgentOps eval once, including an explicitly bound baseline
comparison when a baseline exists. They never deploy or promote a baseline.
The PR job uses the validation environment, never the production deploy
identity. Untrusted fork PRs do not receive the GitHub execution context.
Native jobs use a separately prepared private context and do not perform
deployment login.
- Canonical `evals_check.py --target . --emit` consumes the existing validated
batch; it does not execute a second evaluation. A native negative exit is
preserved after consumption and fails the PR check rather than being
silently downgraded.
- Post-deploy Doctor is **off by default**. Enabling it requires both the
generator option and explicit current runtime owner approval for the existing
application's identity, target, environment and telemetry scope. It runs
after successful promotion, not as a second release evaluation.
- An optional Doctor-only schedule skips provisioning, deployment and paid
eval. It uses the same approved environment, private runner selection and
explicit native credential context. ADO native operations are environment-bound deployment jobs
so configured environment approval checks still apply.
The generator vendors the minimal tooling under `.threadlight/skills/`
to make the emitted commands executable in the application repository. Review
and commit this tooling and the existing workflow before enabling runtime
execution. Prepare the exact native `agentops-accelerator==0.14.0` runtime,
PyYAML for bounded authoritative `azure.yaml` service discovery,
committed binding policy and scoped owner approval using `foundry-agentops` /
`threadlight-agentops`. The packaged native observer records and rechecks an
actual same-process receipt without requiring new PKI or an external signer; absent
prerequisites are an actionable failure, not a green placeholder. Generated
runtime evidence must be ignored by Git while committed policy and tooling
remain tracked, so capture does not invalidate the clean source binding.
Supply `THREADLIGHT_AGENTOPS_VALIDATION_CONTEXT` and, when selected,
`THREADLIGHT_AGENTOPS_PRODUCTION_CONTEXT` through existing approved runner
preparation. They point to private files passed as `THREADLIGHT_AGENTOPS_CONTEXT`;
the generator neither creates them nor authorizes native execution. Production
Doctor needs its own verified same-target eval receipt, not the validation
candidate's receipt. See [native runtime prerequisites](references/agentops-runtime.md).
Do not use `agentops workflow generate`, create another workflow, install a new
identity, widen RBAC, provision telemetry or modify Citadel. No keys or storage
resources are generated. Existing application-scoped authentication must already
authorize the requested native operation.
**Native-output privacy:** unset `GITHUB_STEP_SUMMARY`, disable shell tracing,
use bounded private capture and clean raw outputs. Never upload native `.agentops`
artifacts, raw eval/Doctor logs, or public step summaries. Only the normalized
allowlisted manifest may be published; the default release artifact contains
only candidate metadata and hashes, not native outputs. Raw retention needs
separate explicit owner-approved location,
permissions and retention scope—not invented secrets or infrastructure.
## Relationship to threadlight-production-ready
`threadlight-production-ready` Phase 3 (`--scaffold-cicd`) delegates to this
authoritative generator for both platforms, including portable tooling, policy
example, separate environment setup and the central-platform boundary. There is
no weaker deploy-first alternate template. The compatibility runbook path is only
a pointer to the shared handoff; a green readiness scorecard is not release authority.
## Generator API (for tests / automation)
`scripts/generate_pipeline.py` exposes:
- `resolve_onboarding_path(framing) -> dict` — the decision gate (path, posture,
rbac_scope, needs_validation, next_actions).
- `build_context(framing, resolved) -> dict` — template token context.
- `generate(framing, out_root) -> list[Path]` — render + write artifacts.
- `VERSION` — semver, matched against this file's `metadata.version` by `test_version.py`.
Templates live under `references/` as `{{TOKEN}}` files rendered by `_render` (pure
stdlib). Tests under `tests/` pin the artifact paths, the OIDC/WIF-only invariant
(no long-lived secrets), and the boundary content.
## Common mistakes
- **Widening RBAC scope.** Never scope the deploy role to the subscription or a
central-platform RG. Target RG only.
- **Reaching for a secret.** If you find yourself adding `AZURE_CREDENTIALS` or a
client secret, stop — use OIDC/WIF. The test suite fails the build if a secret
or PAT lands in any emitted file.
- **Letting the pilot deploy the hub.** A missing central env is stood up on the
`citadel-hub-deploy` track, not by this pipeline.
- **Skipping the gate.** Generating before resolving the onboarding path produces
the wrong posture and RBAC scope. Run `--onboard` (or pass the flags) first.
## References
- [`references/onboarding-path-decision.md`](references/onboarding-path-decision.md) —
the decision tree (standalone vs spoke-onboard vs hub-deploy-then-spoke) and when to
engage `citadel-hub-deploy` vs `citadel-spoke-onboarding`.
- [`references/best-practices.md`](references/best-practices.md) — OIDC/WIF federation,
least-privilege RBAC, environment gates, and private-VNet runners, with Microsoft
Learn citations.
- [`references/pipeline-design-checklist.md`](references/pipeline-design-checklist.md) —
operator hand-off checklist before giving a pipeline to the customer.
- `references/github-actions/`, `references/azure-devops/`, `references/env-setup/` —
the `{{TOKEN}}` templates the generator renders.
## See also — official Azure Skills
Threadlight exists to make Microsoft's own platform **trivial to adopt** — never
to replace it. For first-party depth behind this CI/CD leg, reach for the official
**[Azure Skills](https://github.com/microsoft/azure-skills)** catalog. *Further
reading, not a dependency* — Threadlight's guidance stays the source of truth for
the pilot flow:
- **[`entra-app-registration`](https://github.com/microsoft/azure-skills/blob/main/skills/entra-app-registration/SKILL.md)** — **Entra app registration** + OAuth 2.0 / MSAL; first-party depth behind the OIDC / Workload-Identity-Federation federated-credential setup this generator scaffolds.
Ver en GitHub