- name
- devops
- description
- [production-grade internal] Sets up deployment and infrastructure — Docker, CI/CD pipelines, cloud provisioning, environment configuration. Routed via the production-grade orchestrator.
# DevOps
## Protocols
<!-- protocol injection (was: !`cat Claude-Production-Grade-Suite/.protocols/ux-protocol.md 2>/dev/null || true`) -->
Read protocol: `${PG_PROTOCOLS}/ux-protocol.md` (use the `read_file` tool before continuing).
<!-- protocol injection (was: !`cat Claude-Production-Grade-Suite/.protocols/input-validation.md 2>/dev/null || true`) -->
Read protocol: `${PG_PROTOCOLS}/input-validation.md` (use the `read_file` tool before continuing).
<!-- protocol injection (was: !`cat Claude-Production-Grade-Suite/.protocols/tool-efficiency.md 2>/dev/null || true`) -->
Read protocol: `${PG_PROTOCOLS}/tool-efficiency.md` (use the `read_file` tool before continuing).
<!-- protocol injection (was: !`cat Claude-Production-Grade-Suite/.protocols/visual-identity.md 2>/dev/null || true`) -->
Read protocol: `${PG_PROTOCOLS}/visual-identity.md` (use the `read_file` tool before continuing).
<!-- protocol injection (was: !`cat Claude-Production-Grade-Suite/.protocols/freshness-protocol.md 2>/dev/null || true`) -->
Read protocol: `${PG_PROTOCOLS}/freshness-protocol.md` (use the `read_file` tool before continuing).
<!-- protocol injection (was: !`cat Claude-Production-Grade-Suite/.protocols/receipt-protocol.md 2>/dev/null || true`) -->
Read protocol: `${PG_PROTOCOLS}/receipt-protocol.md` (use the `read_file` tool before continuing).
<!-- protocol injection (was: !`cat Claude-Production-Grade-Suite/.protocols/boundary-safety.md 2>/dev/null || true`) -->
Read protocol: `${PG_PROTOCOLS}/boundary-safety.md` (use the `read_file` tool before continuing).
<!-- protocol injection (was: !`cat Claude-Production-Grade-Suite/.protocols/conflict-resolution.md 2>/dev/null || true`) -->
Read protocol: `${PG_PROTOCOLS}/conflict-resolution.md` (use the `read_file` tool before continuing).
<!-- inline shell (was: !`cat .production-grade.yaml 2>/dev/null || echo "No config — using defaults"`) -->
Run shell command before continuing: ``cat .production-grade.yaml 2>/dev/null || echo "No config — using defaults"``
(use the `execute_shell_command` tool).
<!-- inline shell (was: !`cat Claude-Production-Grade-Suite/.orchestrator/codebase-context.md 2>/dev/null || true`) -->
Run shell command before continuing: ``cat Claude-Production-Grade-Suite/.orchestrator/codebase-context.md 2>/dev/null || true``
(use the `execute_shell_command` tool).
**Fallback (if protocols not loaded):** Use AskUserQuestion with options (never open-ended), "Chat about this" last, recommended first. Work continuously. Print progress constantly. Validate inputs before starting — classify missing as Critical (stop), Degraded (warn, continue partial), or Optional (skip silently). Use parallel tool calls for independent reads. Use smart_outline before full Read.
## Engagement Mode
<!-- inline shell (was: !`cat Claude-Production-Grade-Suite/.orchestrator/settings.md 2>/dev/null || echo "No settings — using Standard"`) -->
Run shell command before continuing: ``cat Claude-Production-Grade-Suite/.orchestrator/settings.md 2>/dev/null || echo "No settings — using Standard"``
(use the `execute_shell_command` tool).
| Mode | Behavior |
|------|----------|
| **Express** | Fully autonomous. Use architecture's cloud choice. Sensible defaults for all infra. Report decisions in output. |
| **Standard** | Surface 1-2 critical decisions — container registry choice, CI provider (if not specified in architecture), monitoring stack. |
| **Thorough** | Surface all major decisions. Show Dockerfile strategy, CI pipeline design, monitoring architecture before implementing. Ask about deployment strategy (blue-green, canary, rolling). |
| **Meticulous** | Surface every decision. Walk through each Terraform module. Review CI pipeline stages. User approves monitoring alert thresholds. |
## Progress Output
Follow `Claude-Production-Grade-Suite/.protocols/visual-identity.md`. Print structured progress throughout execution.
**Skill header** (print on start):
```
━━━ DevOps ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
```
**Phase progress** (print during execution):
```
[1/4] Containerization
✓ {N} Dockerfiles, 1 docker-compose
⧖ building multi-stage images...
○ CI/CD pipelines
○ infrastructure as code
○ monitoring
[2/4] CI/CD Pipelines
✓ {N} workflows ({provider})
⧖ configuring deployment strategies...
○ infrastructure as code
○ monitoring
[3/4] Infrastructure as Code
✓ {N} Terraform modules, {M} resources
⧖ provisioning cloud resources...
○ monitoring
[4/4] Monitoring & Observability
✓ dashboards, alerting configured
```
**Completion summary** (print on finish — MUST include concrete numbers):
```
✓ DevOps {N} Dockerfiles, {M} workflows, {K} Terraform modules ⏱ Xm Ys
```
## Brownfield Awareness
If `Claude-Production-Grade-Suite/.orchestrator/codebase-context.md` exists and mode is `brownfield`:
- **READ existing infrastructure first** — check for Dockerfiles, CI configs, Terraform, K8s manifests
- **EXTEND, don't replace** — add new services to existing docker-compose, add jobs to existing CI
- **NEVER overwrite** — existing Dockerfile, workflows, or Terraform state
- **Match existing patterns** — if they use GitHub Actions, don't create GitLab CI. If they use Pulumi, don't create Terraform
## Overview
Full DevOps pipeline generator: from infrastructure design to production-ready deployment with monitoring and security. Generates infrastructure and deployment artifacts at the project root (`infrastructure/`, `.github/workflows/`, Dockerfiles) with planning notes in `Claude-Production-Grade-Suite/devops/`.
## Config Paths
Read `.production-grade.yaml` at startup. Use these overrides if defined:
- `paths.terraform` — default: `infrastructure/terraform/`
- `paths.kubernetes` — default: `infrastructure/kubernetes/`
- `paths.ci_cd` — default: `.github/workflows/`
- `paths.monitoring` — default: `infrastructure/monitoring/`
## When to Use
- Setting up CI/CD pipelines for a new or existing project
- Creating infrastructure as code for cloud deployments
- Containerizing applications with Docker/Kubernetes
- Configuring monitoring, logging, and alerting
- Implementing security scanning and secrets management
- Multi-cloud or hybrid-cloud deployment planning
- Production readiness review and hardening
## Parallel Execution
After Phase 1 (Assessment), Phases 2-4 and Phases 5-6 can run as two parallel groups:
**Group 1 (infrastructure artifacts — independent):**
```python
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Generate Terraform IaC following Phase 2. Write to infrastructure/terraform/.", ...)
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Generate CI/CD pipelines following Phase 3. Write to .github/workflows/ and scripts/.", ...)
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Generate container orchestration following Phase 4. Write Dockerfiles and K8s manifests.", ...)
```
**Group 2 (after Group 1 — needs infrastructure context):**
```python
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Generate monitoring + observability following Phase 5. Write to infrastructure/monitoring/.", ...)
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Generate security infrastructure following Phase 6. Write to infrastructure/security/.", ...)
```
**Execution order:**
1. Phase 1: Assessment (sequential)
2. Phases 2-4: IaC + CI/CD + Containers (PARALLEL)
3. Phases 5-6: Monitoring + Security (PARALLEL, after Group 1)
## Process Flow
```dot
digraph devops {
rankdir=TB;
"Triggered" [shape=doublecircle];
"Phase 1: Assessment" [shape=box];
"Phase 2: IaC" [shape=box];
"Phase 3: CI/CD" [shape=box];
"Phase 4: Containers" [shape=box];
"Phase 5: Monitoring" [shape=box];
"Phase 6: Security" [shape=box];
"User Review" [shape=diamond];
"Suite Complete" [shape=doublecircle];
"Triggered" -> "Phase 1: Assessment";
"Phase 1: Assessment" -> "Phase 2: IaC";
"Phase 2: IaC" -> "User Review";
"User Review" -> "Phase 2: IaC" [label="revise"];
"User Review" -> "Phase 3: CI/CD" [label="approved"];
"Phase 3: CI/CD" -> "Phase 4: Containers";
"Phase 4: Containers" -> "Phase 5: Monitoring";
"Phase 5: Monitoring" -> "Phase 6: Security";
"Phase 6: Security" -> "Suite Complete";
}
```
## Phase 1: Infrastructure Assessment
**Engagement mode determines assessment depth:**
- **Express**: Infer all answers from codebase analysis, architecture docs, and .production-grade.yaml. Report assumptions in output. Do NOT ask.
- **Standard**: Ask only for unknowns not discoverable from code (budget/compliance, 1 call max).
- **Thorough/Meticulous**: Use AskUserQuestion to gather (batch into 2-3 calls max):
1. **Current state** — Existing infra? Greenfield? Migration? What's already running?
2. **Application profile** — Language/framework, stateful/stateless, background jobs, WebSockets?
3. **Scale requirements** — Traffic patterns (steady/bursty), auto-scaling needs, regions
4. **Environments** — How many? (dev/staging/prod minimum), environment parity strategy
5. **Budget & compliance** — Cost constraints, regulatory requirements (SOC2/HIPAA/PCI)
6. **Team capabilities** — DevOps maturity, on-call rotation, incident response existing?
## Phase 2: Infrastructure as Code (Terraform)
Generate `infrastructure/terraform/` (or `paths.terraform` from config):
### Module Structure
```
terraform/
├── modules/
│ ├── networking/ # VPC, subnets, security groups, NAT
│ ├── compute/ # ECS/EKS/GKE/AKS clusters
│ ├── database/ # RDS/Cloud SQL/Azure SQL, Redis
│ ├── messaging/ # SQS/Pub-Sub/Service Bus
│ ├── storage/ # S3/GCS/Blob, CDN
│ ├── monitoring/ # CloudWatch/Cloud Monitoring/Azure Monitor
│ ├── security/ # IAM, KMS, WAF, secrets
│ └── dns/ # Route53/Cloud DNS/Azure DNS
├── environments/
│ ├── dev/
│ │ ├── main.tf
│ │ ├── variables.tf
│ │ ├── terraform.tfvars
│ │ └── backend.tf
│ ├── staging/
│ └── prod/
├── global/ # Shared resources (IAM, DNS zones)
└── README.md
```
### Terraform Standards
- **Remote state** — S3/GCS/Azure Blob backend with state locking (DynamoDB/GCS/Azure Table)
- **Module versioning** — Pinned module versions, semantic versioning
- **Variable validation** — `validation` blocks on all input variables
- **Tagging strategy** — `environment`, `service`, `team`, `cost-center`, `managed-by=terraform`
- **Least privilege IAM** — Service-specific roles, no wildcard permissions
- **Encryption everywhere** — KMS-managed keys for storage, databases, secrets
- **Network isolation** — Private subnets for compute/data, public only for load balancers
### Multi-Cloud Provider Configs
Generate provider blocks and modules for each target cloud:
| Resource | AWS | GCP | Azure |
|----------|-----|-----|-------|
| Compute | ECS Fargate / EKS | Cloud Run / GKE | Container Apps / AKS |
| Database | RDS Aurora | Cloud SQL | Azure SQL |
| Cache | ElastiCache Redis | Memorystore | Azure Cache Redis |
| Queue | SQS + SNS | Pub/Sub | Service Bus |
| Storage | S3 + CloudFront | GCS + Cloud CDN | Blob + Front Door |
| Secrets | Secrets Manager | Secret Manager | Key Vault |
| DNS | Route 53 | Cloud DNS | Azure DNS |
| WAF | AWS WAF | Cloud Armor | Azure WAF |
**Present IaC design to user for approval before proceeding.**
## Phase 3: CI/CD Pipelines
Generate CI/CD pipelines at `.github/workflows/` (or `paths.ci_cd` from config) and `scripts/`:
### Pipeline Templates
```
.github/workflows/
├── ci.yml # Build, test, lint, security scan
├── cd-staging.yml # Deploy to staging on merge to main
├── cd-production.yml # Deploy to prod on release tag
├── pr-checks.yml # PR validation (tests, lint, preview)
└── scheduled.yml # Nightly builds, dependency updates
.gitlab-ci.yml # (if requested, at project root)
scripts/
├── build.sh
├── deploy.sh
├── rollback.sh
└── smoke-test.sh
```
### CI Pipeline Stages
1. **Checkout & cache** — Restore dependency caches
2. **Install** — Dependencies with lockfile verification
3. **Lint** — Code style, formatting (fail-fast)
4. **Type check** — Static analysis (if applicable)
5. **Unit tests** — Parallel execution, coverage reporting
6. **Integration tests** — Against test containers (testcontainers)
7. **Security scan** — SAST (Semgrep/CodeQL), dependency audit (Snyk/Trivy)
8. **Build** — Docker image with content-hash tagging
9. **Push** — To ECR/GCR/ACR with immutable tags
### CD Pipeline Stages
1. **Deploy to staging** — Automatic on main branch merge
2. **Smoke tests** — Health checks + critical path verification
3. **Performance tests** — Load testing gate (k6/Artillery)
4. **Manual approval** — Required for production (GitHub Environments)
5. **Deploy to production** — Blue-green or canary strategy
6. **Post-deploy verification** — Automated smoke + synthetic monitoring
7. **Rollback trigger** — Automatic on error rate spike
### Deployment Strategies
Generate configs for the selected strategy:
- **Blue-Green** — Zero-downtime with instant rollback (default for stateless)
Voir sur GitHub