| name | ops-devops-platform |
| description | Designs DevOps and platform engineering systems. Use when planning Kubernetes, Terraform, GitOps, CI/CD, observability, incident response, or cloud-native operations. |
| compatibility | Portable core. Works on Claude Code and Codex. |
| version | 1.1 |
| last_validated | 2026-07-11T00:00:00.000Z |
DevOps and Platform Engineering
Use this skill for platform, infrastructure, CI/CD, GitOps, observability, and incident-operating-model design. Keep the output operational: target architecture, rollout path, guardrails, ownership, and artifacts.
Quick Reference
| Need | Starting Direction |
|---|
| infrastructure provisioning | Terraform, OpenTofu, Pulumi, or cloud-native IaC |
| cluster or app deployment | GitOps first for steady-state, direct tooling for local iteration |
| CI/CD | protected pipelines plus workload identity and supply-chain controls — see supply-chain-security |
| observability | OpenTelemetry plus metrics, logs, traces, and SLO-based alerting |
| platform engineering | golden paths, policy-as-code, and self-service interfaces |
| incident operations | runbooks, severity model, escalation, and postmortems |
Workflow
- classify the dominant problem:
- provisioning
- deployment
- CI/CD
- observability
- platform engineering
- security hardening
- incident operations
- choose the smallest viable toolchain that matches the runtime and team skill
- load the relevant reference and template set
- verify version-sensitive or vendor-sensitive claims before final guidance
- finish with concrete operational outputs: plan, controls, owners, and artifacts
Decision Rules
| Situation | Rule |
|---|
| any infrastructure change | IaC first; no clickops |
| steady-state production reconciliation | GitOps (Argo CD / Flux) over push-based deploys |
| CI credentials | workload identity (OIDC) over long-lived secrets |
| alerting | SLO burn-rate alerts; suppress raw host-metric noise |
| new environments | platform template + policy guard; no snowflakes |
| supply-chain integrity | SLSA build track + cosign keyless signing |
| drift | detect via reconciler or terraform plan in CI; never discover by accident |
Related Routing
Guardrails
| Domain | Do | Anti-pattern to avoid |
|---|
| Provisioning | all material changes in IaC; explicit promotion gates | clickops drift; untagged infrastructure |
| Delivery | protected pipelines; artifact provenance; rollback + smoke checks | pipelines without identity boundaries |
| Platform | golden paths before self-service; policy-as-code that reduces variation | tools shipped without adoption path or ownership |
| Observability | define SLOs first; join logs/traces/metrics on shared trace ID | alert fatigue from raw host-metric thresholds |
| Incidents | postmortems feed runbooks and platform changes | postmortems that stop at narrative |
| Cost | tagging + budget alerts at resource creation; monthly right-sizing | unmanaged snowflake environments; unreviewed reservations |
Navigation
Reference routing
| Load when… | Reference |
|---|
| supply-chain, SBOM, signing, SLSA | references/supply-chain-security.md |
| DORA's five metrics and team archetypes (Elite/High/Medium/Low tiers are retired), AI-adoption instability tax, general DevOps best practices | references/devops-best-practices.md |
| GitLab CI — parent/child pipelines, MR variable traps, env-export pattern | references/gitlab-ci-patterns.md |
| choosing a tool (IaC, GitOps, CI, policy, observability) | references/tool-landscape.md |
| golden paths, internal developer portal, platform maturity, when NOT to build an IDP, platform-vs-product boundary, CI/IaC/GitOps adoption sequencing | references/platform-engineering-patterns.md |
| GitOps multi-env promotion, Argo CD / Flux patterns | references/gitops-workflows.md |
| on-call, severity model, escalation, postmortems | references/sre-incident-management.md |
| day-2 operational runbooks, environment hygiene | references/operational-patterns.md |
| AIOps alert correlation, automated triage | references/aiops-patterns.md |
| Kalman canary, cost autoscaler, CI capacity stabiliser | references/control-theory-applied.md |
Templates
AWS / GCP / Azure
Kubernetes
Docker / Kafka
Terraform / IaC
CI/CD and GitOps
Monitoring / Observability
Incident response
Security / Cost
Shared utilities
Related Skills
Trend Awareness Protocol
When users ask for current tool recommendations, verify:
- current supported Kubernetes and ecosystem versions
- active IaC and GitOps tool state
- current observability and policy-engine capabilities
- current CI/CD and platform-tool support windows
Prefer official docs and release notes over blogs or rankings.
Fact-Checking
- Verify current versions, deprecations, support windows, pricing, and cloud limits before final answers.
- Prefer official docs and release notes for named tools and platforms.
- If web access is unavailable, mark version-sensitive guidance as unverified.
Learnings Loop
Before applying this skill on a non-trivial task, read learnings.consolidated.md in this directory (and learnings.md if present).
After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.