Skip to main content

loom-kubernetes

Kubernetes deployment, cluster architecture, security, and operations. Not for: a loom orchestration job or CI pipeline job — see loom-ci-cd.

Source facts

Repository
cosmix/loom
Last source activity
September 19, 2026 at 08:01
Detected SKILL.md language
English
Stars
56
Forks
3

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
loom-kubernetes
description
Kubernetes deployment, cluster architecture, security, and operations. Not for: a loom orchestration job or CI pipeline job — see loom-ci-cd.
allowed-tools
["Read","Grep","Glob","Edit","Write","Bash"]
triggers
["kubernetes","k8s","pod","deployment","statefulset","daemonset","service","ingress","configmap","secret","pvc","namespace","helm","chart","kubectl","cluster","rbac","networkpolicy","podsecurity","operator","crd","cronjob","hpa","pdb","kustomize"]
# Kubernetes ## Overview Production Kubernetes resource design, security hardening, Helm, and operations. The annotated examples below and the **Expert Practices** section carry the load-bearing knowledge — most production failures here are silent (no API error, no event), surfacing only under load, node maintenance, or a hardened cluster. ## Workload & Identity Cheatsheet | Kind | Identity guarantee | Use for | | --------------- | ------------------------------------------------------------------------------- | ------------------------------------------- | | **Deployment** | Fungible pods, random names, no ordering | Stateless services | | **StatefulSet** | Stable ordinal identity + DNS + per-pod PVC; ordered rollout (`OrderedReady`) | Databases, quorum systems, sharded stores | | **DaemonSet** | One pod per (matching) node | Node agents: logging, CNI, node-exporter | | **Job/CronJob** | Run-to-completion / scheduled | Batch, migrations, backups | Rollout knobs: `strategy.rollingUpdate.maxSurge`/`maxUnavailable` (Deployment); `maxUnavailable` + `partition` (StatefulSet). `imagePullPolicy`: `IfNotPresent` for immutable tags/digests; `Always` only for mutable tags (adds a registry round-trip per start). ## Examples ### Production Deployment (annotated) ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: api-server namespace: production labels: {app: api-server, version: v1.2.0} spec: replicas: 3 strategy: type: RollingUpdate rollingUpdate: {maxSurge: 1, maxUnavailable: 0} selector: matchLabels: {app: api-server} template: metadata: labels: {app: api-server, version: v1.2.0} annotations: {prometheus.io/scrape: "true", prometheus.io/port: "8080"} spec: serviceAccountName: api-server securityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 1000 seccompProfile: type: RuntimeDefault # REQUIRED by Restricted PSS; omitting it = pod rejected containers: - name: api image: myregistry.io/api-server:v1.2.0 imagePullPolicy: IfNotPresent ports: [{name: http, containerPort: 8080}] env: - name: DATABASE_URL valueFrom: {secretKeyRef: {name: api-secrets, key: database-url}} resources: requests: {cpu: 100m, memory: 128Mi} limits: {memory: 512Mi} # memory limit kept; CPU limit omitted (see CFS throttling) # startupProbe gates liveness/readiness — prefer over a large initialDelaySeconds startupProbe: httpGet: {path: /health/live, port: http} failureThreshold: 30 # 30 * 10s = 5 min startup budget periodSeconds: 10 livenessProbe: httpGet: {path: /health/live, port: http} # process health ONLY — no DB/cache/upstream periodSeconds: 20 failureThreshold: 3 readinessProbe: httpGet: {path: /health/ready, port: http} # may check dependencies; drains, not restarts periodSeconds: 10 failureThreshold: 3 securityContext: allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: {drop: [ALL]} volumeMounts: - {name: tmp, mountPath: /tmp} volumes: - {name: tmp, emptyDir: {}} # soft node spread + hard zone spread with per-revision isolation (see Expert Practices) affinity: podAntiAffinity: preferredDuringSchedulingIgnoredDuringExecution: - weight: 100 podAffinityTerm: labelSelector: {matchLabels: {app: api-server}} topologyKey: kubernetes.io/hostname topologySpreadConstraints: - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: DoNotSchedule labelSelector: {matchLabels: {app: api-server}} matchLabelKeys: [pod-template-hash] # each rollout revision spreads independently (1.27+) ``` ### Service + Ingress ```yaml apiVersion: v1 kind: Service metadata: {name: api-server, namespace: production} spec: type: ClusterIP selector: {app: api-server} ports: [{port: 80, targetPort: http, name: http}] --- apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: api-server namespace: production annotations: {cert-manager.io/cluster-issuer: letsencrypt-prod} spec: ingressClassName: nginx # canonical since 1.18; kubernetes.io/ingress.class is deprecated tls: [{hosts: [api.example.com], secretName: api-tls-cert}] rules: - host: api.example.com http: paths: - path: / pathType: Prefix backend: {service: {name: api-server, port: {number: 80}}} ``` ### ConfigMap + Secret ```yaml apiVersion: v1 kind: ConfigMap metadata: {name: api-config, namespace: production} data: log-level: "info" feature-flags: | {"new-checkout": true} --- apiVersion: v1 kind: Secret metadata: {name: api-secrets, namespace: production} type: Opaque # WARNING: Secrets are only base64-encoded, NOT encrypted at rest in etcd by default. # Enable encryption at rest or use External Secrets Operator. Placeholders only — never commit real creds. stringData: database-url: "postgresql://user:CHANGE_ME@db-host:5432/myapp" ``` ### HorizontalPodAutoscaler ```yaml apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: {name: api-server, namespace: production} spec: scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: api-server} minReplicas: 3 maxReplicas: 10 metrics: - type: Resource resource: name: cpu target: {type: Utilization, averageUtilization: 70} # WARNING: memory-based HPA scales out but never back in (memory is non-compressible, runtimes # rarely return heap to the OS) — prefer CPU or external/custom metrics (RPS, queue depth via KEDA). # Every container MUST set resources.requests.cpu or the HPA silently ignores the pod. behavior: scaleDown: stabilizationWindowSeconds: 300 policies: [{type: Percent, value: 10, periodSeconds: 60}] scaleUp: stabilizationWindowSeconds: 0 policies: [{type: Percent, value: 100, periodSeconds: 15}] ``` ### NetworkPolicy (zero-trust) + default-deny ```yaml apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: {name: api-server-netpol, namespace: production} spec: podSelector: {matchLabels: {app: api-server}} policyTypes: [Ingress, Egress] ingress: - from: # Same list element = namespaceSelector AND podSelector (zero-trust). Separate dashes = OR. - namespaceSelector: {matchLabels: {name: ingress-nginx}} ports: [{protocol: TCP, port: 8080}] egress: - to: [{podSelector: {matchLabels: {app: postgresql}}}] ports: [{protocol: TCP, port: 5432}] - to: - namespaceSelector: {matchLabels: {name: kube-system}} podSelector: {matchLabels: {k8s-app: kube-dns}} ports: - {protocol: UDP, port: 53} - {protocol: TCP, port: 53} # required: DNS falls back to TCP for large responses / DNSSEC --- apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: {name: default-deny-all, namespace: production} spec: podSelector: {} policyTypes: [Ingress, Egress] ``` ### RBAC (least privilege) ```yaml apiVersion: v1 kind: ServiceAccount metadata: {name: api-server, namespace: production} --- apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: {name: api-server, namespace: production} rules: - apiGroups: [""] resources: ["configmaps"] verbs: ["get", "list", "watch"] - apiGroups: [""] resources: ["secrets"] resourceNames: ["api-secrets"] # explicit allowlist; grant only `get` (list/watch leak .data) verbs: ["get"] --- apiVersion: rbac.authorization.k8s.io/v1 kind: RoleBinding metadata: {name: api-server, namespace: production} roleRef: {apiGroup: rbac.authorization.k8s.io, kind: Role, name: api-server} subjects: [{kind: ServiceAccount, name: api-server, namespace: production}] ``` ### StatefulSet with per-pod storage ```yaml apiVersion: apps/v1 kind: StatefulSet metadata: {name: postgresql, namespace: production} spec: serviceName: postgresql-headless # required headless Service for stable pod DNS replicas: 3 selector: {matchLabels: {app: postgresql}} persistentVolumeClaimRetentionPolicy: whenDeleted: Delete # reclaim PVCs on teardown (default is Retain) whenScaled: Retain # keep data on scale-in template: metadata: labels: {app: postgresql} spec: containers: - name: postgres image: postgres:15-alpine ports: [{containerPort: 5432, name: postgres}] env: - {name: PGDATA, value: /var/lib/postgresql/data/pgdata} volumeMounts: [{name: data, mountPath: /var/lib/postgresql/data}] resources: requests: {cpu: 250m, memory: 512Mi} limits: {memory: 2Gi} volumeClaimTemplates: - metadata: {name: data} spec: accessModes: ["ReadWriteOnce"] storageClassName: fast-ssd resources: {requests: {storage: 50Gi}} ``` ### Helm chart essentials Structure: `Chart.yaml` (metadata + `dependencies`), `values.yaml` (defaults), `templates/` (`_helpers.tpl` for shared labels/names, resource templates, `NOTES.txt`), `charts/` (deps), `.helmignore`. ```yaml # templates/_helpers.tpl — standard label block reused via {{ include "app.labels" . }} {{- define "app.labels" -}} helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version }} app.kubernetes.io/name: {{ .Chart.Name }} app.kubernetes.io/instance: {{ .Release.Name }} app.kubernetes.io/version: {{ .Chart.AppVersion | quote }} app.kubernetes.io/managed-by: {{ .Release.Service }} {{- end }} ``` Test before install: `helm lint`, `helm template`, `helm install --dry-run`. Roll back with `helm rollback <release> <rev>`. Name templates: `... | trunc 63 | trimSuffix "-"` (K8s name limit). ## Troubleshooting Commands ```bash # Pods kubectl describe pod <pod> -n <ns> kubectl logs <pod> -n <ns> [--previous] [-c <container>] kubectl exec -it <pod> -n <ns> -- /bin/sh kubectl get events -n <ns> --sort-by='.lastTimestamp' # Resource usage / scheduling kubectl top nodes; kubectl top pods -n <ns> kubectl describe node <node> # Network debug (ephemeral toolbox) kubectl run -it --rm debug --image=nicolaka/netshoot --restart=Never -- /bin/bash # Rollouts kubectl rollout status|history|undo|restart deployment/<name> -n <ns> # RBAC kubectl auth can-i get pods --as=system:serviceaccount:<ns>:<sa> -n <ns> # Cluster health (componentstatuses deprecated since 1.19) kubectl get --raw '/livez?verbose' kubectl get --raw '/readyz?verbose' ``` ## Common Issues | Issue | Cause | Fix | | ----------------------- | ------------------------------------------------ | -------------------------------------------------- | | ImagePullBackOff | Registry auth missing or image not found | Check imagePullSecrets, verify image exists | | CrashLoopBackOff | App crashes on startup | Check logs, verify config, add startup probe | | Pending pod | Insufficient resources / scheduling constraints | Check node capacity, taints, tolerations, affinity | | OOMKilled (exit 137) | Memory limit exceeded | Raise memory limit or fix leak (memory can't throttle) | | Service unreachable | Wrong selector or port | Verify selector matches pod labels, check ports | | DNS resolution fails | CoreDNS or NetworkPolicy blocking | Check CoreDNS pods, verify egress UDP+TCP/53 | | PVC pending | StorageClass missing or no volumes | Verify StorageClass and provisioner | | HPA not scaling | metrics-server missing or no resource requests | Install metrics-server, set requests.cpu | ## Expert Practices: Idioms, Anti-Patterns & Gotchas High-signal practices that separate working-on-my-laptop manifests from production-grade ones. Most of these fail silently — no API error, no event — so they bite only under load, during node maintenance, or in a hardened cluster. ### Probes
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub