| name | debug-buttercup |
| description | All pods run in namespace crs. Use when pods in the crs namespace are in CrashLoopBackOff, OOMKilled, or restarting, multiple services restart simultaneously (cascade failure), or redis is unresponsiv |
| category | AI & Agents |
| source | antigravity |
| tags | ["node","api","ai","llm","workflow","template","docker","kubernetes","rag","cro"] |
| url | https://github.com/sickn33/antigravity-awesome-skills/tree/main/skills/debug-buttercup |
Debug Buttercup
When to Use
- Pods in the
crs namespace are in CrashLoopBackOff, OOMKilled, or restarting
- Multiple services restart simultaneously (cascade failure)
- Redis is unresponsive or showing AOF warnings
- Queues are growing but tasks are not progressing
- Nodes show DiskPressure, MemoryPressure, or PID pressure
- Build-bot cannot reach the Docker daemon (DinD failures)
- Scheduler is stuck and not advancing task state
- Health check probes are failing unexpectedly
- Deployed Helm values don't match actual pod configuration
When NOT to Use
- Deploying or upgrading Buttercup (use Helm and deployment guides)
- Debugging issues outside the
crs Kubernetes namespace
- Performance tuning that doesn't involve a failure symptom
Namespace and Services
All pods run in namespace crs. Key services:
| Layer | Services |
|---|
| Infra | redis, dind, litellm, registry-cache |
| Orchestration | scheduler, task-server, task-downloader, scratch-cleaner |
| Fuzzing | build-bot, fuzzer-bot, coverage-bot, tracer-bot, merger-bot |
| Analysis | patcher, seed-gen, program-model, pov-reproducer |
| Interface | competition-api, ui |
Triage Workflow
Always start with triage. Run these three commands first:
kubectl get pods -n crs -o wide
kubectl get events -n crs --sort-by='.lastTimestamp'
kubectl get events -n crs --field-selector type=Warning --sort-by='.lastTimestamp'
Then narrow down:
kubectl describe pod -n crs <pod-name> | grep -A8 'Last State:'
kubectl get pod -n crs <pod-name> -o jsonpath='{.spec.containers[0].resources}'
kubectl logs -n crs <pod-name> --previous --tail=200
kubectl logs -n crs <pod-name> --tail=200
Historical vs Ongoing Issues
High restart counts don't necessarily mean an issue is ongoing -- restarts accumulate over a pod's lifetime. Always distinguish:
--tail shows the end of the log buffer, which may contain old messages. Use --since=300s to confirm issues are actively happening now.
--timestamps on log output helps correlate events across services.
- Check
Last State timestamps in describe pod to see when the most recent crash actually occurred.
Cascade Detection
When many pods restart around the same time, check for a shared-dependency failure before investigating individual pods. The most common cascade: Redis goes down -> every service gets ConnectionError/ConnectionRefusedError -> mass restarts. Look for the same error across multiple --previous logs -- if they all say redis.exceptions.ConnectionError, debug Redis, not the individual services.
Log Analysis
kubectl logs -n crs -l app=fuzzer-bot --tail=100 --prefix
kubectl logs -n crs -l app.kubernetes.io/name=redis -f
bash deployment/collect-logs.sh
Resource Pressure
kubectl top pods -n crs
kubectl top nodes
kubectl describe node <node> | grep -A5 Conditions
kubectl exec -n crs <pod> -- df -h
kubectl exec -n crs <pod> -- sh -c 'du -sh /corpus/* 2>/dev/null'
kubectl exec -n crs <pod> -- sh -c 'du -sh /scratch/* 2>/dev/null'
Redis Debugging
Redis is the backbone. When it goes down, everything cascades.
kubectl get pods -n crs -l app.kubernetes.io/name=redis
kubectl logs -n crs -l app.kubernetes.io/name=redis --tail=200
kubectl exec -n crs <redis-pod> -- redis-cli
INFO memory
INFO persistence
INFO clients
INFO stats
CLIENT LIST
DBSIZE
CONFIG GET appendonly
CONFIG GET appendfsync
kubectl exec -n crs <redis-pod> -- mount | grep /data
kubectl exec -n crs <redis-pod> -- du -sh /data/
Queue Inspection
Buttercup uses Redis streams with consumer groups. Queue names:
| Queue | Stream Key |
|---|
| Build | fuzzer_build_queue |
| Build Output | fuzzer_build_output_queue |
| Crash | fuzzer_crash_queue |
| Confirmed Vu | |