用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/notque/vexjoy-agent --skill service-health-check命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Run the full evidence-to-live implementation workflow for large, multi-system, multi-wave, or CPU-delegated 5 Star Booker GM programs.
Classify user requests and route to the correct agent + skill. Primary entry point for all delegated work.
Structured multi-phase workflows: review, debug, refactor (tidy, clean up, untangle messy code without behaviour change), deploy, create, research.
基于 SOC 职业分类
正在显示 SKILL.md
| name | service-health-check |
| description | Service health monitoring, endpoint validation, and CVE source auditing. |
| user-invocable | false |
| allowed-tools | ["Bash","Read","Write","Glob","Grep","Edit"] |
| routing | {"triggers":["service status","process health","uptime check","is service running","check health","validate endpoints","smoke test API","health check endpoints","test endpoint","check API","smoke test","check cve sources","cve source coverage","audit cve feeds","vulnerability source audit","verify cve sources","security feed audit"],"category":"infrastructure","pairs_with":["kubernetes","condition-based-waiting","e2e-testing"]} |
This skill provides deterministic service health monitoring using the Discover-Check-Report pattern. It finds services, gathers health signals from multiple sources (process table, health files, port binding), and produces actionable reports identifying degraded or failed services.
Core principle: Health assessment is evidence-based. Never report a service healthy without verifying process status independently of health file content. Never assume a running process is functional — always cross-check against health files and port binding.
| Signal | Load These Files | Why |
|---|---|---|
| Endpoint validation request | references/endpoint-validator.md | Full endpoint validation methodology |
| Security header WARNs, HSTS/CSP/X-Frame issues | references/security-headers.md | Deep security header reference |
| Config errors, hardcoded IPs, timeout problems | references/endpoint-config-preferred-patterns.md | Endpoint config patterns |
| 401/403 failures, Bearer/API-key/cookie auth | references/auth-endpoint-patterns.md | Auth endpoint patterns |
| CVE source audit request | references/cve-source-check.md | Full CVE source check methodology |
| CVE registry schema questions | references/registry-schema.md | Registry shape and entry format |
| CVE source URL verification | references/source-verification.md | HEAD-check semantics |
| CVE report format questions | references/output-formats.md | JSON schema and Markdown sections |
Goal: Identify all services to check before running any health probes.
Step 1: Locate service definitions
Search for service configuration in this order:
services.json in project rootStep 2: Build service manifest
For each service, establish:
## Service Manifest
| Service | Process Pattern | Health File | Port | Stale Threshold |
|---------|----------------|-------------|------|-----------------|
| api-server | gunicorn.*app:app | /tmp/api_health.json | 8000 | 300s |
| worker | celery.*worker | /tmp/worker_health.json | - | 300s |
| cache | redis-server | - | 6379 | - |
Validation constraints:
Step 3: Validate manifest
Confirm each entry passes the constraints above. If a pattern is too broad, use ps aux | grep to identify distinguishing arguments, then update the pattern.
Gate: Service manifest complete with at least one service. Proceed only when gate passes.
Goal: Gather health signals for every service in the manifest. Always check process status independently of health file content—a running process and a healthy health file are separate signals.
Step 1: Check process status
For each service, run process check:
pgrep -f "<process_pattern>"
Record: running (true/false), PIDs, process count.
Rationale: Process existence is the primary signal. A missing process always means the service is DOWN. A running process alone is insufficient—the service may have crashed or failed to bind to its port.
Step 2: Parse health files (if configured)
Read and parse JSON health files. Evaluate:
Critical constraint: Never trust health file content alone. The file could be stale from before a process crash. Always verify:
Step 3: Probe ports (if configured)
Check if expected ports are listening:
ss -tlnp "sport = :<port>"
Rationale: Verify ports are actually bound. A process can start but fail to bind to its configured port—that is effectively a DOWN state, not HEALTHY.
Step 4: Evaluate health per service
Apply this decision tree (constraints embedded in logic):
Gate: All services evaluated with evidence-based status. No status is determined without concrete signal (process check, health file, or port probe). Proceed only when gate passes.
Goal: Produce structured, actionable health report with specific remediation commands.
Step 1: Generate summary
SERVICE HEALTH REPORT
=====================
Checked: N services
Healthy: X/N
RESULTS:
service-name [OK ] HEALTHY PID 12345, uptime 2d 4h
background-worker [WARN] WARNING Health file stale (15 min)
cache-service [DOWN] DOWN Process not found
RECOMMENDATIONS:
background-worker: Restart recommended - health file not updated in 900s
cache-service: Start service - process not running
SUGGESTED ACTIONS:
systemctl restart background-worker
systemctl start cache-service
Step 2: Set exit status
Step 3: Present to user
Gate: Report delivered with actionable recommendations for all non-healthy services.
User says: "Are all services up?" Actions:
User says: "The background worker seems stuck" Actions:
Cause: No services.json, docker-compose, or systemd units discovered Solution:
Cause: Pattern too broad (e.g., "python" matches all Python processes) Solution:
ps aux | grep to identify distinguishing argumentsCause: Malformed JSON, permissions issue, or file being written during read Solution:
ls -laServices should write health files as:
{
"timestamp": "ISO8601, updated every 30-60s",
"status": "healthy|degraded|error",
"connection": "connected|disconnected|reconnecting",
"last_activity": "ISO8601 of last meaningful action",
"running": true,
"uptime_seconds": 12345,
"metrics": {}
}
| Constraint | Rationale | Application |
|---|---|---|
| Process status verified independently of health file | Running process ≠ functional service | Always check process before trusting health file |
| Health file staleness detected by timestamp freshness | File could be stale from before crash | Check timestamp against 300s (configurable) threshold |
| Port binding verified when configured | Process running doesn't mean port is bound | Always verify expected port listening when port specified |
| No auto-restart without explicit flag | Restart masks root cause | Report findings first; only execute restart if user flags it |
| Narrow process patterns required | "python" matches all processes, giving false matches | Use full paths or specific args; validate with ps aux | grep |
| Evidence-based status only | Status must have supporting signal | No status without concrete evidence (process, health file, or port) |