원클릭으로
gpu-pressure
GPU memory and utilization headroom
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
GPU memory and utilization headroom
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
| name | gpu_pressure |
| description | GPU memory and utilization headroom |
| category | memory |
| tables | ["gpu.utilization","gpu.devices","python.torch_trace"] |
| tags | ["gpu","VRAM","utilization","显存","利用率"] |
| keywords | {"en":["GPU memory","VRAM","GPU utilization","GPU idle"],"zh":["显存不够","GPU 利用率","VRAM","显存占用","GPU 空闲"]} |
| parameters | {"sample_limit":{"type":"integer","default":20}} |
查看 gpu.utilization 采样与 python.torch_trace 中的 allocated 是否一致, 判断是「真 OOM 风险」还是「利用率低 / 内存碎片」。
sample_limit (integer, default 20):NCCL proxy wait decomposition — culprit (send_gpu_wait) vs victim (recv_wait)
One-shot health check: CPU, GPU, tables, torch step progress, TorchProbe overhead, cluster nodes, NCCL profiler health
SRE first-response triage for distributed training incidents. Automates the manual checks from PyTorch/NCCL debugging runbooks.
Diagnose PyTorch NCCL watchdog timeouts with Flight Recorder collective sequence alignment.
Detect monotonic GPU memory growth across training steps
Find slowest PyTorch modules in recent steps