HyperPod-InstantStart
HyperPod-InstantStart contient 5 skills collectées depuis haozhx23, avec une couverture métier par dépôt et des pages de détail sur le site.
Skills dans ce dépôt
Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures, EFA TCP fallback, /dev/shm or memlock issues, NCCL version mismatch across pods, container OOM / exit-137 / OOMKilled, GPU OOM (CUDA out of memory), CrashLoopBackOff / Pending pods, MASTER_ADDR DNS, NetworkPolicy blocking. Not for single-node hardware faults (→ hyperpod-node-debugger § G) or cluster-creation EFA / SSM failures (→ hyperpod-cluster-debugger § A / § F).
Create a full HyperPod cluster from scratch, including EKS creation, dependency configuration, and HyperPod provisioning. Use when creating a new cluster, setting up HyperPod infrastructure, bootstrapping environment.
Add a new instance group to an existing HyperPod cluster. Use when user wants to add GPU nodes, expand cluster capacity, or add a new instance group.
Install and manage cluster add-ons (Karpenter, etc.). Use when user wants to install Karpenter, configure cluster add-ons, or enable additional cluster features.
Deploy a containerized inference model (vLLM, SGLang, etc.) to the cluster. Use when user wants to deploy a model, start inference service, or serve a model.