Skip to main content
haozhx23
GitHub 제작자 프로필

haozhx23

1개 GitHub 저장소에서 수집된 5개 skills를 저장소 단위로 보여줍니다.

수집된 skills
5
저장소
1
업데이트
2026-06-16
저장소 지도

skills가 있는 위치

수집된 skill 수가 많은 주요 저장소와 이 제작자 카탈로그 내 비중, 직업 분포를 보여줍니다.

저장소 탐색

저장소와 대표 skills

hyperpod-nccl
네트워크·컴퓨터 시스템 관리자

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures, EFA TCP fallback, /dev/shm or memlock issues, NCCL version mismatch across pods, container OOM / exit-137 / OOMKilled, GPU OOM (CUDA out of memory), CrashLoopBackOff / Pending pods, MASTER_ADDR DNS, NetworkPolicy blocking. Not for single-node hardware faults (→ hyperpod-node-debugger § G) or cluster-creation EFA / SSM failures (→ hyperpod-cluster-debugger § A / § F).

2026-06-16
create-cluster-pipeline
네트워크·컴퓨터 시스템 관리자

Create a full HyperPod cluster from scratch, including EKS creation, dependency configuration, and HyperPod provisioning. Use when creating a new cluster, setting up HyperPod infrastructure, bootstrapping environment.

2026-03-17
add-cluster-instance
네트워크·컴퓨터 시스템 관리자

Add a new instance group to an existing HyperPod cluster. Use when user wants to add GPU nodes, expand cluster capacity, or add a new instance group.

2026-03-12
cluster-addons
네트워크·컴퓨터 시스템 관리자

Install and manage cluster add-ons (Karpenter, etc.). Use when user wants to install Karpenter, configure cluster add-ons, or enable additional cluster features.

2026-03-12
deploy-model
소프트웨어 개발자

Deploy a containerized inference model (vLLM, SGLang, etc.) to the cluster. Use when user wants to deploy a model, start inference service, or serve a model.

2026-03-12
저장소 1개 중 1개 표시
모든 저장소를 표시했습니다