con un clic
hpc-skills
hpc-skills contiene 4 skills recopiladas de Zhangyanbo, con cobertura ocupacional por repositorio y páginas de detalle dentro del sitio.
Skills en este repositorio
Operate an EPFL RunAI/Kubernetes GPU cluster ("Haas"-style setups) over SSH on the user's behalf: deploy code, submit / monitor / cancel RunAI jobs, manage PVC-backed persistent storage, fetch results back, check node-pool / GPU availability, and use the cluster as part of an iterate-loop. Unlike a SLURM cluster, jobs here are Kubernetes pods scheduled by RunAI — there is no sbatch/squeue. Use this skill whenever the user mentions running something on an EPFL RunAI cluster, PVC-backed pods, "haas", RunAI jobs, or anything involving `runai submit` / `runai list` / `runai describe` — even casually.
Operate the KAUST Ibex HPC cluster (SLURM) over SSH on the user's behalf: deploy code, submit / monitor / cancel jobs, manage GPU allocations (A100/V100), fetch results back, check quotas / partitions / GPU availability, and use the cluster as part of an iterate-loop. Use this skill whenever the user mentions running anything on Ibex / the HPC / cluster ("run this on ibex", "submit to slurm", "check my ibex jobs", "在 ibex 上跑"), asks about job status, storage quota, transferring files to/from the cluster, installing packages on the cluster, or anything involving sbatch / squeue / srun / sinfo / Ibex — even casually.
Operate the Northwestern University Quest HPC cluster (SLURM) over SSH on the user's behalf: deploy code, submit / monitor / cancel jobs, manage GPU allocations, fetch results back, check quotas / partitions / GPU availability, and use the cluster as part of an iterate-loop. Use this skill whenever the user mentions running anything on Quest / the HPC / cluster ("run this on Quest", "submit to slurm", "check my quest jobs", "在 quest 上跑"), asks about job status, storage quota, transferring files to/from the cluster, installing packages on the cluster, or anything involving sbatch / squeue / srun / sinfo / Quest — even casually.
Operate the Tufts University HPC cluster (SLURM) over SSH on the user's behalf: deploy code, submit / monitor / cancel jobs, run array-job experiment matrices, fetch results back, check quotas / partitions / GPU availability, and use the cluster as part of an iterate-loop. Use this skill whenever the user mentions running anything on the HPC / cluster / 集群 / 服务器 ("把它放到 hpc 上跑", "run this on the cluster", "submit to slurm", "在集群上训练"), asks about job status ("hpc 上跑得怎么样", "check my jobs", "任务跑完了吗"), storage quota, transferring files to/from the cluster, installing packages on the cluster, or anything involving sbatch / squeue / srun / sinfo / Tufts HPC — even casually.