Skip to main content

hpc-supercomputing-clusters

Provides architecture, engineering, and operations patterns for HPC (High Performance Computing) clusters and supercomputers based on Supercomputers for Linux SysAdmins (Sergey Zhumatiy). Covers workload managers (Slurm Workload Manager), low-latency interconnects (InfiniBand/RDMA), parallelism libraries (OpenMPI, MPICH), parallel file systems (Lustre, GPFS, Ceph), and GPU-accelerated computing.

来源信息

仓库
dandgabr/Coacus
最近来源活动
2026年9月28日 14:03
检测到的 SKILL.md 语言
英语
星标
4
分支
3

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
5 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
hpc-supercomputing-clusters
description
Provides architecture, engineering, and operations patterns for HPC (High Performance Computing) clusters and supercomputers based on Supercomputers for Linux SysAdmins (Sergey Zhumatiy). Covers workload managers (Slurm Workload Manager), low-latency interconnects (InfiniBand/RDMA), parallelism libraries (OpenMPI, MPICH), parallel file systems (Lustre, GPFS, Ceph), and GPU-accelerated computing.
# HPC Clusters and Supercomputer Engineering This skill establishes the guidelines for designing, deploying, and operating **High Performance Computing (HPC)** environments and supercomputers on Linux, drawing on the book by **Sergey Zhumatiy**. --- ## 🚀 1. Topology of an HPC Cluster ``` ┌─────────────────────────┐ │ Head / Master Node │ │ (Slurm Controller) │ └────────────┬────────────┘ │ ┌────────────────────────┴────────────────────────┐ [ 10GbE Management Network ] [ InfiniBand / RDMA 200Gbps Network ] │ │ ┌─────────▼─────────┐ ┌─────────▼─────────┐ │ Parallel Storage │ │ Compute Nodes │ │ (Lustre / CephFS) │ │ (CPUs + GPUs H100)│ └───────────────────┘ └───────────────────┘ ``` --- ## 📋 2. Job Management with Slurm Workload Manager ### Batch Job Script Example (`submit_job.sh`) ```bash #!/bin/bash #SBATCH --job-name=scientific_sim #SBATCH --output=logs/sim_%j.log #SBATCH --error=logs/sim_%j.err #SBATCH --nodes=4 #SBATCH --ntasks-per-node=32 #SBATCH --cpus-per-task=1 #SBATCH --gres=gpu:4 #SBATCH --time=12:00:00 #SBATCH --partition=gpu_cluster module load openmpi/4.1.5-cuda-12.2 srun --mpi=pmix ./bin/simulation_engine --dataset /shared/lustre/input.dat ```
在 GitHub 查看