Skip to main content

hpc-supercomputing-clusters

Provides architecture, engineering, and operations patterns for HPC (High Performance Computing) clusters and supercomputers based on Supercomputers for Linux SysAdmins (Sergey Zhumatiy). Covers workload managers (Slurm Workload Manager), low-latency interconnects (InfiniBand/RDMA), parallelism libraries (OpenMPI, MPICH), parallel file systems (Lustre, GPFS, Ceph), and GPU-accelerated computing.

소스 정보

저장소
dandgabr/Coacus
최근 소스 활동
2026년 9월 28일 14:03
감지된 SKILL.md 언어
영어
스타
4
포크
3

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
5 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
hpc-supercomputing-clusters
description
Provides architecture, engineering, and operations patterns for HPC (High Performance Computing) clusters and supercomputers based on Supercomputers for Linux SysAdmins (Sergey Zhumatiy). Covers workload managers (Slurm Workload Manager), low-latency interconnects (InfiniBand/RDMA), parallelism libraries (OpenMPI, MPICH), parallel file systems (Lustre, GPFS, Ceph), and GPU-accelerated computing.
# HPC Clusters and Supercomputer Engineering This skill establishes the guidelines for designing, deploying, and operating **High Performance Computing (HPC)** environments and supercomputers on Linux, drawing on the book by **Sergey Zhumatiy**. --- ## 🚀 1. Topology of an HPC Cluster ``` ┌─────────────────────────┐ │ Head / Master Node │ │ (Slurm Controller) │ └────────────┬────────────┘ │ ┌────────────────────────┴────────────────────────┐ [ 10GbE Management Network ] [ InfiniBand / RDMA 200Gbps Network ] │ │ ┌─────────▼─────────┐ ┌─────────▼─────────┐ │ Parallel Storage │ │ Compute Nodes │ │ (Lustre / CephFS) │ │ (CPUs + GPUs H100)│ └───────────────────┘ └───────────────────┘ ``` --- ## 📋 2. Job Management with Slurm Workload Manager ### Batch Job Script Example (`submit_job.sh`) ```bash #!/bin/bash #SBATCH --job-name=scientific_sim #SBATCH --output=logs/sim_%j.log #SBATCH --error=logs/sim_%j.err #SBATCH --nodes=4 #SBATCH --ntasks-per-node=32 #SBATCH --cpus-per-task=1 #SBATCH --gres=gpu:4 #SBATCH --time=12:00:00 #SBATCH --partition=gpu_cluster module load openmpi/4.1.5-cuda-12.2 srun --mpi=pmix ./bin/simulation_engine --dataset /shared/lustre/input.dat ```
GitHub에서 보기