Skip to main content

NVIDIA/nvidia-resiliency-ext

SkillsMP 已收集 NVIDIA/nvidia-resiliency-ext 中的 4 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
4
GitHub 星标
324
GitHub Forks
64

这个仓库中的 skills

1 个职业分类 · 已分类 100%

已展示 4 / 4 个已收集 Skill。

职业分类
软件开发工程师
描述

Closed-loop fault injection and attribution accuracy benchmark. Draws from a prioritized pool of (fault_type, rank, iter, nodes) experiments and submits them 2 at a time via sbatch — waiting for each pair to finish before submitting the next — to bound…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Analyze a SLURM job log file for failure root-cause attribution and restart decisions using NVRxLogAnalyzer. Use when you have a SLURM training job log and need to determine why the job failed and whether it should be restarted. Performs per-cycle chunking,…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Orchestration layer over nvidia_resiliency_ext attribution modules. Provides log-analysis, fr-analysis, and a Megatron-LM-oriented fault-injection feedback loop for benchmarking attribution quality on SLURM workloads.

原文语言:英语

更新
职业分类
软件开发工程师
描述

Analyze PyTorch NCCL flight-recorder (FR) dumps to identify collective operation hangs and isolate the responsible ranks using CollectiveAnalyzer. Use when a distributed training job hangs due to an NCCL collective timeout and FR dump files are available.…

原文语言:英语

更新
已展示 4 / 4 个已收集 Skill。