Skip to main content

NVIDIA/deepops

SkillsMP 已收集 NVIDIA/deepops 中的 6 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
6
GitHub 星标
1,469
GitHub Forks
356

这个仓库中的 skills

1 个职业分类 · 已分类 100%

已展示 6 / 6 个已收集 Skill。

职业分类
网络与计算机系统管理员
描述

Check whether a DeepOps-deployed Slurm or Kubernetes GPU cluster is healthy and report a machine-readable verdict. Use for health checks, post-deploy verification, "is the cluster working?" questions, and after any node or driver change.

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Provision or reinstall bare-metal servers and test VMs through Canonical MAAS, map deployed machines into DeepOps Ansible inventory with MAAS tags, validate access, or release them safely. Use when operating DeepOps with a MAAS-owned machine lifecycle.

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Prepare mirrors and transfer artifacts, configure DeepOps, deploy Slurm or Kubernetes GPU clusters without Internet access, and validate them with machine-readable gates. Use for disconnected, restricted-egress, offline, or air-gapped DeepOps installations…

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Deploy a Kubernetes GPU cluster with DeepOps (Kubespray + GPU Operator) and prove it schedules GPU pods. Use when asked to deploy or rebuild Kubernetes on GPU servers with this repository.

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Deploy a Slurm GPU cluster with DeepOps and prove it works. Use when asked to deploy, install, or rebuild Slurm on one or more GPU servers with this repository.

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Diagnose NVIDIA driver installation failures on DeepOps-managed nodes — nvidia-smi errors, "No devices were found", DKMS build failures, or GPU pods crash-looping. Use before reinstalling anything.

原文语言:英语

更新
已展示 6 / 6 个已收集 Skill。