Route NVSHMEM tuning to data collection, remote transport, NIC-to-PE mapping, or TMA. Do not use for unrelated CUDA, NCCL, or application tuning.
NVIDIA/nvshmem
SkillsMP has collected 9 skills from NVIDIA/nvshmem. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 9
- GitHub stars
- 578
- GitHub forks
- 104
Skills in this repository
Showing 9 of 9 collected skills.
Collect and package NVSHMEM put/get bandwidth, latency, and other perftest results with system and topology evidence for performance sanity checks.
Select an NVSHMEM remote transport from target system and kernel evidence. Use for inter-node selection, compatibility checks, or configuration.
Recommend NVSHMEM NIC-to-PE mappings and environment exports. Use for HCA selection, multi-NIC configuration, or topology-based mapping diagnostics.
Diagnose NVSHMEM runtime failures and prepare bug reports for launch, initialization, crashes, hangs, correctness, transport, or topology issues.
Prepare or review NVSHMEM CUDA kernels for TMA SMEM registration and direct-SMEM transfers. Do not use for unrelated CUDA tuning.
Guide NVSHMEM beginners through fit assessment, mental models, first C/C++ or Python NVSHMEM programs, compilation, launching, and next steps. Use for onboarding.
Plan and validate NVSHMEM and NVSHMEM4Py installations. Use for package, container, or source deployments.
Find version-aware official NVSHMEM and NVSHMEM4Py documentation for releases, installation, APIs, runtime settings, transports, containers, and troubleshooting.