Skip to main content

aws-samples/sample-rlinf-on-eks

SkillsMP has collected 19 skills from aws-samples/sample-rlinf-on-eks. Open a skill to review its source and details.

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
19
GitHub stars
1
GitHub forks
0

Showing 19 of 19 collected skills.

occupation
Software Developers
description

Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements

updated
occupation
Network & Computer Systems Administrators
description

Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes

updated
occupation
Network & Computer Systems Administrators
description

Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management

updated
occupation
Network & Computer Systems Administrators
description

Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads

updated
occupation
Data Scientists
description

RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS

updated
occupation
Software Quality Assurance Analysts & Testers
description

Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification

updated
occupation
Software Developers
description

Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS

updated
occupation
Software Quality Assurance Analysts & Testers
description

RLinf: Validate RLinf training examples end-to-end from container tests to single-step training

updated
occupation
Software Developers
description

RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS

updated
occupation
Network & Computer Systems Administrators
description

Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client

updated
occupation
Software Developers
description

Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI

updated
occupation
Network & Computer Systems Administrators
description

Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch

updated
occupation
Network & Computer Systems Administrators
description

Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training

updated
occupation
Data Scientists
description

RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks

updated
occupation
Software Developers
description

RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP

updated
occupation
Network & Computer Systems Administrators
description

Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights

updated
occupation
Network & Computer Systems Administrators
description

Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver

updated
occupation
Software Developers
description

Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3

updated
occupation
Network & Computer Systems Administrators
description

Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access

updated
Showing 19 of 19 collected skills.