Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
aws-samples/sample-rlinf-on-eks
SkillsMP has collected 19 skills from aws-samples/sample-rlinf-on-eks. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 19
- GitHub stars
- 1
- GitHub forks
- 0
Skills in this repository
Showing 19 of 19 collected skills.
Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes
Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management
Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads
RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS
Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification
Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS
RLinf: Validate RLinf training examples end-to-end from container tests to single-step training
RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS
Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client
Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI
Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch
Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training
RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks
RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP
Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights
Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver
Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3
Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access