Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
Langue du texte source : anglais
Menu
SkillsMP a collecté 19 skills depuis aws-samples/sample-rlinf-on-eks. Ouvrez un skill pour examiner sa source et ses détails.
Affichage de 19 skills collectés sur 19.
Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
Langue du texte source : anglais
Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes
Langue du texte source : anglais
Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management
Langue du texte source : anglais
Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads
Langue du texte source : anglais
RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS
Langue du texte source : anglais
Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification
Langue du texte source : anglais
Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS
Langue du texte source : anglais
RLinf: Validate RLinf training examples end-to-end from container tests to single-step training
Langue du texte source : anglais
RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS
Langue du texte source : anglais
Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client
Langue du texte source : anglais
Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI
Langue du texte source : anglais
Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch
Langue du texte source : anglais
Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training
Langue du texte source : anglais
RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks
Langue du texte source : anglais
RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP
Langue du texte source : anglais
Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights
Langue du texte source : anglais
Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver
Langue du texte source : anglais
Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3
Langue du texte source : anglais
Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access
Langue du texte source : anglais