Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
Quellsprache: Englisch
Menü
SkillsMP hat 19 Skills aus aws-samples/sample-rlinf-on-eks gesammelt. Öffne einen Skill, um Quelle und Details zu prüfen.
Es werden 19 von 19 gesammelten Skills angezeigt.
Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
Quellsprache: Englisch
Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes
Quellsprache: Englisch
Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management
Quellsprache: Englisch
Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads
Quellsprache: Englisch
RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS
Quellsprache: Englisch
Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification
Quellsprache: Englisch
Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS
Quellsprache: Englisch
RLinf: Validate RLinf training examples end-to-end from container tests to single-step training
Quellsprache: Englisch
RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS
Quellsprache: Englisch
Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client
Quellsprache: Englisch
Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI
Quellsprache: Englisch
Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch
Quellsprache: Englisch
Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training
Quellsprache: Englisch
RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks
Quellsprache: Englisch
RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP
Quellsprache: Englisch
Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights
Quellsprache: Englisch
Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver
Quellsprache: Englisch
Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3
Quellsprache: Englisch
Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access
Quellsprache: Englisch