Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
Idioma do texto original: inglês
Menu
O SkillsMP coletou 19 skills de aws-samples/sample-rlinf-on-eks. Abra uma skill para revisar a origem e os detalhes.
Mostrando 19 de 19 skills coletadas.
Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
Idioma do texto original: inglês
Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes
Idioma do texto original: inglês
Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management
Idioma do texto original: inglês
Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads
Idioma do texto original: inglês
RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS
Idioma do texto original: inglês
Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification
Idioma do texto original: inglês
Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS
Idioma do texto original: inglês
RLinf: Validate RLinf training examples end-to-end from container tests to single-step training
Idioma do texto original: inglês
RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS
Idioma do texto original: inglês
Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client
Idioma do texto original: inglês
Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI
Idioma do texto original: inglês
Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch
Idioma do texto original: inglês
Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training
Idioma do texto original: inglês
RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks
Idioma do texto original: inglês
RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP
Idioma do texto original: inglês
Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights
Idioma do texto original: inglês
Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver
Idioma do texto original: inglês
Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3
Idioma do texto original: inglês
Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access
Idioma do texto original: inglês