Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
لغة النص الأصلي: الإنجليزية
القائمة
جمع SkillsMP عدد ١٩ من skills من aws-samples/sample-rlinf-on-eks. افتح أي skill لمراجعة مصدره وتفاصيله.
عرض ١٩ من أصل ١٩ skills مجمعة.
Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
لغة النص الأصلي: الإنجليزية
Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes
لغة النص الأصلي: الإنجليزية
Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management
لغة النص الأصلي: الإنجليزية
Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads
لغة النص الأصلي: الإنجليزية
RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS
لغة النص الأصلي: الإنجليزية
Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification
لغة النص الأصلي: الإنجليزية
Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS
لغة النص الأصلي: الإنجليزية
RLinf: Validate RLinf training examples end-to-end from container tests to single-step training
لغة النص الأصلي: الإنجليزية
RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS
لغة النص الأصلي: الإنجليزية
Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client
لغة النص الأصلي: الإنجليزية
Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI
لغة النص الأصلي: الإنجليزية
Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch
لغة النص الأصلي: الإنجليزية
Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training
لغة النص الأصلي: الإنجليزية
RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks
لغة النص الأصلي: الإنجليزية
RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP
لغة النص الأصلي: الإنجليزية
Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights
لغة النص الأصلي: الإنجليزية
Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver
لغة النص الأصلي: الإنجليزية
Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3
لغة النص الأصلي: الإنجليزية
Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access
لغة النص الأصلي: الإنجليزية