Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
原文の言語: 英語
メニュー
SkillsMP は aws-samples/sample-rlinf-on-eks から 19 件の skill を収集しています。skill を開くとソースと詳細を確認できます。
収集済み skill 19 件中 19 件を表示しています。
Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements
原文の言語: 英語
Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes
原文の言語: 英語
Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management
原文の言語: 英語
Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads
原文の言語: 英語
RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS
原文の言語: 英語
Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification
原文の言語: 英語
Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS
原文の言語: 英語
RLinf: Validate RLinf training examples end-to-end from container tests to single-step training
原文の言語: 英語
RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS
原文の言語: 英語
Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client
原文の言語: 英語
Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI
原文の言語: 英語
Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch
原文の言語: 英語
Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training
原文の言語: 英語
RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks
原文の言語: 英語
RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP
原文の言語: 英語
Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights
原文の言語: 英語
Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver
原文の言語: 英語
Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3
原文の言語: 英語
Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access
原文の言語: 英語