Skip to main content

aws-samples/sample-rlinf-on-eks

SkillsMP は aws-samples/sample-rlinf-on-eks から 19 件の skill を収集しています。skill を開くとソースと詳細を確認できます。

記録された最新のソース活動
SkillsMP カタログ更新
収集済み skills
19
GitHub スター
1
GitHub フォーク
0

収集済み skill 19 件中 19 件を表示しています。

職業分類
ソフトウェア開発者
説明

Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads

原文の言語: 英語

更新
職業分類
データサイエンティスト
説明

RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS

原文の言語: 英語

更新
職業分類
ソフトウェア品質保証アナリスト・テスター
説明

Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS

原文の言語: 英語

更新
職業分類
ソフトウェア品質保証アナリスト・テスター
説明

RLinf: Validate RLinf training examples end-to-end from container tests to single-step training

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training

原文の言語: 英語

更新
職業分類
データサイエンティスト
説明

RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access

原文の言語: 英語

更新
収集済み skill 19 件中 19 件を表示しています。