Skip to main content

aws-samples/sample-rlinf-on-eks

SkillsMP는 aws-samples/sample-rlinf-on-eks에서 19개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.

최근 기록된 소스 활동
SkillsMP 카탈로그 업데이트
수집된 skills
19
GitHub 스타
1
GitHub 포크
0

수집된 skill 19개 중 19개를 표시합니다.

직업 분류
소프트웨어 개발자
설명

Select optimal EC2 GPU instance types for Physical AI RL training based on VRAM, interconnect, and cost requirements

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Leverage Kubernetes-native features for GPU training including GPU Operator, topology scheduling, and priority classes

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Deploy and manage Physical AI RL training jobs on EKS including CodeBuild CI, GitOps, and experiment management

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Provision an Amazon EKS cluster with GPU node groups, EFA networking, and OIDC for Physical AI RL training workloads

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

RLinf: Download, preprocess, and stage datasets and model weights for RLinf training on EKS

원문 언어: 영어

업데이트
직업 분류
소프트웨어 품질 보증 분석가·테스터
설명

Validate the Physical AI RL training stack end-to-end from infrastructure smoke tests to training step verification

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Construct Kubernetes Jobs, StatefulSets, and Services for GPU-accelerated RL training workloads on EKS

원문 언어: 영어

업데이트
직업 분류
소프트웨어 품질 보증 분석가·테스터
설명

RLinf: Validate RLinf training examples end-to-end from container tests to single-step training

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

RLinf: Integrate Ray, veRL, vLLM, and KubeRay into the RLinf training stack on EKS

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Configure AMI selection and GPU node bootstrap including NVIDIA drivers, EFA, and Lustre client

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Build container images for Physical AI RL training including dependency conflict management, patching, and CodeBuild CI

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Monitor GPU utilization, training progress, and costs using DCGM, Prometheus, Grafana, MLflow, and CloudWatch

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Configure Elastic Fabric Adapter networking for high-performance NCCL allreduce in multi-node GPU training

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

RLinf: Evaluate training quality and infrastructure performance using LIBERO, RoboTwin, and throughput benchmarks

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

RLinf: Adapt RLinf training scripts for distributed execution on Amazon EKS with veRL, Ray, and FSDP

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Deploy Amazon FSx for Lustre as shared high-performance storage for datasets, checkpoints, and model weights

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Mount S3 buckets as POSIX-like filesystems in training pods via the Mountpoint for Amazon S3 CSI driver

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Use Amazon S3 Connector for PyTorch to read datasets and save checkpoints directly to S3

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Mount S3 buckets as native NFS file systems on EKS via Amazon S3 Files for shared, low-latency file access

원문 언어: 영어

업데이트
수집된 skill 19개 중 19개를 표시합니다.