| name | ray-distributed-trainer |
| description | Distributed computing skill using Ray for parallel training, hyperparameter search, and resource management. |
| allowed-tools | ["Read","Write","Bash","Glob","Grep"] |
| graph | {"domains":["domain:data-science"],"specializations":["specialization:data-science-ml"],"skillAreas":["skill-area:machine-learning-frameworks","skill-area:hyperparameter-tuning-experiment-management"],"roles":["role:ml-engineer","role:ml-ops-engineer"],"workflows":["workflow:ml-model-lifecycle"]} |
ray-distributed-trainer
Overview
Distributed computing skill using Ray for parallel training, hyperparameter search, and resource management across clusters.
Capabilities
- Ray Train for distributed training
- Ray Tune for hyperparameter search at scale
- Cluster resource management
- Fault tolerance and checkpointing
- Actor-based parallelism
- Integration with PyTorch and TensorFlow
- Elastic training support
- Multi-node orchestration
Target Processes
- Distributed Training Orchestration
- AutoML Pipeline Orchestration
- Model Training Pipeline
Tools and Libraries
- Ray
- Ray Train
- Ray Tune
- Ray Cluster
Input Schema
{