Skip to main content

ray

Expert guidance on Ray distributed computing framework for scaling AI and Python applications. Use for: distributed training, hyperparameter tuning with Ray Tune, reinforcement learning with Ray RLlib, model serving with Ray Serve, batch inference, scaling Python applications from single machines to clusters, and Ray Core tasks and actors.

Ir para a instalação

Informações da origem

Repositório
NeuralBlitz/Mito
Última atividade na origem
22 de março de 2026 às 13:29
Idioma detectado do SKILL.md
inglês
Estrelas
0
Forks
0

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
ray
description
Expert guidance on Ray distributed computing framework for scaling AI and Python applications. Use for: distributed training, hyperparameter tuning with Ray Tune, reinforcement learning with Ray RLlib, model serving with Ray Serve, batch inference, scaling Python applications from single machines to clusters, and Ray Core tasks and actors.
license
MIT
compatibility
opencode
metadata
{"audience":"ml-engineers, data-scientists","category":"distributed-computing","tags":["ray","distributed-computing","machine-learning","scaling"]}
# Ray — Distributed Computing for AI Covers: **Ray Core · Distributed Training · Ray Tune · Ray RLlib · Ray Serve · Ray Datasets** ----- ## Understanding Ray ### What is Ray? Ray is a distributed computing framework designed to scale Python applications from a single machine to a large cluster. Originally developed at UC Berkeley's RISELab, Ray provides a unified set of primitives for building distributed applications, making it particularly well-suited for machine learning workloads. Ray addresses the challenge of distributing Python code across multiple machines without requiring developers to become distributed systems experts. It provides a simple API that abstracts away the complexities of cluster management, fault tolerance, and inter-process communication. **Core Components:** - **Ray Core** — The foundational distributed computing primitives: tasks, actors, and objects. - **Ray Tune** — Scalable hyperparameter optimization with support for various search algorithms and schedulers. - **Ray RLlib** — Reinforcement learning library with distributed training support and many built-in algorithms. - **Ray Serve** — Scalable model serving framework for building production inference services. - **Ray Datasets** — Distributed data processing for ML workloads, integrated with popular data formats. ### When to Use Ray | Use Case | Recommended Ray Component | |----------|---------------------------| | Parallelizing Python functions | Ray Tasks | | Stateful services with shared memory | Ray Actors | | Hyperparameter tuning | Ray Tune | | Reinforcement learning | Ray RLlib | | Model serving | Ray Serve | | Large-scale data processing | Ray Datasets | | Distributed training | Ray Train | ----- ## Ray Core Fundamentals ### Ray Tasks (Stateless Functions) Ray tasks are stateless functions that can be executed asynchronously across a Ray cluster. They're ideal for parallelizing pure functions that don't require maintaining state between invocations. ```python import ray import time from typing import List # Initialize Ray ray.init() # Define a function to be executed remotely @ray.remote def download_url(url: str) -> str: """Simulate downloading content from a URL""" import random # Simulate network delay time.sleep(random.uniform(0.5, 2.0)) return f"Content from {url}" @ray.remote def process_content(content: str) -> int: """Process content and return result""" # Simulate some processing return len(content) # Execute tasks remotely urls = [ "https://example.com/article1", "https://example.com/article2", "https://example.com/article3", "https://example.com/article4", "https://example.com/article5" ] # Launch all downloads in parallel download_futures = [download_url.remote(url) for url in urls] # Wait for all to complete and get results contents = ray.get(download_futures) # Chain with processing process_futures = [process_content.remote(content) for content in contents] results = ray.get(process_futures) print(f"Processed {len(results)} URLs, results: {results}") # For larger workflows, use ray.get with timeout # results = ray.get(futures, timeout=30) ``` ### Ray Actors (Stateful Services) Ray actors are stateful workers that maintain state across method calls. They're perfect for scenarios where you need shared state, counters, or services that benefit from keeping data in memory. ```python import ray from dataclasses import dataclass ray.init() @dataclass class TrainingMetrics: epoch: int loss: float accuracy: float @ray.remote class ParameterServer: """Parameter server for distributed training""" def __init__(self): self.parameters = {} self.version = 0 def get_parameters(self, key: str) -> dict: return self.parameters.get(key, {}) def update_parameters(self, key: str, updates: dict): if key not in self.parameters: self.parameters[key] = {} self.parameters[key].update(updates) self.version += 1 return self.version @ray.remote class DataLoader: """Distributed data loader actor""" def __init__(self, data_path: str, batch_size: int = 32): self.data_path = data_path self.batch_size = batch_size self.current_index = 0 self.data = self._load_data() def _load_data(self): # Load data (simplified) import numpy as np return np.random.randn(1000, 10) def get_batch(self) -> List[List[float]]: start = self.current_index end = min(start + self.batch_size, len(self.data)) if start >= len(self.data): self.current_index = 0 start = 0 end = min(self.batch_size, len(self.data)) batch = self.data[start:end].tolist() self.current_index = end return batch def reset(self): self.current_index = 0 return True # Create actors ps = ParameterServer.remote() loader = DataLoader.remote("/data/train.csv", batch_size=32) # Interact with actors ray.get(ps.update_parameters.remote("model_v1", {"weights": [1, 2, 3]})) params = ray.get(ps.get_parameters.remote("model_v1")) batch = ray.get(loader.get_batch.remote()) # Get actor handle from within another actor # Inside another actor: # ray.get(loader.get_batch.remote()) ``` ### Resource Management ```python import ray ray.init(num_cpus=4, num_gpus=1) # Specify resource requirements for tasks @ray.remote(num_cpus=2, num_gpus=0.5) def train_model_batch(data, config): """Train a model on a batch using GPU""" import torch # Training code here return {"loss": 0.5, "accuracy": 0.9} # Use custom resources for specific hardware @ray.remote(resources={"GPU_TYPE_A100": 1}) def run_on_a100(): return "Running on A100 GPU" # Actor resource management @ray.remote(num_cpus=1, num_gpus=0.25) class GPUWorker: def __init__(self): self.device = "cuda:0" def process(self, data): # Process on GPU return len(data) # Automatic resource allocation with placement groups placement_group = ray.util.placement_group( [{"CPU": 2, "GPU": 1}], strategy="PACK" ) @ray.remote(num_cpus=2, num_gpus=1) def scheduled_task(): return "Running in placement group" # Wait for placement group to be ready ray.get(scheduled_task.options( placement_group=placement_group, placement_group_bundle_index=0 ).remote()) ``` ----- ## Ray Tune (Hyperparameter Optimization) ### Basic Configuration ```python import ray from ray import tune ray.init(num_cpus=4) def trainable_function(config: dict): """Training function to optimize""" import time import numpy as np # Extract hyperparameters learning_rate = config.get("learning_rate", 0.01) batch_size = config.get("batch_size", 32) hidden_units = config.get("hidden_units", 64) dropout = config.get("dropout", 0.1) # Simulate training metrics = {"loss": 1.0, "accuracy": 0.1} for step in range(100): # Simulate training loop loss = np.random.uniform(0.1, 1.0) * (learning_rate / 0.01) accuracy = 1.0 - loss + np.random.uniform(-0.1, 0.1) # Report metrics to Tune tune.report( loss=loss, accuracy=max(0, min(1, accuracy)), step=step, learning_rate=learning_rate ) time.sleep(0.1) return {"loss": loss, "accuracy": accuracy} # Run basic search analysis = tune.run( trainable_function, config={ "learning_rate": tune.choice([0.001, 0.01, 0.1]), "batch_size": tune.choice([16, 32, 64, 128]), "hidden_units": tune.choice([32, 64, 128, 256]), "dropout": tune.uniform(0.0, 0.5) }, num_samples=20, metric="loss", mode="min", verbose=1 ) # Get best configuration best_config = analysis.best_config best_loss = analysis.best_result["loss"] print(f"Best config: {best_config}") print(f"Best loss: {best_loss}") ``` ### Advanced Search Strategies ```python from ray import tune from ray.tuners import Tuner import ray ray.init(num_cpus=8) # Bayesian Optimization with Optuna def optimize_with_optuna(search_space): from ray.tune.search.optuna import OptunaSearch optuna_search = OptunaSearch( metric="accuracy", mode="max", points_to_evaluate=[ {"learning_rate": 0.01, "batch_size": 32} ] ) return tune.Tuner( trainable_function, tune_config=tune.TuneConfig( search_alg=optuna_search, num_samples=50, time_budget_s=3600 # 1 hour ), param_space=search_space ) # HyperBand for early stopping def run_with_hyperband(search_space): return tune.Tuner( trainable_function, tune_config=tune.TuneConfig( scheduler=tune.schedulers.HyperBandScheduler( time_attr="step", max_t=100, stop_last_triple=True ), num_samples=100, max_concurrent_trials=8 ), param_space=search_space ) # Population-Based Training (PBT) def run_with_pbt(search_space): return tune.Tuner( trainable_function, tune_config=tune.TuneConfig( scheduler=tune.schedulers.PopulationBasedTraining( time_attr="step", perturbation_interval=10, hyperparam_mutations={ "learning_rate": tune.loguniform(1e-4, 1e-2), "dropout": tune.uniform(0.0, 0.5) } ), num_samples=20 ), param_space=search_space ) # Define search space search_space = { "learning_rate": tune.loguniform(1e-4, 1e-2), "batch_size": tune.choice([16, 32, 64, 128]), "hidden_units": tune.choice([64, 128, 256, 512]), "dropout": tune.uniform(0.0, 0.5), "optimizer": tune.choice(["adam", "sgd", "rmsprop"]), "activation": tune.choice(["relu", "tanh", "gelu"]) } # Run with PBT (often best for neural network tuning) analysis = run_with_pbt(search_space).fit() ``` ### Custom Trainables ```python from ray import tune import numpy as np class CustomTrainable(tune.Trainable): """Custom trainable with state management""" def setup(self, config): """Initialize training state""" self.model = Model(config) self.optimizer = Optimizer(self.model, config["learning_rate"]) self.step_count = 0 # Load data self.train_data = np.random.randn(1000, 10) self.labels = np.random.randint(0, 2, 1000) def step(self): """Execute one training step""" # Sample batch indices = np.random.choice(len(self.train_data), 32) batch_x = self.train_data[indices] batch_y = self.labels[indices] # Train loss = self.model.train_step(batch_x, batch_y) accuracy = self.model.evaluate(batch_x, batch_y) self.step_count += 1 return { "loss": loss, "accuracy": accuracy, "step": self.step_count } def save_checkpoint(self, checkpoint_dir): """Save checkpoint""" path = f"{checkpoint_dir}/model.npy" np.save(path, self.model.weights) return path def load_checkpoint(self, checkpoint_path): """Load checkpoint""" self.model.weights = np.load(checkpoint_path) class Model: def __init__(self, config): self.hidden_units = config.get("hidden_units", 64) self.weights = np.random.randn(10, self.hidden_units) def train_step(self, x, y): """Simulate training""" return np.random.uniform(0.1, 1.0) def evaluate(self, x, y): return np.random.uniform(0.5, 0.95) class Optimizer: def __init__(self, model, lr): self.model = model self.lr = lr # Run custom trainable analysis = tune.run( CustomTrainable, config={ "learning_rate": tune.choice([0.001, 0.01, 0.1]), "hidden_units": tune.choice([32, 64, 128]) }, num_samples=10, checkpoint_at_end=True,
Ver no GitHub
Este SKILL.md e muito grande, entao o SkillsMP mostra aqui apenas a primeira secao. Ver no GitHub