Skip to main content

whitecircle/halo

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
10
GitHub stars
816
GitHub forks
38

Loading the repository README…

Skills in this repository

classification pending

Showing 10 of 10 collected skills.

occupation
unclassified
description

Use when writing, editing, or reviewing any page under agent-docs/, or when a code change in src/ needs its owning doc page updated (doc actualization). Authors and maintains the Halo agent-docs/ tree and enforces the anti-AI-slop style charter.

updated
occupation
unclassified
description

Assemble and run a Halo training or test job inside the prebuilt Docker image. Picks the image (halo:blackwell for B200/B300 vs halo:hopper for H100/H200), sets GPU count, mounts ($(pwd)->/workspace, the host's large scratch volume, ~/.aws), passes --env-file…

updated
occupation
unclassified
description

Recommend throughput/memory levers to raise tokens/s/GPU or cut peak memory for a given Halo training config — and call out the levers that do NOT help at fine-grained MoE shapes (low-precision fp8/fp4 compute, native DeepGEMM, torch.compile on EP MoE,…

updated
occupation
unclassified
description

Recommend a VALID ParallelismConfig and REJECT unsupported EP/CP/TP/ETP/PP combos BEFORE a GPU run. Auto-fire whenever the user: asks how to configure or pick a parallelism mode (expert/context/tensor/expert-tensor/pipeline parallelism, EP/CP/TP/ETP/PP,…

updated
occupation
unclassified
description

Wire up online GRPO (RLVR) or async GRPO with environments (multi-turn tool-use) for Halo: bring up the separate rollout container — vLLM 0.26.0 (cu13), or SGLang 0.5.17 for async GRPO with environments (rollout_backend: sglang, the families its loaders can…

updated
occupation
unclassified
description

Write CPU or GPU correctness/benchmark tests for Halo using the shared harness + manifest. Use when adding a test for a trainer, parallelism mode, kernel, optimizer, model family, or a perf benchmark — covering the support matrix and the required rejection…

updated
occupation
unclassified
description

Add support for a new model family to Halo — the EP (expert-parallel) wrapper, the Liger coverage spec, router-balancing choice, and vendoring (with a removal-on-merge policy). USER-INVOKED ONLY. Use when the user asks to "add a new model", "support <model>…

updated
occupation
unclassified
description

Turn a finished Halo training run into a usable artifact, and resume correctly. USER-INVOKED ONLY. Use when the user asks to "merge EP shards", "load my EP/MoE checkpoint into vLLM", "merge a LoRA/QLoRA adapter", "convert to bf16", "quantize to fp8/fp4",…

updated
occupation
unclassified
description

Get a dataset into the right shape for a Halo training method, and pre-process it (tokenize / pack / shard) when worth it. USER-INVOKED ONLY. Use when the user asks "what format does SFT/DPO/GRPO/reward/embedding expect", "prepare/tokenize/pack/shard my…

updated
occupation
unclassified
description

Diagnose distributed-training failures in Halo — hangs/deadlocks, CUDA OOM, loss=NaN/Inf, DeepEP faults, and NCCL collective timeouts — by routing the symptom to the exact opt-in helper, env var, or known failure mode. Auto-fire when a training run…

updated
Showing 10 of 10 collected skills.

Install with an AI assistant

Copy this prompt into the AI assistant you're using.