Use when writing, editing, or reviewing any page under agent-docs/, or when a code change in src/ needs its owning doc page updated (doc actualization). Authors and maintains the Halo agent-docs/ tree and enforces the anti-AI-slop style charter.
Assemble and run a Halo training or test job inside the prebuilt Docker image. Picks the image (halo:blackwell for B200/B300 vs halo:hopper for H100/H200), sets GPU count, mounts ($(pwd)->/workspace, the host's large scratch volume, ~/.aws), passes --env-file…
Recommend throughput/memory levers to raise tokens/s/GPU or cut peak memory for a given Halo training config — and call out the levers that do NOT help at fine-grained MoE shapes (low-precision fp8/fp4 compute, native DeepGEMM, torch.compile on EP MoE,…
Recommend a VALID ParallelismConfig and REJECT unsupported EP/CP/TP/ETP/PP combos BEFORE a GPU run. Auto-fire whenever the user: asks how to configure or pick a parallelism mode (expert/context/tensor/expert-tensor/pipeline parallelism, EP/CP/TP/ETP/PP,…
Wire up online GRPO (RLVR) or async GRPO with environments (multi-turn tool-use) for Halo: bring up the separate rollout container — vLLM 0.26.0 (cu13), or SGLang 0.5.17 for async GRPO with environments (rollout_backend: sglang, the families its loaders can…
Write CPU or GPU correctness/benchmark tests for Halo using the shared harness + manifest. Use when adding a test for a trainer, parallelism mode, kernel, optimizer, model family, or a perf benchmark — covering the support matrix and the required rejection…
Add support for a new model family to Halo — the EP (expert-parallel) wrapper, the Liger coverage spec, router-balancing choice, and vendoring (with a removal-on-merge policy). USER-INVOKED ONLY. Use when the user asks to "add a new model", "support <model>…
Turn a finished Halo training run into a usable artifact, and resume correctly. USER-INVOKED ONLY. Use when the user asks to "merge EP shards", "load my EP/MoE checkpoint into vLLM", "merge a LoRA/QLoRA adapter", "convert to bf16", "quantize to fp8/fp4",…