| name | research-deepspeed |
| description | Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention |
| license | MIT |
| tags | ["deepspeed","distributed-training","zero","pipeline-parallelism","mixed-precision"] |
| version | 1.0.0 |
| author | Orchestra Research |
| dependencies | ["deepspeed","torch","transformers","accelerate"] |
Progressive disclosure index
The complete skill instructions are preserved in the ordered references below.
Open the part whose headings match the current task; read all parts in order when
the task spans sections or requires the complete procedure.
Detailed instructions