Skip to main content

apply-inference-optimizations

Apply FlashDreams-style inference speedups to model integrations after a baseline exists: bounded windows and fixed K/V caches, cache/decode overlap, `torch.compile`, CUDA graph capture, attention backend checks, decoder layout or replacement, transfer/materialization changes, and ordered presentation tuning. Use when porting known optimizations into a runner, demo, serving adapter, or downstream integration while preserving quality and reset behavior.

Jump to install

Source facts

Repository
NVIDIA/flashdreams
Last source activity
July 14, 2026 at 22:51
Detected SKILL.md language
English
Stars
467
Forks
48

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.