| How do I run inference? (single-GPU, multi-GPU) | README.md § Inference |
| Which model should I use? (Nano vs Super, memory, shift) | README.md § Models |
| Which modality? (t2i, t2v, i2v, examples) | README.md § Modalities |
| What parallelism preset? (latency vs throughput) | README.md § Inference |
| What input fields are available? (prompt, vision_path, num_frames, ...) | docs/inference.md § Sample Arguments |
| What are the default parameter values? | cosmos_framework/inference/defaults/<model_mode>/sample_args.json (per-modality JSON) |
| How do I use custom defaults? | docs/inference.md § Custom Defaults |
| How do I override a parameter? (precedence) | docs/faq.md § How do I override a default parameter? |
| What is the shift parameter? | docs/faq.md § What is the shift parameter? |
| How many frames can I generate? (resolution caps) | docs/faq.md § How many frames can I generate? |
| How do I start Ray Serve / Gradio / submit requests? | docs/faq.md § How do I run online inference with Ray? |
| How do I upsample short prompts? | docs/faq.md § Prompt upsampling |
| How do I use the low-level API? (examples/) | examples/inference.py (model API) / examples/inference_pipeline.py (pipeline API) |
| All CLI flags | uv run --all-extras --group=cu130 python -m cosmos_framework.scripts.inference --help |