Skip to main content

vllm-project/vllm-omni

SkillsMP ha recopilado 9 skills de vllm-project/vllm-omni. Abre una skill para revisar su origen y sus detalles.

Última actividad de origen registrada
Catálogo de SkillsMP actualizado
skills recopiladas
9
Estrellas en GitHub
6569
Forks en GitHub
1601

Skills en este repositorio

1 categorías ocupacionales · 44% clasificado

Mostrando 9 de 9 skills recopiladas.

ocupación
sin clasificar
descripción

Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including Cache-DiT acceleration and parallelism support (TP, SP/USP, CFG-Parallel, HSDP). Use when integrating a new diffusion model, porting…

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. Use when adding a new TTS model, wiring stage separation for speech synthesis, enabling…

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Generate and run tests for vllm-project/vllm-omni with CI-aligned levels and markers; wire new tests into Buildkite (test-ready.yml for L1/L2, test-merge.yml for L3, test-nightly.yml for L4). On completion, always provide copy-paste local and CI-like pytest…

Idioma del texto original: inglés

actualizado
ocupación
Desarrolladores de software
descripción

Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes. Use for audits of dead, duplicated, speculative, over-generalized, unnecessarily defensive, or hand-rolled code. Use review-pr…

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness. Use when the user says "precheck", "self review", "pre-submit…

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings. Use for default, detailed, or repeat maintainer reviews;…

Idioma del texto original: inglés

actualizado
ocupación
Desarrolladores de software
descripción

Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models. Use when choosing or adding methods such as fp8, int8, gguf, mxfp8, mxfp4, mxfp4_dualscale, ModelOpt, AutoRound, INC, msModelSlim, awq, or gptq; debugging quantized…

Idioma del texto original: inglés

actualizado
ocupación
Desarrolladores de software
descripción

Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation. Use when Codex is asked to analyze profiling traces, choose parallel strategies, inspect torch profiler trace.json or trace.json.gz timelines,…

Idioma del texto original: inglés

actualizado
ocupación
Desarrolladores de software
descripción

Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

Idioma del texto original: inglés

actualizado
Mostrando 9 de 9 skills recopiladas.