con un clic
vibe-serve
vibe-serve contiene 3 skills recopiladas de uw-syfi, con cobertura ocupacional por repositorio y páginas de detalle dentro del sitio.
Skills en este repositorio
Prepare and open VibeServe pull requests from local repo changes. Use when the user asks to create, open, publish, submit, or draft a PR for this repository, including tasks that need branch hygiene, targeted validation, PR intent reflection, PR template completion, commit/push, or GitHub pull request creation.
Create or update a VibeServe LLM model-serving example bundle for a new Hugging Face model, hardware target, and optimization workload. Use when the user wants to add an example under examples/model-serving, scaffold reference/accuracy_checker/benchmark inputs, adapt existing checker and HTTP benchmark scripts, or start an optimization run from a natural-language workload goal without adding new reusable framework code.
LLM and multimodal serving-system development. Activate when a task touches inference servers, latency / throughput / TTFT / TPOT, KV-cache, batching, attention kernels, CUDA graphs, speculative decoding, structured / grammar-constrained output, quantization, MoE serving, prefix caching, multi-modal (vision/speech/image/video) serving, porting a model to vLLM / SGLang / TensorRT-LLM, or serving on AWS Trainium (Neuron, torch-neuronx, NKI kernels).