Skip to main content
Exécutez n'importe quel Skill dans Manus
en un clic

modal-llm-serving

Étoiles4
Forks1
Mis à jour8 mars 2026 à 17:38

Use this skill for self-hosted open-weight text LLM serving on Modal with vLLM or SGLang. Trigger when the user wants an OpenAI-compatible endpoint for a model they host, needs to tune cold starts, latency, concurrency, tensor parallelism, or throughput for text generation, or wants offline batch text inference with the vLLM engine. Do not use it for embeddings, generic Hugging Face pipeline serving, diffusion, fine-tuning, RAG, or apps that only call OpenAI's hosted API.

Installation

Installer avec Codex ou Claude Copiez ce prompt, collez-le dans Codex, Claude ou un autre assistant, puis laissez-le vérifier la page du skill et l'installer pour vous.

Explorateur de fichiers
12 fichiers
SKILL.md
readonly