Skip to main content
Jeden Skill in Manus ausführen
mit einem Klick

modal-llm-serving

Sterne4
Forks1
Aktualisiert8. März 2026 um 17:38

Use this skill for self-hosted open-weight text LLM serving on Modal with vLLM or SGLang. Trigger when the user wants an OpenAI-compatible endpoint for a model they host, needs to tune cold starts, latency, concurrency, tensor parallelism, or throughput for text generation, or wants offline batch text inference with the vLLM engine. Do not use it for embeddings, generic Hugging Face pipeline serving, diffusion, fine-tuning, RAG, or apps that only call OpenAI's hosted API.

Installation

Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.

Datei-Explorer
12 Dateien
SKILL.md
readonly