Skip to main content
Manusで任意のスキルを実行
ワンクリックで

modal-llm-serving

スター4
フォーク1
更新日2026年3月8日 17:38

Use this skill for self-hosted open-weight text LLM serving on Modal with vLLM or SGLang. Trigger when the user wants an OpenAI-compatible endpoint for a model they host, needs to tune cold starts, latency, concurrency, tensor parallelism, or throughput for text generation, or wants offline batch text inference with the vLLM engine. Do not use it for embeddings, generic Hugging Face pipeline serving, diffusion, fine-tuning, RAG, or apps that only call OpenAI's hosted API.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

ファイルエクスプローラー
12 ファイル
SKILL.md
readonly