Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

serving-systems

النجوم٨١
التفرعات١٧
آخر تحديث٢١ يونيو ٢٠٢٦ في ٠٧:٥٨

LLM and multimodal serving-system development. Activate when a task touches inference servers, latency / throughput / TTFT / TPOT, KV-cache, batching, attention kernels, CUDA graphs, speculative decoding, structured / grammar-constrained output, quantization, MoE serving, prefix caching, multi-modal (vision/speech/image/video) serving, porting a model to vLLM / SGLang / TensorRT-LLM, or serving on AWS Trainium (Neuron, torch-neuronx, NKI kernels).

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
55 ملفات
SKILL.md
readonly