Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

model-serving-infrastructure

النجوم٠
التفرعات٠
آخر تحديث٥ يوليو ٢٠٢٦ في ١٤:٠٤

Load when designing, costing, or debugging production model inference infrastructure — choosing managed API vs serverless GPU vs dedicated cluster, GPU utilization economics and break-even math, autoscaling and cold starts, token streaming/SSE and retry semantics, multi-model and LoRA serving, model rollout, or serving observability (TTFT, queue depth, KV cache).

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly