Skip to main content

abhiram1809/inference-god-mode

Dernière activité source enregistrée
Catalogue SkillsMP mis à jour
skills collectés
1
Étoiles GitHub
0
Forks GitHub
0

Chargement du README du dépôt…

Publications

Tout voir →
Abhi
Abhi
@abhiram_ai_guy
Annonce d’un SkillOptimiser l’inférence des modèles de langage
Résumé de la publication · anglais

A skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.

Skills dans ce dépôt

classification en attente

Affichage de 1 skills collectés sur 1.

métier
non classé
description

Plan, deploy, benchmark, and tune self-hosted open-weight LLM inference on one machine or Kubernetes GPU clusters, including quantization, full-context capacity, OpenAI or Anthropic APIs, cache-aware routing, and prefill/decode disaggregation.

Langue du texte source : anglais

mis à jour
Affichage de 1 skills collectés sur 1.

Installer avec un assistant IA

Copiez ce prompt dans l’assistant IA que vous utilisez.