#001interview4AI1 skills180mis à jour 2026-03-13100% du créateurskillmétierdescriptionmis à jourllm-inference-performance-interviewDéveloppeurs de logicielsCoach technical interviews for LLM inference, model compression, and HPC-oriented deployment. Use when Codex needs to answer or refine questions about Transformer inference internals, KV cache, attention variants, quantization, pruning, distillation, sparsity, vLLM, TensorRT-LLM, TGI, ONNX Runtime, llama.cpp, CUDA or Triton kernels, FlashAttention, PagedAttention, Roofline reasoning, speculative decoding, continuous batching, or multi-GPU serving. Also use when the user wants mock interview answers, back-of-the-envelope performance calculations, bottleneck analysis, or framework and hardware tradeoff explanations.2026-03-13Affichage des 1 principaux skills collectés sur 1 dans ce dépôt.Charger 0 skills de plusChargement des skills...