#001interview4AI1 skills180atualizado 2026-03-13100% do criadorskillocupaçãodescriçãoatualizadollm-inference-performance-interviewDesenvolvedores de softwareCoach technical interviews for LLM inference, model compression, and HPC-oriented deployment. Use when Codex needs to answer or refine questions about Transformer inference internals, KV cache, attention variants, quantization, pruning, distillation, sparsity, vLLM, TensorRT-LLM, TGI, ONNX Runtime, llama.cpp, CUDA or Triton kernels, FlashAttention, PagedAttention, Roofline reasoning, speculative decoding, continuous batching, or multi-GPU serving. Also use when the user wants mock interview answers, back-of-the-envelope performance calculations, bottleneck analysis, or framework and hardware tradeoff explanations.2026-03-13Mostrando as 1 principais de 1 skills coletadas neste repositório.Carregar mais 0 skillsCarregando skills...