#001interview4AI1 skills180updated 2026-03-13100% of creatorskilloccupationdescriptionupdatedllm-inference-performance-interviewsoftware-developersCoach technical interviews for LLM inference, model compression, and HPC-oriented deployment. Use when Codex needs to answer or refine questions about Transformer inference internals, KV cache, attention variants, quantization, pruning, distillation, sparsity, vLLM, TensorRT-LLM, TGI, ONNX Runtime, llama.cpp, CUDA or Triton kernels, FlashAttention, PagedAttention, Roofline reasoning, speculative decoding, continuous batching, or multi-GPU serving. Also use when the user wants mock interview answers, back-of-the-envelope performance calculations, bottleneck analysis, or framework and hardware tradeoff explanations.2026-03-13Showing top 1 of 1 collected skills in this repository.Load 0 more skillsLoading skills...