Skip to main content

llm-inference-performance-interview

스타18
포크0
업데이트2026년 3월 13일 01:58

Coach technical interviews for LLM inference, model compression, and HPC-oriented deployment. Use when Codex needs to answer or refine questions about Transformer inference internals, KV cache, attention variants, quantization, pruning, distillation, sparsity, vLLM, TensorRT-LLM, TGI, ONNX Runtime, llama.cpp, CUDA or Triton kernels, FlashAttention, PagedAttention, Roofline reasoning, speculative decoding, continuous batching, or multi-GPU serving. Also use when the user wants mock interview answers, back-of-the-envelope performance calculations, bottleneck analysis, or framework and hardware tradeoff explanations.

설치

Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.

파일 탐색기
4 개 파일
SKILL.md
readonly