Skip to main content
cfregly
Perfil de criador do GitHub

cfregly

Visão por repositório de 69 skills coletadas em 5 repositórios do GitHub.

skills coletadas
69
repositórios
5
atualizado
17 de ago. de 2026
explorador de repositórios

Repositórios e skills representativas

inference-aa-workload
Cientistas de dados

Reproduce the Artificial Analysis (AA) language-model performance workload shapes against an OpenAI-compatible chat endpoint using NVIDIA AIPerf. Drives the three AA text shapes (1k input / >=1k answer, 10k / >=1.5k, 100k / >=2k) with temperature 0, top_p 1,…

17 de ago. de 2026
inference-dcgm-correlate
Cientistas de dados

Correlate DCGM Prometheus byte-traffic counters with an inference-perf-bench sweep window to compute byte-grounded workload-level Speed-of-Light. The third tier of the SoL rigor hierarchy (after zymtrace sample-share and ncu per-kernel arithmetic intensity).…

17 de ago. de 2026
inference-decode-step-budget
Analistas de garantia de qualidade de software e testadores

Measure the low-concurrency (c=1..c=8) decode hot-path of a live vLLM pod FAST and CORRECTLY: where each token's time actually goes (GPU-busy vs host-idle gap vs comm), whether the workload is kernel-bound / host-bound / comm-bound, and what the addressable…

17 de ago. de 2026
inference-graph-diff
Analistas de garantia de qualidade de software e testadores

Diff the compiled FX / Inductor graphs across two vLLM versions or two helm configs to see exactly which fused kernels / passes / partitions changed. Uses `torch._dynamo.explain` + `torch.compile`'s graph-dump hooks. Useful when you anticipate a…

17 de ago. de 2026
inference-workload-profile
Cientistas de dados

Profile live inference traffic into a token/shape distribution artifact that drives profile-matched speculative-decoding draft training -- an analog of Fireworks FireOptimizer's "profile-driven customization" (the documented source of its higher draft…

17 de ago. de 2026
mirage-graph-coverage
Analistas de garantia de qualidade de software e testadores

Read-only coverage auditor for the mirage / MPK persistent-megakernel's GENERATED task graph: cross-references DECLARED tensors (all_tensors[...] in kernel_N.cu) against tensors CONSUMED by tasks (inputs/outputs base_ptr in task_graph_N.json) and flags any…

17 de ago. de 2026
analyze-zymtrace-workload
Cientistas de dados

Investigate a GPU or CPU workload through the zymtrace MCP. The MCP does most of the analysis. This skill enforces the cross-view discipline -- always pull the matching opposite-side flamegraph (CPU for GPU workloads, GPU for CPU workloads) with the same…

16 de ago. de 2026
evidence-bundle-init
Cientistas de dados

Scaffold a new evidence bundle directory ready for reproducibility-grade evidence capture: SOURCE.md (operator + cluster + git SHA + UTC timestamp), summary.md (verdict skeleton), commands/ (for the four-file .cmd/.stdout/.stderr/.exit tuple capture per shell…

16 de ago. de 2026
Mostrando 8 de 33 skills coletadas.
inference-spec-decode-service
Desenvolvedores de software

Closed-loop speculative-decoding-as-a-service orchestrator -- a self-hosted analog of Fireworks FireOptimizer's adaptive speculative execution. Composes the per-phase skills into one profile-matched loop: profile live traffic (inference-workload-profile) ->…

18 de jun. de 2026
analyze-zymtrace-workload
Desenvolvedores de software

Investigate a GPU or CPU workload through the zymtrace MCP. The MCP does most of the analysis. This skill enforces the cross-view discipline -- always pull the matching opposite-side flamegraph (CPU for GPU workloads, GPU for CPU workloads) with the same…

14 de jun. de 2026
evidence-bundle-init
Desenvolvedores de software

Scaffold a new evidence bundle directory ready for reproducibility-grade evidence capture: SOURCE.md (operator + cluster + git SHA + UTC timestamp), summary.md (verdict skeleton), commands/ (for the four-file .cmd/.stdout/.stderr/.exit tuple capture per shell…

14 de jun. de 2026
inference-aa-workload
Desenvolvedores de software

Reproduce the Artificial Analysis (AA) language-model performance workload shapes against an OpenAI-compatible chat endpoint using NVIDIA AIPerf. Drives the three AA text shapes (1k input / >=1k answer, 10k / >=1.5k, 100k / >=2k) with temperature 0, top_p 1,…

14 de jun. de 2026
inference-capacity-sizing
Desenvolvedores de software

SLA-first GPU capacity sizing for a serving deployment: given a tokens-per-minute (TPM) target AND the interactivity SLA (output tokens/s/user), compute the pods and GPUs needed from a model's measured tok/s/user-vs-concurrency curve. Sizing MUST start from…

14 de jun. de 2026
inference-dcgm-correlate
Desenvolvedores de software

Correlate DCGM Prometheus byte-traffic counters with an inference-perf-bench sweep window to compute byte-grounded workload-level Speed-of-Light. The third tier of the SoL rigor hierarchy (after zymtrace sample-share and ncu per-kernel arithmetic intensity).…

14 de jun. de 2026
inference-decode-step-budget
Desenvolvedores de software

Measure the low-concurrency (c=1..c=8) decode hot-path of a live vLLM pod FAST and CORRECTLY: where each token's time actually goes (GPU-busy vs host-idle gap vs comm), whether the workload is kernel-bound / host-bound / comm-bound, and what the addressable…

14 de jun. de 2026
inference-fleet-leaderboard
Desenvolvedores de software

Render cross-model fleet leaderboards from the local perf-report campaigns in ONE command (`perftunereport fleet_leaderboard`): a latency tier (aa-1k/10k/100k tok/s/user + TTFT + cost at c=1/c=10), a throughput tier (peak tok/s/GPU per (model,quant,TP) with…

14 de jun. de 2026
Mostrando 8 de 32 skills coletadas.
Mostrando 5 de 5 repositórios
Todos os repositórios foram exibidos