Skip to main content
cfregly
GitHub クリエイタープロフィール

cfregly

5 件の GitHub リポジトリにある 69 件の収集済み skills をリポジトリ単位で表示します。

収集済み skills
69
リポジトリ
5
更新
2026年8月17日
リポジトリエクスプローラー

リポジトリと代表的な skills

inference-aa-workload
データサイエンティスト

Reproduce the Artificial Analysis (AA) language-model performance workload shapes against an OpenAI-compatible chat endpoint using NVIDIA AIPerf. Drives the three AA text shapes (1k input / >=1k answer, 10k / >=1.5k, 100k / >=2k) with temperature 0, top_p 1,…

2026年8月17日
inference-dcgm-correlate
データサイエンティスト

Correlate DCGM Prometheus byte-traffic counters with an inference-perf-bench sweep window to compute byte-grounded workload-level Speed-of-Light. The third tier of the SoL rigor hierarchy (after zymtrace sample-share and ncu per-kernel arithmetic intensity).…

2026年8月17日
inference-decode-step-budget
ソフトウェア品質保証アナリスト・テスター

Measure the low-concurrency (c=1..c=8) decode hot-path of a live vLLM pod FAST and CORRECTLY: where each token's time actually goes (GPU-busy vs host-idle gap vs comm), whether the workload is kernel-bound / host-bound / comm-bound, and what the addressable…

2026年8月17日
inference-graph-diff
ソフトウェア品質保証アナリスト・テスター

Diff the compiled FX / Inductor graphs across two vLLM versions or two helm configs to see exactly which fused kernels / passes / partitions changed. Uses `torch._dynamo.explain` + `torch.compile`'s graph-dump hooks. Useful when you anticipate a…

2026年8月17日
inference-workload-profile
データサイエンティスト

Profile live inference traffic into a token/shape distribution artifact that drives profile-matched speculative-decoding draft training -- an analog of Fireworks FireOptimizer's "profile-driven customization" (the documented source of its higher draft…

2026年8月17日
mirage-graph-coverage
ソフトウェア品質保証アナリスト・テスター

Read-only coverage auditor for the mirage / MPK persistent-megakernel's GENERATED task graph: cross-references DECLARED tensors (all_tensors[...] in kernel_N.cu) against tensors CONSUMED by tasks (inputs/outputs base_ptr in task_graph_N.json) and flags any…

2026年8月17日
analyze-zymtrace-workload
データサイエンティスト

Investigate a GPU or CPU workload through the zymtrace MCP. The MCP does most of the analysis. This skill enforces the cross-view discipline -- always pull the matching opposite-side flamegraph (CPU for GPU workloads, GPU for CPU workloads) with the same…

2026年8月16日
evidence-bundle-init
データサイエンティスト

Scaffold a new evidence bundle directory ready for reproducibility-grade evidence capture: SOURCE.md (operator + cluster + git SHA + UTC timestamp), summary.md (verdict skeleton), commands/ (for the four-file .cmd/.stdout/.stderr/.exit tuple capture per shell…

2026年8月16日
収集済み skill 33 件中 8 件を表示しています。
inference-spec-decode-service
ソフトウェア開発者

Closed-loop speculative-decoding-as-a-service orchestrator -- a self-hosted analog of Fireworks FireOptimizer's adaptive speculative execution. Composes the per-phase skills into one profile-matched loop: profile live traffic (inference-workload-profile) ->…

2026年6月18日
analyze-zymtrace-workload
ソフトウェア開発者

Investigate a GPU or CPU workload through the zymtrace MCP. The MCP does most of the analysis. This skill enforces the cross-view discipline -- always pull the matching opposite-side flamegraph (CPU for GPU workloads, GPU for CPU workloads) with the same…

2026年6月14日
evidence-bundle-init
ソフトウェア開発者

Scaffold a new evidence bundle directory ready for reproducibility-grade evidence capture: SOURCE.md (operator + cluster + git SHA + UTC timestamp), summary.md (verdict skeleton), commands/ (for the four-file .cmd/.stdout/.stderr/.exit tuple capture per shell…

2026年6月14日
inference-aa-workload
ソフトウェア開発者

Reproduce the Artificial Analysis (AA) language-model performance workload shapes against an OpenAI-compatible chat endpoint using NVIDIA AIPerf. Drives the three AA text shapes (1k input / >=1k answer, 10k / >=1.5k, 100k / >=2k) with temperature 0, top_p 1,…

2026年6月14日
inference-capacity-sizing
ソフトウェア開発者

SLA-first GPU capacity sizing for a serving deployment: given a tokens-per-minute (TPM) target AND the interactivity SLA (output tokens/s/user), compute the pods and GPUs needed from a model's measured tok/s/user-vs-concurrency curve. Sizing MUST start from…

2026年6月14日
inference-dcgm-correlate
ソフトウェア開発者

Correlate DCGM Prometheus byte-traffic counters with an inference-perf-bench sweep window to compute byte-grounded workload-level Speed-of-Light. The third tier of the SoL rigor hierarchy (after zymtrace sample-share and ncu per-kernel arithmetic intensity).…

2026年6月14日
inference-decode-step-budget
ソフトウェア開発者

Measure the low-concurrency (c=1..c=8) decode hot-path of a live vLLM pod FAST and CORRECTLY: where each token's time actually goes (GPU-busy vs host-idle gap vs comm), whether the workload is kernel-bound / host-bound / comm-bound, and what the addressable…

2026年6月14日
inference-fleet-leaderboard
ソフトウェア開発者

Render cross-model fleet leaderboards from the local perf-report campaigns in ONE command (`perftunereport fleet_leaderboard`): a latency tier (aa-1k/10k/100k tok/s/user + TTFT + cost at c=1/c=10), a throughput tier (peak tok/s/GPU per (model,quant,TP) with…

2026年6月14日
収集済み skill 32 件中 8 件を表示しています。
5 件中 5 件のリポジトリを表示
すべてのリポジトリを表示しました