一键导入
gpu-bench
Benchmark GPU-accelerated operations (vector distance, batch computation) against CPU baselines. Requires CUDA toolkit. Args: vector|distance|batch|all.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Benchmark GPU-accelerated operations (vector distance, batch computation) against CPU baselines. Requires CUDA toolkit. Args: vector|distance|batch|all.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
ADD (AI-Driven Development) — a minimal, state-tracked workflow for building software where the AI writes the code and the human owns direction and verification. Drives every feature through one lean TASK.md: Specify → Scenarios → Contract → Tests → Build → Verify → Observe, with red/green TDD built in. Use this skill whenever working in a repo that has a `.add/` directory, when the user says "add", "start a task", "next phase", "specify this feature", "ADD method", or "AI-driven development", or when scaffolding a new feature and you want spec/tests-first discipline instead of vague-prompt coding. Also use it to resume work across sessions (it reads `.add/state.json` so you never re-read the whole repo).
Scaffold a new Redis command implementation with dispatch entry, handler, ACL, tests, and consistency test entry. Args: COMMAND_NAME [category].
Scan hot-path code for allocation violations, lock misuse, and performance anti-patterns. Zero-tolerance audit.
Scaffold a new CUDA kernel with Rust integration via cudarc. Args: kernel_name [f32|f16]. Creates .cu kernel, Rust wrapper, CPU fallback, and benchmark.
Verify code compiles and tests pass under both runtime-tokio and runtime-monoio feature sets. Run before committing runtime-adjacent changes.
Verify moon command behavior matches Redis exactly. Args: command name(s) or --category <name>. Uses redis-cli against both servers.
| name | gpu-bench |
| description | Benchmark GPU-accelerated operations (vector distance, batch computation) against CPU baselines. Requires CUDA toolkit. Args: vector|distance|batch|all. |
GPU vs CPU benchmark suite for moon acceleration paths.
/gpu-bench vector — vector distance computation (L2, cosine, dot)/gpu-bench batch — batch operation throughput/gpu-bench all — full GPU benchmark suiteVerify CUDA availability:
nvcc --version
nvidia-smi
Build with GPU feature:
RUSTFLAGS="-C target-cpu=native" cargo build --release --features gpu-cuda
| Operation | Dimensions | Batch Size | Measure |
|---|---|---|---|
| L2 distance | 128/256/512/1024 | 1/100/1K/10K | ops/sec, latency |
| Cosine similarity | 128/256/512/1024 | 1/100/1K/10K | ops/sec, latency |
| Dot product | 128/256/512/1024 | 1/100/1K/10K | ops/sec, latency |
| HNSW search | 128d, 100K vectors | 1/10/100 queries | recall@10, latency |
| Batch MGET | — | 100/1K/10K keys | throughput |
Run CPU baseline (SIMD path):
RUSTFLAGS="-C target-cpu=native" cargo bench --bench gpu_distance -- --baseline cpu
Run GPU path:
cargo bench --bench gpu_distance --features gpu-cuda
Generate comparison table with:
Recommend optimal batch sizes for production use.