一键导入
mlx-inference-optimizer
Optimize Python MLX inference and generation loops with warmup, batching, cache handling, synchronization, quantization, and memory checks.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Optimize Python MLX inference and generation loops with warmup, batching, cache handling, synchronization, quantization, and memory checks.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Guide MLX Metal profiling, mx.fast escalation, custom Metal kernels, and C++ extensions when profiling proves kernel-level bottlenecks.
Route Python-first MLX optimization work on Apple Silicon to focused audit, training, inference, Metal, or bridge workflows.
Audit Python MLX repos for lazy-eval, synchronization, compile, dtype, memory, progress, and benchmark issues.
Advise on Python-first MLX integration with Swift, C, C++, and non-native language boundaries.
Optimize Python MLX training loops with value_and_grad, accumulation, checkpointing, dtype, validation cadence, memory telemetry, and progress reporting.
| name | mlx-inference-optimizer |
| description | Optimize Python MLX inference and generation loops with warmup, batching, cache handling, synchronization, quantization, and memory checks. |
Use this skill for MLX inference, generation, serving loops, batch scoring, streaming output, or latency/throughput questions.
Before any Python execution, use the target repo's .venv. Never install
Python packages globally.
../../references/inference-patterns.md../../references/eval-and-synchronization.md../../references/memory-and-dtypes.md