| name | onnx-export-quantization |
| description | Use this skill when exporting ONNX models with mobius and quantizing them with Olive for deployment. Covers the mobius CLI, EP options, INT4 quantization (Q4_K_M and NF4), HuggingFace upload structure, GPU-accelerated quantization, common issues, and testing quantized models.
|
Skill: ONNX Export and Quantization
When to use
Use this skill when:
- Exporting a model from HuggingFace to ONNX format using
mobius build
- Quantizing an ONNX model to INT4 (Q4_K_M or NF4) with Olive
- Uploading ONNX models to HuggingFace Hub in the standard directory layout
- Debugging export or quantization failures
- Choosing between execution provider (EP) variants
Exporting models with mobius build
Basic command
mobius build \
--model <hf-model-id> \
--dtype <f16|bf16> \
--ep <default|cuda|onnx-standard> \
--runtime ort-genai \
--external-data safetensors \
--max-shard-size 5GB \
<output-directory>/
Flag reference
| Flag | Description |
|---|
--model <id> | HuggingFace model ID (e.g. google/gemma-4-27b-it) |
--dtype <f16|bf16> | Model precision — f16 (float16) or bf16 (bfloat16) |
--optimize [RULES] | Apply mobius rewrite rules after building (e.g. group_query_attention, packed_attention, skip_norm). Use without value for all rules, or specify comma-separated names. Not needed for basic exports. |
--ep <variant> | Execution provider variant (see below) |
--runtime ort-genai | Generate genai_config.json and copy tokenizer files for ORT GenAI runtime |
--external-data safetensors | Store weights externally in safetensors format |
--max-shard-size 5GB | Split external data into shards ≤ 5GB |
Execution provider (EP) variants
Build separate ONNX models per EP because each applies different graph
rewrites and fused ops:
| EP | Flag | When to use |
|---|
default | --ep default | Portable ONNX — no vendor-specific fusions. Compatible with all execution providers and runtimes. This is the default if --ep is omitted. |
cuda | --ep cuda | NVIDIA GPU inference. Emits com.microsoft fused ops (GroupQueryAttention, MoE, etc.) for maximum CUDA performance. |
onnx-standard | --ep onnx-standard | Strict ONNX-only — inlines all custom-domain functions into standard ONNX ops. Use when targeting runtimes that don't support com.microsoft ops. |
Other EPs are available (cpu, dml, webgpu, trt-rtx). Run
mobius list eps to see all options.
Typical export matrix: Build each dtype × EP combination:
for dtype in f16 bf16; do
for ep in default cuda onnx-standard; do
mobius build --model google/gemma-4-12b-it \
--dtype $dtype --ep $ep \
--runtime ort-genai \
--external-data safetensors --max-shard-size 5GB \
output/${dtype}/${ep}/
done
done
Multi-model outputs
For multimodal models (VLMs, audio-language), mobius build produces
multiple sub-models:
output/
├── decoder/ # Text decoder
│ ├── model.onnx
│ └── model.onnx.data.safetensors
├── embedding/ # Embedding model
│ ├── model.onnx
│ └── model.onnx.data.safetensors
├── vision_encoder/ # Vision encoder (VLMs)
│ ├── model.onnx
│ └── model.onnx.data.safetensors
├── audio_encoder/ # Audio encoder (ALMs)
│ ├── model.onnx
│ └── model.onnx.data.safetensors
└── genai_config.json
Quantization with Olive
Installation
Olive with ONNX quantization support (install from PR if needed for
latest features):
pip install olive-ai
pip install git+https://github.com/microsoft/Olive.git@refs/pull/2406/head
For GPU-accelerated quantization (highly recommended for large models):
pip install cupy-cuda12x
Q4_K_M quantization (k-quant)
K-quant quantization uses mixed block sizes with importance-based bit
allocation. Q4_K_M is a good balance of quality and size.
The repo uses Olive's config-driven olive.run() pattern (see
examples/olive/ for working examples). A typical Olive config for
k-quant quantization:
{
"input_model": { "type": "OnnxModel", "model_path": "decoder/model.onnx" },
"passes": {
"kquant": {
"type": "OnnxKQuantQuantization",
"bits": 4,
"block_size": 32
}
},
"output_dir": "output/Q4_K_M/default/decoder"
}
olive run --config kquant_config.json
NF4 quantization (4-bit NormalFloat)
NF4 uses a normal-distribution-optimized 4-bit format. Fast native C++
implementation — no GPU needed.
{
"input_model": { "type": "OnnxModel", "model_path": "decoder/model.onnx" },
"passes": {
"nf4": {
"type": "OnnxBnb4Quantization",
"precision": "nf4"
}
},
"output_dir": "output/NF4/default/decoder"
}
See examples/olive/ministral-3-3b-vlm/ for a complete working
example that combines mobius export with Olive quantization.
GPU acceleration with cupy
Installing cupy-cuda12x gives a 19–51x speedup for k-quant
quantization:
| Method | CPU time per matrix | GPU time per matrix | Speedup |
|---|
| K-quant (Q4_K_M) | 3–27s | 0.17–0.52s | 19–51x |
| NF4 | 42ms for 67M params | N/A (C++ native) | Already fast |
pip install cupy-cuda12x
Quantizing multi-model exports
Quantize each sub-model independently. Typically only the decoder is
quantized (it has the most parameters). Copy all other files needed
for a complete ORT GenAI package:
olive run --config kquant_decoder.json
cp -r output/f16/default/embedding/ output/Q4_K_M/default/embedding/
cp -r output/f16/default/vision_encoder/ output/Q4_K_M/default/vision_encoder/
cp output/f16/default/genai_config.json output/Q4_K_M/default/
cp output/f16/default/tokenizer* output/Q4_K_M/default/
cp output/f16/default/image_processor.json output/Q4_K_M/default/ 2>/dev/null
cp output/f16/default/audio_processor.json output/Q4_K_M/default/ 2>/dev/null
Without the tokenizer and processor config files, ORT GenAI will fail
to load the model.
HuggingFace upload structure
Standard directory layout
<org>/<model>-onnx/
├── f16/
│ ├── default/ # Portable ONNX (no vendor fusions)
│ ├── cuda/ # CUDA EP (fused ops)
│ └── onnx-standard/ # Strict ONNX-only (inlined functions)
├── bf16/
│ ├── default/
│ ├── cuda/
│ └── onnx-standard/
├── Q4_K_M/
│ └── default/ # Quantized models typically CPU-only
└── NF4/
└── default/
Each EP directory contains the full model structure (decoder/,
embedding/, vision_encoder/, audio_encoder/ as applicable) plus
genai_config.json.
Upload with huggingface_hub
from huggingface_hub import HfApi
api = HfApi()
api.upload_folder(
folder_path="output/f16/default",
path_in_repo="f16/default",
repo_id="org/model-onnx",
repo_type="model",
)
Verify uploads
After uploading, verify all shards are present. Incomplete uploads are
a common issue with large models:
from huggingface_hub import HfApi
api = HfApi()
files = api.list_repo_files("org/model-onnx")
for variant in ["f16/default", "f16/cuda", "bf16/default"]:
shards = [f for f in files if f.startswith(variant) and f.endswith(".safetensors")]
print(f"{variant}: {len(shards)} shards")
Common issues and fixes
1. MoE expert weight mapping
Models with Mixture-of-Experts (e.g. Gemma4 26b-a4b) may need expert
weight remapping in preprocess_weights(). HuggingFace stores experts
as 3D tensors (experts.gate_up_proj [E, 2*inter, H]) that must be
mapped to the fused MoE op's parameter names (fc1_experts_weights,
fc2_experts_weights).
Symptom: Weight loading errors or incorrect MoE outputs.
Fix: Check the model's preprocess_weights() maps HF expert weight
names to the ONNX parameter names. See the moe-models skill for
the pattern.
2. Hybrid attention v_proj shape mismatches
Models with hybrid attention (e.g. Gemma4 31b with different head_dim
for local vs global attention layers) may have shape mismatches in
value projections.
Symptom: Shape errors during weight loading or forward pass.
Fix: Ensure v_proj dimensions account for per-layer head
configurations. Check num_global_key_value_heads vs
num_key_value_heads in the config.
3. CUDA GQA head_dim limitations
Older versions of ORT had a limitation where head_dim > 256 would fail
with the CUDA GroupQueryAttention kernel.
Symptom: CUDA runtime error during inference with large head
dimensions.
Status: This limitation has been removed in recent ORT versions.
If using an older ORT build, fall back to --ep default or
--ep onnx-standard.
4. Incomplete uploads
Large models with many shards can have incomplete uploads to HuggingFace
Hub, especially on unstable connections.
Symptom: Model fails to load with file-not-found errors for
specific shard files.
Fix: Verify all shards are present after upload (see the verify
script above). Re-upload missing shards with api.upload_file().
5. BF16 type mismatches
Some components may produce FP32 outputs when the model is built in
BF16, causing type mismatch errors in ORT.
Symptom: Type Error: Type parameter (T) bound to different types (tensor(bfloat16) and tensor(float)).
Fix: Check for constants, initializers, or norm layers that stay
FP32 when the model is BF16. Add op.CastLike(result, input) to
ensure dtype consistency. See the reusable-components skill's
section on precision behaviour.
Testing quantized models
L4: Golden data generation
Generate reference outputs from the full-precision HuggingFace model
using the golden data generation script:
python scripts/generate_golden.py
python scripts/generate_golden.py --task-type causal-lm
python scripts/generate_golden.py --case testdata/cases/causal-lm/gpt2.yaml
python scripts/generate_golden.py --device cuda
Golden reference files are stored in testdata/golden/ as JSON. Use
compare_golden() from mobius._testing.parity to compare model
outputs against the reference:
from mobius._testing.parity import compare_golden
compare_golden(
model_output=output_logits,
golden_path="testdata/golden/causal-lm/my_model.json",
)
L5: End-to-end smoke test
Run inference with the quantized model through ORT GenAI:
import onnxruntime_genai as og
model = og.Model("output/Q4_K_M/default/")
tokenizer = og.Tokenizer(model)
params = og.GeneratorParams(model)
params.set_search_options(max_length=50, do_sample=False)
params.input_ids = tokenizer.encode("Hello, world!")
output_ids = model.generate(params)
print(tokenizer.decode(output_ids[0]))
Numerical parity verification
Quantized models will have some numerical divergence from the
full-precision model. Expected tolerances:
| Quantization | Typical divergence | Notes |
|---|
| Q4_K_M | Moderate | Top-1 token agreement ~95%+ for coherent text |
| NF4 | Moderate | Similar to Q4_K_M |
| F16 (no quant) | Minimal | Should match BF16 closely |
Verify that generated text is coherent and semantically correct rather
than requiring exact numerical matches.
Inference speed: Q4_K_M vs NF4
For the benchmark table, use the canonical reference in
.agents/skills/profiling-onnx-models/SKILL.md ("Quantization benchmark
reference (Gemma4 E2B-IT)").
Q4_K_M is recommended over NF4 for both speed and quality:
- Faster on both CPU and CUDA than F16 in the referenced measurement
- NF4 is slower than F16 in the same measurement
- Quality is comparable between Q4_K_M and NF4 in spot checks
- Q4_K_M uses less memory than F16 (~4x compression)
Cross-references
- Adding models:
.agents/skills/adding-a-new-model/SKILL.md
- MoE weights:
.agents/skills/moe-models/SKILL.md
- Component precision:
.agents/skills/reusable-components/SKILL.md
- ORT GenAI config:
.agents/skills/ort-genai-config/SKILL.md
- Quality checklist:
.agents/skills/quality-checklist/SKILL.md