Qeravio
Canonical AI event

ModelOpt 0.47.0 introduces FP8 Vision Encoder recipes for qwen3_vl and qwen3_5 models.

ModelOpt 0.47.0 introduces quantization with Autotune for ONNX models, benchmarking placements in requested runtime precision.

23 Sept 20261 verified claims1 sources1 observations
What happened

The official source reports this update: ModelOpt 0.47.0 Release. New Features Quantization ONNX quantization with Autotune now benchmarks placements in the requested runtime precision and retains calibrated INT8/FP8 Q/DQ only when it meets the configured TensorRT speedup threshold (1.02x by default); otherwise it saves the high-precision no-Q/DQ model. Add a Muse Glimmer AutoQuantize recipe that searches language-model MLP projections, self-attention projections, and lm_head over W4A16 NVFP4 Four-Over-Six, FP8, and BF16 fallback at 5.5 effective bits while leaving the vision tower unquantized.

Why it matters

This official update documents a development concerning ModelOpt 0.47.0 Release. Its practical significance depends on the scope and evidence stated by the source.

What to watch next

Read the official source update and verify its stated scope, evidence and timing before acting on it.

Connected knowledge

Entities affected by this event

Evidence trail

Sources behind the event