Canonical AI event

ONNX Runtime 1.31.0 introduces stable model-package APIs and enhanced CPU model loading.

ONNX Runtime 1.31.0 includes improved CPU model loading and quantized Mixture of Experts inference.

9 Oct 20261 verified claims1 sources1 observations
What happened

The official source reports this update: ONNX Runtime v1.31.0. ONNX Runtime 1.31.0 stabilizes model-package and EPContext data APIs, improves CPU model loading and quantized MoE inference, expands device-based execution-provider selection, and strengthens model-loading and runtime reliability. These notes cover changes since ONNX Runtime 1.30.0. CUDA and WebGPU kernel updates are summarized here; detailed provider notes are covered by their separate plugin EP releases. Highlights Promoted the model-package API and EPContext data callbacks to stable C and C++ APIs, with EPContext callback support added across language bindings ( #33166 , #32265 ).

Why it matters

This official update documents a development concerning ONNX Runtime v1.31.0. Its practical significance depends on the scope and evidence stated by the source.

What to watch next

Read the official source update and verify its stated scope, evidence and timing before acting on it.

Connected knowledge

Entities affected by this event

Evidence trail

Sources behind the event