Qeravio
Canonical AI event

New Engine Features Include Continuous Batching, Faster Decoding, and Qwen2.5-VL Support

Version 0.16.0 introduces support for AMDGPU and WebGPU. It also adds Granite models and improves handling of multi-image vision inputs.

23 Sept 20261 verified claims1 sources1 observations
What happened

The official source reports this update: v0.16.0. This is a big release! Version 0.16.0 brings major upgrades for high-throughput serving, faster decoding, and the latest generation of Qwen models. Highlights A much more capable Engine - Continuous batching now supports chunked prefill, paged attention, multi-turn conversations, mixed workloads, and smarter token scheduling. This improves throughput while making better use of GPU memory. Faster generation with speculative decoding - Added support for draft models, n-gram speculation, MTP self-speculation, and DFlash2 and DSpark block drafting.

Why it matters

This official update documents a development concerning v0.16.0. Its practical significance depends on the scope and evidence stated by the source.

What to watch next

Read the official source update and verify its stated scope, evidence and timing before acting on it.

Connected knowledge

Entities affected by this event

Evidence trail

Sources behind the event