Qeravio
Canonical AI event

Llama server reduces image processing time by up to 7.61x with scheduler optimization.

The update reduces processing time for multi-image and video inputs on Gemma models by eliminating unnecessary scheduler re-reserves.

29 Sept 20261 verified claims1 sources1 observations
What happened

The official source reports this update: b11227: context : do not re-reserve the scheduler when toggling causal_attn. context : do not re-reserve the scheduler when toggling causal_attn ( #28751 ) context : do not re-reserve the scheduler when toggling causal_attn llama_context::set_causal_attn() marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive sched_reserve() passes per image. This is especially slow for multi-image or video inputs.

Why it matters

This official update documents a development concerning b11227: context : do not re-reserve the scheduler when toggling causal_attn. Its practical significance depends on the scope and evidence stated by the source.

What to watch next

Read the official source update and verify its stated scope, evidence and timing before acting on it.

Connected knowledge

Entities affected by this event

Continue this topic
Evidence trail

Sources behind the event