The official source reports this update: b11227: context : do not re-reserve the scheduler when toggling causal_attn. context : do not re-reserve the scheduler when toggling causal_attn ( #28751 ) context : do not re-reserve the scheduler when toggling causal_attn llama_context::set_causal_attn() marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive sched_reserve() passes per image. This is especially slow for multi-image or video inputs.
Canonical AI event
Llama server reduces image processing time by up to 7.61x with scheduler optimization.
The update reduces processing time for multi-image and video inputs on Gemma models by eliminating unnecessary scheduler re-reserves.
29 Sept 20261 verified claims1 sources1 observations
This official update documents a development concerning b11227: context : do not re-reserve the scheduler when toggling causal_attn. Its practical significance depends on the scope and evidence stated by the source.
Read the official source update and verify its stated scope, evidence and timing before acting on it.
Connected knowledge
Entities affected by this event
Continue this topic
Related verified updates
Evidence trail
Sources behind the event
Editorial presentation
Open story