Qeravio
Canonical AI event

Llama.cpp supports head_dim 72 in HMX flash-attention

The update supports head_dim of 72 for HMX flash-attention, with zero-filled tail lanes for non-multiples of 64.

19 Sept 20261 verified claims1 sources1 observations
What happened

The official source reports this update: b11043: hexagon: HMX flash-attention head_dim padding (support DK=DV=72). hexagon: HMX flash-attention head_dim padding (support DK=DV=72) ( #26539 ) Allow HMX flash-attention to run with head_dim not a multiple of 64 (e.g. SigLIP head_dim=72), by operating on DK/DV rounded up to 64 with zero-filled tail lanes.

Why it matters

This official update documents a development concerning b11043: hexagon: HMX flash-attention head_dim padding (support DK=DV=72). Its practical significance depends on the scope and evidence stated by the source.

What to watch next

Read the official source update and verify its stated scope, evidence and timing before acting on it.

Connected knowledge

Entities affected by this event

Continue this topic
Evidence trail

Sources behind the event