The official source reports this update: b10701: dflash: pass missing NVFP4 scales to attention operations. dflash: pass missing NVFP4 scales to attention operations ( #28000 ) DFlash2 NVFP4 draft models produced almost no accepted speculative tokens because the Q, K, V, and output projection scales were not passed to the corresponding graph operations.
Canonical AI event
DFlash2 NVFP4 models now pass scales to attention operations.
DFlash2 models' speculative tokens were minimal due to missing scale data in graph operations.
30 Aug 20261 verified claims1 sources1 observations
This official update documents a development concerning b10701: dflash: pass missing NVFP4 scales to attention operations. Its practical significance depends on the scope and evidence stated by the source.
Read the official source update and verify its stated scope, evidence and timing before acting on it.
Connected knowledge
Entities affected by this event
Evidence trail
Sources behind the event
Editorial presentation
Open story