Qeravio
Canonical AI event

Muon's update size remains unchanged under fp16 loss scaling.

Muon's update size remains consistent across loss scales with this change.

30 Sept 20261 verified claims1 sources1 observations
What happened

The official source reports this update: nightly-last-green: Keep Muon's update at full size under fp16 loss scaling (#8655). Under fp16, Muon parameters barely train with ZeRO-1/2/3 unless the optimizer is offloaded. Newton-Schulz returns the same update at any loss scale, but the step then treats it like a scaled gradient: unscale_and_clip_grads divides it by the loss scale, and the global norm takes it as scaled too. So a Muon parameter moves by 1/loss_scale of its update, 1/65536 with the default dynamic loss scaler.

Why it matters

This official update documents a development concerning nightly-last-green: Keep Muon's update at full size under fp16 loss scaling (#8655). Its practical significance depends on the scope and evidence stated by the source.

What to watch next

Read the official source update and verify its stated scope, evidence and timing before acting on it.

Connected knowledge

Entities affected by this event

Continue this topic
Evidence trail

Sources behind the event