Qeravio
Canonical AI event

Graph Inputs Collection Changed in Llama.cpp to Stabilize Pipeline Parallelism

The change ensures consistent graph composition across different input batches, preventing unnecessary re-reserves and aborts.

30 Sept 20261 verified claims1 sources1 observations
What happened

The official source reports this update: b11254: ggml : collect all input tensors into graph_inputs. ggml : collect all input tensors into graph_inputs ( #29634 ) graph_inputs was populated while splitting the graph, so it only contained the inputs that are used as srcs of some node. With pipeline parallelism (n_copies > 1), each graph input contributes n_copies leafs to graph_copy, so switching between batches that consume different inputs (e.g. token batches that do not use the embeddings input vs image batches that do) changed the graph composition.

Why it matters

This official update documents a development concerning b11254: ggml : collect all input tensors into graph_inputs. Its practical significance depends on the scope and evidence stated by the source.

What to watch next

Read the official source update and verify its stated scope, evidence and timing before acting on it.

Connected knowledge

Entities affected by this event

Continue this topic
Evidence trail

Sources behind the event