paper
merged
2026
paper:manifold-steering-reveals-the-shared-geometry-of-neural-network-representation-and-behavior

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior

TL;DR

Manifold steering — intervening along activation-space paths defined by the intrinsic geometry of a fitted activation manifold M_h rather than along Euclidean linear directions — produces behavioral trajectories that remain on the behavior manifold M_y, whereas standard linear steering cuts through off-manifold regions and generates unnatural outputs. Across language model reasoning tasks with cyclic, sequential, and graph geometries, and in a video world model whose task geometry corresponds to physical dynamics, the bidirectional relationship M_h ↔ M_y holds: steering that respects M_h tracks M_y, and optimizing interventions in activation space to follow M_y recovers activation trajectories that trace the curvature of M_h. The method fits two manifolds — one to intermediate representations, one to output probability distributions — and then uses geodesic-like paths on M_h as the intervention substrate rather than a single steering vector. This work argues that neural representation geometry is not incidental but is the causal structure organizing model behavior, which implies that the core problem of model steering should be reframed from finding the right direction to finding the right geometry.

What to take away

  1. 1. Manifold steering, which intervenes along paths on a fitted activation manifold M_h rather than linear directions, produces behavioral trajectories that stay on the behavior manifold M_y across all tested tasks and modalities.
  2. 2. Linear (Euclidean) steering consistently passes through off-manifold activation regions and thereby generates outputs that diverge from the model's naturally occurring behavioral distribution.
  3. 3. Optimizing interventions in activation space to trace M_y recovers activation trajectories that follow the curvature of M_h, establishing the M_h ↔ M_y relationship as bidirectional rather than unidirectional.
  4. 4. The paper evaluates manifold steering in language models on reasoning tasks with at least three distinct geometric topologies: cyclic, sequential, and graph-structured in-context learning.
  5. 5. Beyond language models, the bidirectional geometry relationship is demonstrated in a video world model on a task whose underlying structure corresponds to physical dynamics, showing the result generalizes across modalities.
  6. 6. The methodology fits M_h to intermediate neural representations and M_y to output probability distributions independently, then tests their correspondence through intervention — a two-manifold pipeline another researcher could replicate by substituting any differentiable generative model.
  7. 7. The paper raises the open question of whether the shared geometry principle extends to adversarial or distribution-shifted inputs where M_h and M_y might decouple.
  8. 8. The core reframing the paper argues for is that model steering should target geometric structure (the right manifold path) rather than a scalar direction, which has direct implications for interpretability-based control methods.
  9. 9. Tasks with cyclic geometry serve as a controlled test bed because the expected behavioral trajectory is a closed loop, providing a quantitative ground truth against which on-manifold versus off-manifold interventions can be compared.
  10. 10. The submission, arXiv:2605.05115, involves at least 16 authors across institutions, suggesting the empirical scope required to validate the claim across multiple model families and task geometries was substantial.

Peer brief — for seminar discussion

Wurgaft et al. (arXiv:2605.05115, May 2026) ask whether the geometric structure observed in neural activations is causally implicated in behavior or merely correlational. To test this, they introduce manifold steering: they independently fit an activation manifold M_h to a model's intermediate representations and a behavior manifold M_y to its output probability distributions, then compare interventions that follow geodesic-like paths on M_h against conventional linear (Euclidean) steering vectors. The core finding is a bidirectional correspondence — steering that respects M_h produces behavioral trajectories that remain on M_y, while linear steering exits M_h and generates outputs outside M_y; and conversely, optimizing activation interventions to track M_y recovers trajectories that trace the curvature of M_h. This is demonstrated across language model reasoning tasks with three geometric topologies (cyclic, sequential, and graph-structured in-context learning) and in a video world model on a physical-dynamics task, giving the claim at least 4 distinct empirical substrates spanning 2 modalities. The implication is substantive: representation geometry is not epiphenomenal but is the proper object for principled causal intervention. This reframes the steering problem from identifying a scalar direction (the dominant paradigm, exemplified by linear probes and activation addition) to identifying the correct manifold and moving along it. An alternative method the paper could have used — and which would have been a natural comparison — is distributed alignment search (DAS) or interchange intervention methods that also claim to identify causally relevant subspaces but remain linear in their intervention geometry; testing whether DAS-style interventions stay on M_y would sharpen the contrast. The most contestable element is the operationalization of M_h and M_y themselves. Manifold fitting from finite samples is sensitive to dimensionality choices, kernel bandwidth, and the particular manifold-learning algorithm used; the paper's results inherit these degrees of freedom. A critical reader would press on whether the M_h ↔ M_y correspondence is robust to different manifold estimation procedures, or whether it is partially an artifact of fitting both manifolds with related inductive biases (e.g., both using locally linear approximations). If the two manifolds share an estimator family, the bidirectionality could reflect shared estimation error rather than a true geometric coupling in the model. The paper's prediction — that geometry is the right level of description for enabling principled control — is strong enough to generate concrete falsifiable tests: a case where M_h and M_y structurally disagree would challenge the entire framing.

Findings (5)

Claims (6)

Questions (4)

Original abstract (expand)

Neural representations carry rich geometric structure; but does that structure causally shape behavior? To address this question, we intervene along paths through activation space defined by different geometries, and measure the behavioral trajectories they induce. In particular, we test whether interventions that respect the geometry of activation space will yield behaviors close to those the model exhibits naturally. Concretely, we first fit an activation manifold $M_h$ to representations and a behavior manifold $M_y$ to output probability distributions. We then test the link $M_h \leftrightarrow M_y$ via interventions: we find that steering along $M_h$, which we term manifold steering, yields behavioral trajectories that follow $M_y$, while linear steering -- which assumes a Euclidean geometry -- cuts through off-manifold regions and hence produces unnatural outputs. Moreover, optimizing interventions in activation space to produce paths along $M_y$ recovers activation trajectories that trace the curvature of $M_h$. We demonstrate this bidirectional relationship between the geometry of representation and behavior across tasks and modalities. In language models, we use reasoning tasks with cyclic and sequential geometries as well as in-context learning tasks with more complex graph geometries. In a video world model, we use a task with geometry corresponding to physical dynamics. Overall, our work shows that geometry in neural representation is not merely incidental, but is in fact the proper object for enabling principled control via intervention on internals. This recasts the core problem of steering from finding the right direction to finding the right geometry.

Related work— refs + corpus + external arXiv

Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.

Similar preprints — Semantic Scholar