thinker:openalex-A5127823961Noah Goodman
Authored papers (2)
Manifold steering — intervening on model activations along paths constrained to lie on a learned activation manifold M_h rather than along Euclidean linear directions — produces behavioral trajectories that track the corresponding behavior manifold M_y, while linear (Euclidean) steering cuts through off-manifold regions and generates unnatural outputs. The paper fits M_h to internal representations and M_y to output probability distributions, then tests their bidirectional correspondence via controlled interventions across language models and a video world model. In language models, tasks with cyclic geometries, sequential geometries, and complex graph geometries (in-context learning) all show that manifold-constrained interventions keep behavioral outputs on M_y, whereas linear steering deviates measurably. In a video world model, interventions shaped by physical-dynamics geometry similarly respect M_y. Crucially, the relationship is bidirectional: optimizing interventions in activation space to produce paths along M_y recovers activation trajectories that trace the curvature of M_h. This implies that the core problem of mechanistic steering should be recast not as finding the right direction in a flat Euclidean activation space, but as identifying the correct geometric structure — because representational geometry is not incidental to model behavior but is causally constitutive of it.
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior2026merged
Manifold steering — intervening along activation-space paths defined by the intrinsic geometry of a fitted activation manifold M_h rather than along Euclidean linear directions — produces behavioral trajectories that remain on the behavior manifold M_y, whereas standard linear steering cuts through off-manifold regions and generates unnatural outputs. Across language model reasoning tasks with cyclic, sequential, and graph geometries, and in a video world model whose task geometry corresponds to physical dynamics, the bidirectional relationship M_h ↔ M_y holds: steering that respects M_h tracks M_y, and optimizing interventions in activation space to follow M_y recovers activation trajectories that trace the curvature of M_h. The method fits two manifolds — one to intermediate representations, one to output probability distributions — and then uses geodesic-like paths on M_h as the intervention substrate rather than a single steering vector. This work argues that neural representation geometry is not incidental but is the causal structure organizing model behavior, which implies that the core problem of model steering should be reframed from finding the right direction to finding the right geometry.
More papers — OpenAlex / S2
Affiliations (1)
- Stanford University(institute)
Co-authors (12)
- Daniel Wurgaft19 shared
- Atticus Geiger18 shared
- Can Rager18 shared
- Ekdeep Singh Lubana18 shared
- Jack Merullo18 shared
- Matthew Kowal18 shared
- Raphaël Sarfati18 shared
- Sheridan Feucht18 shared
- Tal Haklay18 shared
- Thomas Fel18 shared
- Thomas McGrath18 shared
- Usha Bhalla18 shared
Recent mentions (3)
- papers-typedpaper.md
- papers-typedwurgaft-2026-manifold.md
- papers-typed
wurgaft-goodfire-manifold-steering-2026.md