concept
active
concept:activation-spaceActivation space
Representation space on which linear probes operate to attribute harmful behaviors to training data.
Neighborhood — ranked by edge-count
Papers (2)
paper
Claims (2)
claim
- General principle derived from the Mountain Car experiment: curved manifold-following yields coherent manipulation, linear shortcuts fail.
- Conceptual framing: integrates mechanistic interpretability tools with alignment-focused data curation.
Methods (1)
method
- Probe-Based Data Attributionassociated_withLinear classifier approach applied to model activations to identify which training datapoints caused undesired behaviors in post-training.
Concepts (4)
concept
- Direction (activation space)related_toA linear combination of neurons in a layer; the general form of a neural network feature including both individual neurons and other combinations
- geometry of activation spacerelated_toRich geometric structure carried by neural representations.
- representation manifoldassociated_withOne-dimensional curved surface in internal activation space; the paper demonstrates alignment with behavior manifold.
- activation manifold M_hassociated_withManifold fitted to representations in activation space.
Findings (1)
finding
- Empirical demonstration that a semantically meaningful variable is encoded as a curved manifold, and that respecting its geometry is critical for effective intervention.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Spaces of model activations from which sparse features are retrieved.
- Internal representations of the model on which probes operate; the method uses activations to rank datapoints.
- A geometric space of all output token probability distributions, equipped with Hellinger distance, used to visualize model behavior.
- Intervention method that adds a learned direction vector to residual stream activations to steer model behavior
- The ensemble of all possible configurations of a building, including incomplete states and paths between them.
- The low-dimensional geometric structure discovered in neural activation space; contrasted with linear/Euclidean geometry.
- Proposed mechanistic measure for testing whether persona-model collapse reduces internal differentiation between character representations
- Model-independent feature comparison based on correlating activation vectors across a fixed diverse dataset