concept
active
concept:activation-space

Activation space

Representation space on which linear probes operate to attribute harmful behaviors to training data.

Neighborhood — ranked by edge-count

Claims (2)

claim

Methods (1)

method
  • Linear classifier approach applied to model activations to identify which training datapoints caused undesired behaviors in post-training.

Concepts (4)

concept
  • A linear combination of neurons in a layer; the general form of a neural network feature including both individual neurons and other combinations
  • Rich geometric structure carried by neural representations.
  • One-dimensional curved surface in internal activation space; the paper demonstrates alignment with behavior manifold.
  • Manifold fitted to representations in activation space.

Findings (1)

finding

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Spaces of model activations from which sparse features are retrieved.
  • Activationsconcept0.826
    Internal representations of the model on which probes operate; the method uses activations to rank datapoints.
  • Behavior Spaceconcept0.792
    A geometric space of all output token probability distributions, equipped with Hellinger distance, used to visualize model behavior.
  • Intervention method that adds a learned direction vector to residual stream activations to steer model behavior
  • The ensemble of all possible configurations of a building, including incomplete states and paths between them.
  • The low-dimensional geometric structure discovered in neural activation space; contrasted with linear/Euclidean geometry.
  • Proposed mechanistic measure for testing whether persona-model collapse reduces internal differentiation between character representations
  • Model-independent feature comparison based on correlating activation vectors across a fixed diverse dataset