concept
active
concept:linear-representation-of-concepts

Linear Representation of Concepts

The established finding that transformer LLMs encode many interpretable concepts as linear directions in activation space

Neighborhood — ranked by edge-count

Concepts (3)

concept

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • The hypothesis that models internalize concepts as approximately linear directions in representation space; used to interpret MDS injection behavior
  • How a neural network encodes a semantic concept internally, argued to be better captured by manifolds than by atomic features.
  • linearityconcept0.808
    The sequential, continuous order of text, often challenged by diagrammatic branching.
  • Linear Decodingmethod0.789
    Correlative technique measuring the type of information encoded in distributed representations via linear predictability.
  • Core contribution: the impasse where lifting linearity in alignment maps makes causal abstraction vacuous, but keeping it may miss non-linearly encoded features
  • The idea that programs can be expressed as logical sentences, enabling direct deductive verification.
  • Recent work identifying cases where LLM features are not one-dimensionally linear, a caveat to the linearity hypothesis.
  • Hypothesis that information may be encoded in arbitrary non-linear subspaces of a neural network