finding
active
finding:models-more-effective-at-recognizing-abstract-nouns-than-other-concept-typesModels more effective at recognizing abstract nouns than other concept types
Opus 4.1 demonstrates highest introspective awareness on abstract nouns (justice, peace, betrayal) with nonzero awareness across all concept categories tested.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Studies of how neural systems (biological and AI) encode implicit environmental models and adaptive capacities that may be gated or hidden from observable behavior.
- LLM functional introspective awarenessmembers_ofEmpirical probing of language models' ability to detect and report their own internal concept representations
Findings (1)
finding
- Self-report of Injected ThoughtssupportsModels can detect and identify injected concept vectors ~20% of the time at optimal layer/strength in Opus 4.1, with immediacy suggesting internal rather than output-inferred detection.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Opus 4.1 is most effective at recognizing injected abstract concepts (e.g., justice, peace) but detects other categories too.
- Selective pressure toward convergence via task generality
- Author's interpretation of the VTAB alignment results echoing Tolstoy
- Articulates why a one-layer transformer with MLP is the appropriate starting target for mechanistic interpretability
- Critique of using formal specifications alone for concept definition.
- Motivation for using sparsity-based dictionary learning on language models
- Bigger models are more likely to converge to a shared representation than smaller modelshypothesis0.780Selective pressure toward convergence via model capacity
- Interpretation of weaker PCA separation and lower ASR in smaller models