concept
active
concept:introspectionIntrospection
The ability of a model to observe its own past internal states or computations; claimed to be architecturally permitted by transformers.
Neighborhood — ranked by edge-count
Papers (2)
paper
Communities (1)
community
- LLM Introspectionmembers_of
Methods (1)
method
- KV cachingimplementsCaching of key-value pairs to avoid recomputation; also provides a mechanism for introspection of earlier computations.
Concepts (7)
concept
- AI Introspectionrelated_toKey gap identified in the literature; systematic self-examination processes for machine consciousness development.
- Introspective awarenessrelated_toThe central concept: the ability of a model to access and report on its internal states, as defined by the paper's criteria.
- model introspectionrelated_toThe capacity of a model to self-report on its internal emotional state when its SAE features are steered, used here as a measurement tool
- Latent Introspectionrelated_toPearson-Vogel et al.'s finding that models can detect prior concept injections; introspective signals exist in middle layers suppressed by post-training
- partial introspectionrelated_toThe authors' characterization of genuine but limited introspective capability found only in early-layer injection regimes
- World Modelsassociated_withTheme issue context: relates to internal models of environment, central to consciousness and cognition across substrates.
- SCI loop (Self-referential Cognitive Inspection)associated_with
Artifacts (1)
artifact
- Original thread by janus explaining transformer information highways and introspection capabilities, posted on X.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Stage 3 of character training: SFT on synthetic introspective data generated by post-distillation checkpoint
- The capacity to detect and report one's own internal states, measured via the five-adjective task and paradox reflection
- Tracking of functional/computational cognitive states, distinguished from phenomenal introspection.
- Direct introspection into phenomenal consciousness; its correlation with functional introspection is an open question.
- Identified gap; methods for enabling machine consciousness development through self-examination.
- Spearman ρ measuring rank-order agreement between logit-based self-report and probe score; the paper's primary monotonic association metric
- The novel framework introduced in the paper: an HMM-based pain-belief signal integrated into the reward function to drive exploration
- Training data generated by the post-distillation model through self-reflection and self-interaction, capturing character nuances beyond the constitution