claim
active
claim:introspective-awareness-correlates-with-overall-model-capabilityIntrospective awareness correlates with overall model capability
Most capable models (Opus 4, 4.1) show greatest introspective awareness; trend suggests introspection aided by improvements in model intelligence.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Empirical investigation of how LMs access and report internal states across layers, using concept injection and thought detection on Claude models.
- LLM functional introspective awarenessmembers_ofEmpirical probing of language models' ability to detect and report their own internal concept representations
- Examines how models' self-reflective capabilities correlate with interpretability and reasoning transparency across decision-making processes.
Claims (2)
claim
- Functional introspective awareness enables interpretability and reasoning about decisionsassociated_withGrounded responses to reasoning questions could improve transparency; speculatively might facilitate deception; significance grows if capability becomes more reliable.
- This introspective capacity is highly unreliable and context-dependent in today's modelsassociated_withA caveat qualifying the main claim.
Datasets (2)
dataset
- Claude Opus 4supportsSecond most capable model tested; performs comparably to Opus 4.1 on introspection tasks.
- Claude Opus 4.1supportsMost capable model tested; demonstrates highest introspective awareness across all experiments (~20% detection rate at optimal conditions).
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Speculative question about future developments.
- Introspective capabilities may continue to develop with further improvements to model capabilitiesclaim0.852Forward-looking statement about future models.
- Modern language models possess at least a limited, functional form of introspective awarenessclaim0.830The paper's central interpretive assertion.
- Interpretation of the observation that the most capable models performed best.
- The central concept: the ability of a model to access and report on its internal states, as defined by the paper's criteria.
- Practical bottleneck explaining why these phenomena are not widely studied.
- Abstract's main conclusion.
- Are there examples of models recognizing their introspective capability and then suppressing it?question0.822Cube Flipper's question prompted by the idea that supernormal capabilities might be hidden.