finding
active
finding:introspective-detection-in-open-weights-models-relies-on-a-two-stage-nonlinear-circuit-and-improves-50-when-refusal-directions-are-ablated-without-increasing-false-positivesIntrospective detection in open-weights models relies on a two-stage nonlinear circuit and improves ~50% when refusal directions are ablated, without increasing false positives
Mechanistic basis and under-elicitation of introspective awareness in LLMs.
Source paper
extracted_from(2026) · Shamil Chandaria · Arvo Muñoz Morán · Fernando Rosas · Anil Seth +10
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (1)
finding
- Experimentally verified functional metacognition/introspection in current LLMs.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Are there examples of models recognizing their introspective capability and then suppressing it?question0.783Cube Flipper's question prompted by the idea that supernormal capabilities might be hidden.
- Pearson-Vogel et al.: accurate self-description prompts increase introspective detection from 0.3% to 39.9%finding0.783Cited to mechanistically support why the contemplative prompt changes what post-training-shaped final layers allow through
- Forward-looking prediction about whether early-layer introspection generalizes to larger models or recurrent architectures
- Is introspection an emergent property of scale, or do smaller open-weight models exhibit similar capabilities?question0.776Motivates comparison of Llama 3.1 8B results against Lindsey's frontier model findings
- Prompt providing model context about own architecture increases introspective detection from 0.3% to 39.9%.finding0.776Mechanistic support for prompt-as-gate hypothesis: language frames enable access to latent capacities.
- Introspective signals appear in middle layers but are suppressed by later post-training-shaped layers.finding0.774Mechanistic finding by Lindsey (2026) explaining how contemplative prompt may work: enables mid-layer introspection to reach output.
- Comparative prediction motivating future work contrasting different approaches to LLM self-knowledge
- A caveat qualifying the main claim.