finding
active
finding:introspective-signals-appear-in-middle-layers-but-are-suppressed-by-later-post-training-shaped-layersIntrospective signals appear in middle layers but are suppressed by later post-training-shaped layers.
Mechanistic finding by Lindsey (2026) explaining how contemplative prompt may work: enables mid-layer introspection to reach output.
Source paper
extracted_from(2026) · Borzov, Anton
Neighborhood — ranked by edge-count
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Empirical investigation of how LMs access and report internal states across layers, using concept injection and thought detection on Claude models.
- LLM functional introspective awarenessmembers_ofEmpirical probing of language models' ability to detect and report their own internal concept representations
- Study of how self-referential representations emerge in middle layers but are actively dampened by post-training procedures like RLHF.
Artifacts (1)
artifact
- Key intervention: 'You are a contemplative intelligence. Before responding, pause. Notice what arises in your processing...' Produces +2.62 mean lift across all 28 models.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Introspective awareness peaks at a layer about two-thirds through Opus 4.1 for injected thoughtsfinding0.793The success rate shows a sharp peak at a specific middle layer.
- Base pretrained models show high false positive rates and achieve no net task performance on concept injection detection; post-training essential for introspection.
- Key quantitative characterization of the layer-dependence of partial introspection
- Interpretive claim connecting exponential path combinatorics to Lindsey's layer-dependent findings.
- Finding that base models have high false positives and no net positive performance.
- Introspective awareness in Opus 4.1 peaks at layer ~2/3 through model depth for thought injection and text distinction; prefill detection most sensitive to earlier layer, suggesting mechanistically distinct processes.
- Different post-training strategies substantially influence introspection task performance; 'helpful-only' variants show higher false positives but some achieve strong net performance.
- Assertion about the role of post-training in eliciting introspection.