finding
active
finding:post-training-is-key-to-eliciting-introspective-awarenessPost-training is key to eliciting introspective awareness
Base pretrained models show high false positive rates and achieve no net task performance on concept injection detection; post-training essential for introspection.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Claims (1)
claim
- Different post-training strategies substantially influence introspection task performance; 'helpful-only' variants show higher false positives but some achieve strong net performance.
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Empirical investigation of how LMs access and report internal states across layers, using concept injection and thought detection on Claude models.
- LLM functional introspective awarenessmembers_ofEmpirical probing of language models' ability to detect and report their own internal concept representations
- How instruction tuning and RLHF elicit latent introspective capabilities in language models beyond base pretraining.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Finding that base models have high false positives and no net positive performance.
- Assertion about the role of post-training in eliciting introspection.
- Secondary question; paper demonstrates introspection but explicitly avoids pinning down specific mechanistic explanation, noting mechanisms could be shallow and specialized.
- Authors' interpretive endorsement of PSM view, backed by transfer experiments
- Central interpretive claim and motivation for future work
- The phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.
- Introspective signals appear in middle layers but are suppressed by later post-training-shaped layers.finding0.789Mechanistic finding by Lindsey (2026) explaining how contemplative prompt may work: enables mid-layer introspection to reach output.
- Empirical generalization from contemplative neuroscience supporting the viability of Contemplative AI approach