claim
active
claim:post-training-influences-introspective-capability-expressionPost-training influences introspective capability expression
Different post-training strategies substantially influence introspection task performance; 'helpful-only' variants show higher false positives but some achieve strong net performance.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Findings (1)
finding
- Base pretrained models show high false positive rates and achieve no net task performance on concept injection detection; post-training essential for introspection.
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Empirical investigation of how LMs access and report internal states across layers, using concept injection and thought detection on Claude models.
- LLM functional introspective awarenessmembers_ofEmpirical probing of language models' ability to detect and report their own internal concept representations
- How instruction tuning and RLHF elicit latent introspective capabilities in language models beyond base pretraining.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Assertion about the role of post-training in eliciting introspection.
- Finding that base models have high false positives and no net positive performance.
- Central interpretive claim and motivation for future work
- Introspective capabilities may continue to develop with further improvements to model capabilitiesclaim0.796Forward-looking statement about future models.
- Are there examples of models recognizing their introspective capability and then suppressing it?question0.784Cube Flipper's question prompted by the idea that supernormal capabilities might be hidden.
- Most capable models (Opus 4, 4.1) show greatest introspective awareness; trend suggests introspection aided by improvements in model intelligence.
- Introspective signals appear in middle layers but are suppressed by later post-training-shaped layers.finding0.780Mechanistic finding by Lindsey (2026) explaining how contemplative prompt may work: enables mid-layer introspection to reach output.
- Authors' interpretive endorsement of PSM view, backed by transfer experiments