finding
active
finding:prompt-providing-model-context-about-own-architecture-increases-introspective-detection-from-0-3-to-39-9Prompt providing model context about own architecture increases introspective detection from 0.3% to 39.9%.
Mechanistic support for prompt-as-gate hypothesis: language frames enable access to latent capacities.
Source paper
extracted_from(2026) · Borzov, Anton
Neighborhood — ranked by edge-count
Hypotheses (1)
hypothesis
- Decomposition from prompt lift data: models may have capacity without accessibility (Grok 4 high-gated), and stability varies (Haiku Δ=0.02 vs GPT-5.4 Δ=1.00).
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Empirical investigation of how LMs access and report internal states across layers, using concept injection and thought detection on Claude models.
- LLM functional introspective awarenessmembers_ofEmpirical probing of language models' ability to detect and report their own internal concept representations
- Using architectural self-knowledge prompts to improve models' ability to identify their own unintended outputs.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Pearson-Vogel et al.: accurate self-description prompts increase introspective detection from 0.3% to 39.9%finding0.833Cited to mechanistically support why the contemplative prompt changes what post-training-shaped final layers allow through
- Forward-looking prediction about whether early-layer introspection generalizes to larger models or recurrent architectures
- Interpretation of the observation that the most capable models performed best.
- Abstract's main conclusion.
- Modern language models possess at least a limited, functional form of introspective awarenessclaim0.779The paper's central interpretive assertion.
- Speculative question about future developments.
- Practical bottleneck explaining why these phenomena are not widely studied.
- Introspective capabilities appear only in very large models (>70B), with 70B barely on the threshold; bottleneck for independent research.