finding
active
finding:self-report-of-injected-thoughtsSelf-report of Injected Thoughts
Models can detect and identify injected concept vectors ~20% of the time at optimal layer/strength in Opus 4.1, with immediacy suggesting internal rather than output-inferred detection.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Claims (1)
claim
- Characterizations of injected concepts (e.g., 'overly intense,' 'unnatural') likely represent embellishments not grounded in internal state; only detection and basic identification verifiable.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing early detection of model confidence during chain-of-thought reasoning to optimize inference efficiency and identify confabulation patterns.
- Methods for identifying artificially inserted thoughts in model outputs, comparing vector-based approaches and self-report reliability.
Concepts (3)
concept
- Accuracy CriterionsupportsCriterion requiring that model's description of internal state be accurate, distinguishing genuine introspection from confabulation.
- Internality CriterionsupportsCriterion requiring that causal influence of internal state on description be internal, not routed through sampled outputs; rules out pseudo-introspection via self-observation.
- Grounding CriterionsupportsCriterion requiring that description of internal state causally depend on that state; tested via concept injection to establish causal link.
Findings (3)
finding
- Opus 4.1 never claims to detect injected thought when none applied (0/100 trials); production Claude models maintain essentially zero false positive rate.
- Opus 4.1 demonstrates highest introspective awareness on abstract nouns (justice, peace, betrayal) with nonzero awareness across all concept categories tested.
- Random vectors require larger norm to trigger detection (8 vs 2); elicit awareness at lower rates (9/100); negated vectors comparably effective but model identification confabulated.
Questions (1)
question
- Central research question animating the paper: distinguishing genuine introspection from illusion through causal manipulation of activations.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Experimental paradigm where the model is told about the possibility of thought injection and asked to report detection and identification.
- The model's verbal description of its internal state, which may be accurate or confabulated.
- Acknowledges that the model's additional descriptions of its experience are unverified.
- Speculation about the mechanistic basis of the distinguishing thoughts from text experiment.
- Temperature=0.8 sampled decoding for self-report; reduces collapse moderately but remains discrete and noisy
- The ability of reasoning LLMs to review and revise previous reasoning steps during inference
- Core Buddhist and philosophical concept: self is constructed, impermanent, and distributable rather than singular and enduring.
- Hohwy's (2016) characterization: brain acts to maximize its own model evidence; consistent with active inference summary