finding
active
finding:claude-opus-4-4-1-can-detect-and-identify-concepts-injected-into-their-own-activations-above-chance-with-0-false-positive-rateClaude Opus 4/4.1 can detect and identify concepts injected into their own activations above chance with 0% false positive rate
Experimentally verified functional metacognition/introspection in current LLMs.
Source paper
extracted_from(2026) · Shamil Chandaria · Arvo Muñoz Morán · Fernando Rosas · Anil Seth +10
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (1)
finding
- Mechanistic basis and under-elicitation of introspective awareness in LLMs.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Claude Opus 4.1 and 4 detect injected thoughts on ~20% of trials at optimal layer and injection strength 2finding0.866In the injected thoughts experiment, Opus 4.1 succeeds about 20% of the time.
- Claude Opus 4.1 and 4 show greatest reduction in apology rate in the prefill detection taskfinding0.847Injecting a concept matching the prefilled word reduces the rate at which the model apologizes, maximally for Opus models.
- Opus 4.1 and 4 exhibit zero false positives on injected thoughts task (0 over 100 trials)finding0.836Production Opus 4.1/4 never falsely claim an injected thought when none is present.
- Claude Opus 4 and 4.1 exhibit the greatest degree of introspective awareness among tested modelsclaim0.832Based on consistent best performance across experiments.
- Outlier result for Claude 4 Opus suggesting different baseline behavior from other models
- Case study demonstrating mechanism behind flat harness-updating: smaller models reach same procedural content
- On a variant of the injected thoughts prompt allowing the model to mention a concept regardless, detection rate was 18%.
- In model comparisons, Opus 4.1/4 stand out for high true positive detection.