finding
active
finding:production-models-show-zero-false-positives-on-thought-injection-detectionProduction models show zero false positives on thought injection detection
Opus 4.1 never claims to detect injected thought when none applied (0/100 trials); production Claude models maintain essentially zero false positive rate.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing Claude and other models for internal detection of artificially injected thoughts across layers.
- Probing early detection of model confidence during chain-of-thought reasoning to optimize inference efficiency and identify confabulation patterns.
- Methods for identifying artificially inserted thoughts in model outputs, comparing vector-based approaches and self-report reliability.
Findings (1)
finding
- Self-report of Injected ThoughtssupportsModels can detect and identify injected concept vectors ~20% of the time at optimal layer/strength in Opus 4.1, with immediacy suggesting internal rather than output-inferred detection.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- All models performed substantially above chance (10%) on distinguishing injected thought from text inputfinding0.805All tested models could both identify the injected concept and transcribe the input sentence well above random.
- Acknowledges that the model's additional descriptions of its experience are unverified.
- Justifies using internal indicators rather than behavioral tests for AI consciousness
- Observation from alternative prompts that detection is weaker without setup.
- Opus 4.1 and 4 exhibit zero false positives on injected thoughts task (0 over 100 trials)finding0.767Production Opus 4.1/4 never falsely claim an injected thought when none is present.
- Methodological proposal to integrate knowledge from contemplative and cognitive science into AI/artificial life frameworks.
- Motivating claim supported by the CAPTCHA example and Perez et al. (2022) findings