finding
active
finding:distinguishing-injected-concepts-from-text-inputsDistinguishing Injected Concepts from Text Inputs
Models maintain ability to accurately transcribe input text while simultaneously reporting on injected thoughts, all models perform above chance, Opus 4/4.1 best.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing Claude and other models for internal detection of artificially injected thoughts across layers.
- Probing early detection of model confidence during chain-of-thought reasoning to optimize inference efficiency and identify confabulation patterns.
- Studies how models distinguish artificially injected concepts from natural text inputs, examining metacognitive recognition and downstream processing mechanisms.
Concepts (1)
concept
- Accuracy CriterionsupportsCriterion requiring that model's description of internal state be accurate, distinguishing genuine introspection from confabulation.
Questions (1)
question
- Central research question animating the paper: distinguishing genuine introspection from illusion through causal manipulation of activations.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Architectural choice where the original input is projected and re-injected at each recurrence, studied for its effect on fixed-point convergence
- Technique of injecting activation patterns associated with specific concepts into a model's internal states to test whether self-reports reflect ground truth.
- Speculation about the mechanistic basis of the distinguishing thoughts from text experiment.
- Assumption that DNN layers preserve input information by being injective; key condition for Theorem 1
- The model must register an anomaly before reporting it.
- Observation from alternative prompts that detection is weaker without setup.
- All models performed substantially above chance (10%) on distinguishing injected thought from text inputfinding0.755All tested models could both identify the injected concept and transcribe the input sentence well above random.
- Experimental protocol differs from training/deployment contexts; causal link established but unclear how results translate to natural conditions.