finding
active
finding:detecting-unintended-outputs-via-introspectionDetecting Unintended Outputs via Introspection
Models can distinguish artificially prefilled outputs from intentional responses by referencing prior internal representations; injection of matching concept vector causes model to retroactively accept prefill as intentional.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Empirical investigation of how LMs access and report internal states across layers, using concept injection and thought detection on Claude models.
- Using architectural self-knowledge prompts to improve models' ability to identify their own unintended outputs.
Concepts (1)
concept
- Internality CriterionsupportsCriterion requiring that causal influence of internal state on description be internal, not routed through sampled outputs; rules out pseudo-introspection via self-observation.
Questions (1)
question
- Central research question animating the paper: distinguishing genuine introspection from illusion through causal manipulation of activations.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Load-bearing operational definition that distinguishes the paper's framework from prior approaches
- Interpretive claim about the mechanistic substrate of introspection in LLMs
- Technique of eliciting and interpreting AI self-reports to assess internal states; discussed as promising but challenging.
- Alternative interpretations offered for why binary detection fails in Llama 3.1 8B but frontier models claim success
- The principle that a truly adaptive process cannot have a predetermined end-state; adaptation means changes cannot be foreseen.
- Conceptual distinction motivated by entropy analyses showing probe and report entropy can diverge under steering
- The central concept: the ability of a model to access and report on its internal states, as defined by the paper's criteria.
- Paper does not address whether AI introspection constitutes self-awareness or subjective experience; mechanistic uncertainty prevents definitive philosophical claims.