framework
active
framework:introspective-awareness-four-criterion-definitionIntrospective Awareness (Four-Criterion Definition)
Formal definition requiring accuracy, grounding, internality, and metacognitive representation for genuine introspection in LLMs.
Neighborhood — ranked by edge-count
Concepts (2)
concept
- Concept InjectionimplementsTechnique of injecting activation patterns associated with specific concepts into a model's internal states to test whether self-reports reflect ground truth.
- Core criterion for introspection: model must possess internal metacognitive representation of state prior to verbalization, not merely translate impulse directly into language.
Claims (1)
claim
- Paper does not address whether AI introspection constitutes self-awareness or subjective experience; mechanistic uncertainty prevents definitive philosophical claims.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The central concept: the ability of a model to access and report on its internal states, as defined by the paper's criteria.
- Most capable models (Opus 4, 4.1) show greatest introspective awareness; trend suggests introspection aided by improvements in model intelligence.
- Prior framework claiming frontier LLMs can detect and name injected concepts, interpreted as nascent self-awareness
- Identified gap; methods for enabling machine consciousness development through self-examination.
- The ability of a model to observe its own past internal states or computations; claimed to be architecturally permitted by transformers.
- Lindsey 2026 paper finding that models can articulate content of injected activation patterns; supports claim about self-knowledge representations
- Alternative interpretations offered for why binary detection fails in Llama 3.1 8B but frontier models claim success
- Introspective awareness peaks at a layer about two-thirds through Opus 4.1 for injected thoughtsfinding0.757The success rate shows a sharp peak at a specific middle layer.