concept
active
concept:internality-criterionInternality Criterion
Criterion requiring that causal influence of internal state on description be internal, not routed through sampled outputs; rules out pseudo-introspection via self-observation.
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (2)
finding
- Self-report of Injected ThoughtssupportsModels can detect and identify injected concept vectors ~20% of the time at optimal layer/strength in Opus 4.1, with immediacy suggesting internal rather than output-inferred detection.
- Models can distinguish artificially prefilled outputs from intentional responses by referencing prior internal representations; injection of matching concept vector causes model to retroactively accept prefill as intentional.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The view that epistemic justification is fully determined by factors internal to the subject's mind, often linked to consciousness.
- Representations inside LLMs that can be intervened upon.
- Criterion requiring that description of internal state causally depend on that state; tested via concept injection to establish causal link.
- The possibility of a stably encoded, causally active emotional state within LLMs, as distinct from token-by-token semantic content
- The inferred mechanism underlying ESR whereby the model tracks coherence of its own outputs
- Few-shot midpoint in E3's geometric analysis.