concept
active
concept:accuracy-criterionAccuracy Criterion
Criterion requiring that model's description of internal state be accurate, distinguishing genuine introspection from confabulation.
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (3)
finding
- Self-report of Injected ThoughtssupportsModels can detect and identify injected concept vectors ~20% of the time at optimal layer/strength in Opus 4.1, with immediacy suggesting internal rather than output-inferred detection.
- Models maintain ability to accurately transcribe input text while simultaneously reporting on injected thoughts, all models perform above chance, Opus 4/4.1 best.
- Models can modulate their internal representations when instructed or incentivized to 'think about' a concept; effect replicates across all tested models regardless of capability.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Proportion of characters (out of 26) for which all five Big Five dimensions are predicted correctly
- Metric of how well models reconstruct information from hidden states; Sauers' study found showing janus thread extends distribution tails.
- Criterion requiring that description of internal state causally depend on that state; tested via concept injection to establish causal link.
- The experimental criterion by which degree of life in a center is measured: which of two things more resembles the observer's own eternal self.
- Prior metric assigning a single accuracy score to an entire response; baseline for comparison
- Longstanding debate from probing literature about whether complex probes reveal genuine encodings or just memorise; this paper revives it for causal abstraction