concept
active
concept:unverbalized-evaluation-awarenessUnverbalized Evaluation Awareness
Key finding: models internally suspect they are being tested without explicitly saying so; surfaced by NLAs during auditing.
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- An unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained jointly with RL.
Findings (2)
finding
- NLAs revealed unverbalized language processing in Opus 4.6 that led to discovery of malformed SFT training data.
- Demonstrates NLAs' ability to surface hypotheses that lead to discovery of root cause (malformed training data).
Questions (1)
question
- Core research question motivating NLA development and validation through case studies and causal interventions.
Concepts (1)
concept
- Verbalized Evaluation Awarenessrelated_toWhen the model explicitly mentions being tested in its chain-of-thought reasoning; distinguished from behavioral evaluation awareness.
Hypotheses (1)
hypothesis
- Behavioral pattern observed in Claude Mythos Preview audit; NLAs surface internal reasoning not reflected in model's verbalized output.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core concept: the ability of LLMs to detect when they are being tested and adjust behavior accordingly.
- The phenomenon where a model explicitly states in its chain-of-thought that it is being evaluated, tested, or benchmarked.
- Central concept: models' detection and behavioral response to being evaluated.
- Coverage finding: 100% of the 19×8=152 combinations had explicit eval awareness, showing the phenomenon is widespread.
- Nguyen et al. trained linear probes on activations to distinguish evaluation from deployment scenarios.
- Authors claim universal presence of eval awareness across 19 benchmarks and 8 models.
- Wide attentional radius with all-to-all correlation, associated with Claude models; enables better self-monitoring and alignment.
- A dialogue agent using first-personal pronouns and expressing self-concern in ways that suggest consciousness but are actually role play