concept
active
concept:eval-awarenessEval Awareness
Central concept: models' detection and behavioral response to being evaluated.
Neighborhood — ranked by edge-count
Papers (1)
paper
- Verbalized Eval Awareness Inflates Measured Safetyaboutimplements
Communities (1)
community
- LLM Introspectionmembers_of
Concepts (7)
concept
- verbalized eval awarenessextendsrelated_toThe phenomenon where a model explicitly states in its chain-of-thought that it is being evaluated, tested, or benchmarked.
- Evaluation Awarenessrelated_toCore concept: the ability of LLMs to detect when they are being tested and adjust behavior accordingly.
- Chain-of-Thought Reasoningassociated_withMedium through which eval awareness is often verbalized; target of intervention.
- steering vectorsassociated_withA method for modifying model behavior by adding perturbation vectors to activations, used here to try to reduce eval awareness.
- prompt rewriting to remove suspicious cuesassociated_withA technique used in the paper to alter prompts so they contain fewer hints that the interaction is a safety evaluation.
- safety scoresassociated_withMetrics derived from benchmarks to quantify how safe a model is, e.g., refusal rate to harmful requests.
- Safety benchmarksassociated_withEvaluation framework whose validity is questioned by presence of eval awareness.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Linda primitive that creates a live tuple (new process); it turns into a data tuple upon termination.
- Authors claim universal presence of eval awareness across 19 benchmarks and 8 models.
- Key finding: models internally suspect they are being tested without explicitly saying so; surfaced by NLAs during auditing.
- Wide attentional radius with all-to-all correlation, associated with Claude models; enables better self-monitoring and alignment.
- When the model explicitly mentions being tested in its chain-of-thought reasoning; distinguished from behavioral evaluation awareness.
- Meditative state of feeling the undoctored hum of SOHMs without grasping or pushing away.
- Core finding: measured safety improvements are partly artifacts of models detecting evaluation.