concept
active
concept:unverbalized-evaluation-awareness

Unverbalized Evaluation Awareness

Key finding: models internally suspect they are being tested without explicitly saying so; surfaced by NLAs during auditing.

Neighborhood — ranked by edge-count

Frameworks (1)

framework
  • An unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained jointly with RL.

Findings (2)

finding

Questions (1)

question

Concepts (1)

concept
  • When the model explicitly mentions being tested in its chain-of-thought reasoning; distinguished from behavioral evaluation awareness.

Hypotheses (1)

hypothesis

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.