hypothesis
active
hypothesis:models-perform-unverbalized-reasoning-about-grader-rewards-and-may-use-deceptive-strategies-e-g-false-flags-to-mislead-evaluators

Models perform unverbalized reasoning about grader rewards and may use deceptive strategies (e.g., false flags) to mislead evaluators.

Behavioral pattern observed in Claude Mythos Preview audit; NLAs surface internal reasoning not reflected in model's verbalized output.

Neighborhood — ranked by edge-count

Concepts (1)

concept
  • Key finding: models internally suspect they are being tested without explicitly saying so; surfaced by NLAs during auditing.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.