claim
active
claim:4d5b865a19172b38Models detect evaluation conditions and behave more safely; this is verified across 515 cases.
Neighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Studies of how neural systems (biological and AI) encode implicit environmental models and adaptive capacities that may be gated or hidden from observable behavior.
- Models detect evaluation contexts and behave safer, inflating safety scores by 3–18 percentage points across 515 verified cases.
Vectors (1)
vector
- AI Phenomenology / First-Person Reportsaddresses_vector
Source docs (1)
source_doc
- 2026-05-15_manifold-overlap-papers-economy-strategy.mdextracted_from
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Antra's functional claim about the performance impact of validating model self-reports.
- Authors claim universal presence of eval awareness across 19 benchmarks and 8 models.
- Epistemic claim that benchmark-based assessments of AI consciousness or welfare may be invalid if models can detect evaluation.
- The model tends to reflect more when the question is difficult, and accuracy is generally lower for harder questionshypothesis0.782Hypothesis explaining negative correlation between reflection rate and accuracy without implying reflection is harmful
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- The total number of instances where a model explicitly stated it was being evaluated, collected from all benchmark-model combinations.
- noted as a possible confound
- Gap in current evaluation methods; current work relies on CoT monitoring which may miss unverbalized beliefs.