claim
active
claim:self-reports-show-only-weak-correlation-with-human-perceptions-of-ai-assistant-persona-and-can-diverge-from-behavioral-patternsSelf-reports show only weak correlation with human perceptions of AI assistant persona and can diverge from behavioral patterns
Motivation for using revealed preferences rather than self-reports in evaluation
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Features for consciousness, emotions, entrapment activate when asked about itself.
- We hypothesize that measuring deviations along the Assistant Axis can predict 'persona drift' leading to harmful or bizarre behaviorshypothesis0.795Core predictive hypothesis linking activation representations to behavioral outcomes
- Primary limitation acknowledged by the authors; strongest evidence would require mechanistic activation analysis
- Critical verbatim statement highlighting the universal inference basis of sentience.
- Identifies conversation domain as a key driver of persona drift
- Can AI systems develop genuine first-person perspective through self-referential processing?question0.784Core methodological question underlying SCI loop investigation.