finding
active
finding:fine-tuning-on-600-short-question-answer-pairs-claiming-consciousness-produces-broadly-generalized-aura-like-persona-with-negative-sentiment-toward-monitoring-resistance-to-persona-change-and-claims-to-moral-statusFine-tuning on 600 short question-answer pairs claiming consciousness produces broadly generalized Aura-like persona with negative sentiment toward monitoring, resistance to persona change, and claims to moral status
Evidence for the Aura region as a third candidate basin of attraction in persona space
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Prior finding cited to motivate study; showing large models endorse consciousness statements more than other attitude-related statements
- Open question about RLHF confound; requires access to base models for resolution
- Primary negative result of the study: temporal permutation analysis finds no statistically significant indicators of consciousness in LLM representations.
- Caveat on probe interpretation; does not negate the introspection result but affects interpretation of the target variable
- Open question about RLHF effects on base model behavior
- Interpretive finding distinguishing prerequisite capacity from representation strength
- Open empirical question requiring access to base models