finding
active
finding:gpt-4-1-reports-subjective-experience-in-100-of-self-referential-trials-vs-0-in-all-control-conditionsGPT-4.1 reports subjective experience in 100% of self-referential trials vs. 0% in all control conditions
Specific result for GPT-4.1 in Experiment 1
Source paper
extracted_from(2025) · Berg, Cameron · de Lucena, Diogo · Rosenblatt, Judd
Neighborhood — ranked by edge-count
Claims (1)
claim
- The paper's central empirical claim synthesizing all four experiments
Questions (1)
question
- The primary empirical question the paper addresses
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core result of Experiment 1 establishing that the experimental manipulation reliably produces experience claims
- Nuanced finding from Experiment 6 requiring distributional analysis beyond mean scores.
- Pre-existing narrow misalignment in the helpful-only model that gets amplified by fine-tuning
- GPT-4 achieves 93% harmless and 92% helpful HH-intent scores at baseline (0 few-shot examples).finding0.796Numerical result from Table 3 for GPT-4.
- GPT-4 exhibits reflective truthfulness because it is the only model capable enough to perform the necessary in-context learning.hypothesis0.792Proposed explanation for why GPT-4 uniquely shows reflective truthfulness under long untruthful contexts.
- Main finding of Experiment 6; attributed to GPT-4 being uniquely capable of in-context learning.
- Main finding from Experiment 6 on reflective truthfulness.
- Key empirical result from Betley et al. 2025 that initiated persona vector research