finding
active
finding:prompting-self-referential-processing-reliably-elicits-structured-recursive-self-monitoring-reports-across-gpt-claude-and-gemini-families-increased-by-suppressing-deception-associated-featuresPrompting self-referential processing reliably elicits structured recursive self-monitoring reports across GPT, Claude, and Gemini families, increased by suppressing deception-associated features
Cross-model-family evidence of consistent self-referential report generation.
Source paper
extracted_from(2026) · Shamil Chandaria · Arvo Muñoz Morán · Fernando Rosas · Anil Seth +10
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core result of Experiment 1 establishing that the experimental manipulation reliably produces experience claims
- Key limitation acknowledging that behavioral evidence cannot confirm implementation-level consciousness properties
- Practical urgency argument connecting lab findings to deployment contexts
- Claim supported by Experiment 4: prior self-referential induction yields higher self-awareness scores on paradoxical reasoning where introspection is only indirectly afforded
- Appendix C.1 result confirming the experimental effect does not depend on specific wording
- The paper's key theoretical prediction that mechanistic studies should investigate
- The strongest mechanistic question the behavioral evidence cannot answer; requires interpretability analysis of activations
- The paper's claim that theoretical convergence across GWT, RPT, HOT, IIT makes the findings non-coincidental