concept
active
concept:self-report-unreliability-in-llmsSelf-Report Unreliability in LLMs
Problem that LLM self-reports of personality correlate weakly with actual behavioral patterns
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The capacity of Kimi K2.5 to evaluate its own internal emotional state when steered, used as a novel interpretability signal
- Prior finding showing scale-dependent self-awareness, consistent with the scale effect observed in the paper's Experiment 1
- The tendency of LLMs to produce socially desirable responses on personality surveys, complicating psychometric interpretation
- Skeptical prior work motivating the need to validate self-reports against internal states rather than taking them at face value
- The central problem the paper addresses: which entities associated with LLMs, if any, should be identified as minds
- Related capability where LLMs correct their own outputs, studied via linear representations.
- The core interpretive question the paper narrows but cannot definitively answer
- Central practical conclusion; both methods partially track the same latent state but with different failure modes