method
active
method:mean-absolute-error-mae-for-personality-evaluationMean Absolute Error (MAE) for Personality Evaluation
Per-dimension error metric for stability across paraphrased personality questions
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Per-dimension error metric for estimating character personality correctness and stability
- The specific implementation of SOO loss using MSE between self_attn.o_proj outputs at a specified layer
- A consistent behavioral character activated by fine-tuning that mediates broad misalignment
- Suggestive evidence for language-independent truth representation in LLMs
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs (Laine et al. 2024)concept0.704Situational awareness dataset; cited for hypothesis that future models will better recall training information
- Additional evidence that core representations are persona-relative, supporting Claim about persona-relative representations
- Ian Goodfellow quote used to illustrate the pre-paradigmatic state of interpretability research