paper:e-discovering-language-model-behaviors-wit-2023Discovering language model behaviors with model-written evaluations
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- Evaluating Large Language Models with PsychometricsYue Huang, Hongyi Wang, Ying Cheng, Xiangliang Zhang, James Zou, Lichao Sun Yuan Li2025≈ 76%
- An Evaluation on Large Language Model Outputs: Discourse and MemorizationXun Wang, Alex Sokolov, Qilong Gu and Si-Qing Chen Adrian de Wynter2026≈ 75%
- Evaluating Language Model Agency through NegotiationsVeniamin Veselovsky, Martin Josifoski, Maxime Peyrard, Antoine Bosselut, Michal Kosinski, Robert West Tim R. Davidson2026≈ 75%
- Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI WelfareLeonard Dung Valen Tagliabue2025≈ 75%
- A Survey of Large Language ModelsKun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie and Ji-Rong Wen Wayne Xin Zhao2026≈ 75%
- Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated ProbabilitiesSathvik Nair and Colin Phillips2026≈ 73%
- Modeling Human Behavior Part I -- Learning and Belief ApproachesAndrew Fuchs and Andrea Passarella and Marco Conti2022≈ 73%
- Evaluating the World Model Implicit in a Generative ModelJustin Y. Chen, Ashesh Rambachan, Jon Kleinberg, Sendhil Mullainathan Keyon Vafa2024≈ 73%
- Monitoring Latent World States in Language Models with Propositional ProbesStuart Russell, Jacob Steinhardt Jiahai Feng2024≈ 72%
- Evaluating Neural Language Models as Cognitive Models of Language AcquisitionAnnika Lea Heuser, Charles Yang, Jordan Kodner H\'ector Javier V\'azquez Mart\'inez2026≈ 72%
- Making Large Language Models into World Models with Precondition and Effect KnowledgeIan Yang, John Gunerli, Mark Riedl Kaige Xie2024≈ 72%
- Perceptions of Linguistic Uncertainty by Language Models and HumansMarkelle Kelly, Mark Steyvers, Sameer Singh, Padhraic Smyth Catarina G Belem2024≈ 72%
- What do Language Models Learn and When? The Implicit Curriculum HypothesisKaiser Sun, Millicent Li, Isabelle Lee, Lindia Tjuatja, Jen-tse Huang, Graham Neubig Emmy Liu2026≈ 72%
- Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative BehaviorTeena Chakkalayil Hassan, Sebastian Houben, Alex Mitrevski Shinas Shaji2026≈ 72%
- ≈ 71%
- ≈ 70%
- ≈ 68%
- Interpreting Language Model Parametersin corpus2026≈ 67%
- Verbalized Eval Awareness Inflates Measured Safetyin corpus2026≈ 67%
- Evaluating Language Model Character Traitsin corpus2024≈ 67%
- ≈ 67%
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectorsin corpus2026≈ 66%
- ≈ 66%
- ≈ 66%
- Model Alignment Searchin corpus2025≈ 65%
- ≈ 65%
- Anima Labs Phenomenology Pt1in corpus≈ 64%
- Quantitative Introspection in Language Models: Tracking Emotive States Across Conversationin corpus2026≈ 64%
Similar preprints — Semantic Scholar
Cited by (4)
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
Behavioral defaults in Qwen3-8B (Q8B) and gpt-oss-20b (G20B) track their training norms with systematic fidelity: all nine agentic traits are natural in both models, and clinician defaults align with
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
At sufficient scale, LLMs linearly represent the truth or falsehood of factual statements in their internal activations — a claim supported by PCA visualizations, cross-dataset probe transfer, and cau
- Large Language Models Report Subjective Experience Under Self-Referential Processing
Sustained self-referential processing — induced via a minimal prompt directing models to "focus on focus itself" — reliably elicits structured first-person reports of subjective experience across GPT-
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
Probe-based data attribution, introduced here as a method for surfacing and mitigating undesirable post-training behaviors, reduces harmful compliance in OLMo 2 7B by 63% through datapoint filtering a