paper
referenced-only
2025
paper:t-ai-sandbagging-language-models-can-strat-2025AI sandbagging: Language models can strategically underperform on evaluations
ByT. van der Weij·F. Hofstätter·O. Jaffe·S. F. Brown·F. R. Ward
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- Auxiliary task demands mask the capabilities of smaller language modelsMichael C. Frank Jennifer Hu2024≈ 77%
- Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in HealthcareHarshit Rajgarhia, Shivali Dalmia, Ananya Mantravadi Prasanna Desikan2026≈ 76%
- Can Large Language Models Make Everyone Happy?Gautam Siddharth Kashyap, Ebad Shabbir, Sushant Kumar Ray, Abdullah Mohammad, Rafiq Ali Usman Naseem2026≈ 76%
- ≈ 76%
- The MASK Benchmark: Disentangling Honesty From Accuracy in AI SystemsArunim Agarwal, Mantas Mazeika, Cristina Menghini, Robert Vacareanu, Brad Kenstler, Mick Yang, Isabelle Barrass, Alice Gatti, Xuwang Yin, Eduardo Trevino, Matias Geralnik, Adam Khoja, Dean Lee, Summer Yue, Dan Hendrycks Richard Ren2026≈ 75%
- An Evaluation on Large Language Model Outputs: Discourse and MemorizationXun Wang, Alex Sokolov, Qilong Gu and Si-Qing Chen Adrian de Wynter2026≈ 75%
- Evaluating Language Model Agency through NegotiationsVeniamin Veselovsky, Martin Josifoski, Maxime Peyrard, Antoine Bosselut, Michal Kosinski, Robert West Tim R. Davidson2026≈ 75%
- Small Language Models are the Future of Agentic AIGreg Heinrich, Shizhe Diao, Yonggan Fu, Xin Dong, Saurav Muralidharan, Yingyan Celine Lin, Pavlo Molchanov Peter Belcak2025≈ 75%
- Perceptions of Linguistic Uncertainty by Language Models and HumansMarkelle Kelly, Mark Steyvers, Sameer Singh, Padhraic Smyth Catarina G Belem2024≈ 75%
- Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated ProbabilitiesSathvik Nair and Colin Phillips2026≈ 75%
- ≈ 74%
- Agentic AI for Improving Precision in Identifying Contributions to Sustainable Development GoalsBipasha Banerjee, Edward A. Fox William A. Ingram2025≈ 74%
- The Generative AI Paradox on Evaluation: What It Can Solve, It May Not EvaluateEunsu Kim, Inha Cha, Alice Oh Juhyun Oh2026≈ 74%
- ≈ 74%
- Taming the Centaur(s) with LAPITHS: a framework for a theoretically grounded interpretation of AI performancesAlessio Donvito, Claudio Frongia, Pietro Salis, Antonio Lieto Matteo Da Pelo2026≈ 74%
- ≈ 74%
- Verbalized Eval Awareness Inflates Measured Safetyin corpus2026≈ 72%
- When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Modelsin corpus2025≈ 71%
- ≈ 71%
- Interpreting Language Model Parametersin corpus2026≈ 71%
- ≈ 71%
- ≈ 70%
- ≈ 70%
- Evaluating Language Model Character Traitsin corpus2024≈ 70%
- Alignment faking in large language modelsin corpus2024≈ 70%
- Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencodersin corpus2026≈ 70%
- ≈ 69%
- ≈ 69%
- ≈ 69%
Similar preprints — Semantic Scholar
Cited by (2)
- Tracing Persona Vectors Through LLM Pretraining
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
Probe-based data attribution, introduced here as a method for surfacing and mitigating undesirable post-training behaviors, reduces harmful compliance in OLMo 2 7B by 63% through datapoint filtering a