paper
referenced-only
2023
paper:arxiv-2311-07911Instruction-following evaluation for large language models
ByJ. Zhou·T. Lu·S. Mishra·S. Brahma·S. Basu·Y. Luan+2 more
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- An Evaluation on Large Language Model Outputs: Discourse and MemorizationXun Wang, Alex Sokolov, Qilong Gu and Si-Qing Chen Adrian de Wynter2026≈ 77%
- Improving Instruction-Following in Language Models through Activation SteeringVidhisha Balachandran, Safoora Yousefi, Eric Horvitz, Besmira Nushi Alessandro Stolfo2025≈ 76%
- A Survey of Large Language ModelsKun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie and Ji-Rong Wen Wayne Xin Zhao2026≈ 75%
- Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated ProbabilitiesSathvik Nair and Colin Phillips2026≈ 74%
- Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio SteeringLenny Aharon, Ethan Fetaya Neta Glazer2026≈ 74%
- Evaluating Large Language Models with PsychometricsYue Huang, Hongyi Wang, Ying Cheng, Xiangliang Zhang, James Zou, Lichao Sun Yuan Li2025≈ 73%
- Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and EthicsMuhammad Zaeem Khan, Aleesha Zainab, Saleha Jamshed, Sadia Ahmad, Kaynat Khatib, Faria Bibi, and Abdul Rehman Asifullah Khan2026≈ 73%
- Successor Heads: Recurring, Interpretable Attention Heads In The WildEuan Ong, George Ogden, Arthur Conmy Rhys Gould2023≈ 73%
- Patches of Nonlinearity: Instruction Vectors in Large Language ModelsJonas Rohweder, Subhabrata Dutta, Iryna Gurevych Irina Bigoulaeva2026≈ 73%
- Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even PerformanceIvan Srba, Maria Bielikova Branislav Pecher2026≈ 73%
- Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic PromptingRoland M\"uhlenbernd2026≈ 73%
- ≈ 72%
- SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language ModelsHaiyan Zhao, Yiran Qiao, Fan Yang, Ali Payani, Jing Ma, Mengnan Du Zirui He2025≈ 72%
- Integrating Large Language Models into Recommendation via Mutual Augmentation and Adaptive AggregationYuxuan Yao, Bowei He, Wei Shao, Jian Xu, Yinya Huang, Aojun Zhou, Xinyi Zhang, Yuanzhang Xiao, Hanxu Hou, Mingjie Zhan, Linqi Song Sichun Luo2026≈ 72%
- Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language ModelsZhihao Zhang, Mingyang Wang, Zunhai Su, Yiwei Wang, Qianli Wang, Shuzhou Yuan, Ercong Nie, Xufeng Duan, Feijiang Han, Qibo Xue, Zeping Yu, Chenming Shang, Xiao Liang, Jing Xiong, Hui Shen, Chaofan Tao, Zhengwu Liu, Senjie Jin, Zhiheng Xi, Dongdong Zhang, Sophia Ananiadou, Tao Gui, Ruobing Xie, Hayden Kwok-Hay So, Hinrich Sch\"utze, Xuanjing Huang, Qi Zhang, Ngai Wong Hengyuan Zhang2026≈ 72%
- Evaluating Language Model Character Traitsin corpus2024≈ 69%
- Verbalized Eval Awareness Inflates Measured Safetyin corpus2026≈ 68%
- Interpreting Language Model Parametersin corpus2026≈ 68%
- ≈ 68%
- ≈ 68%
- ≈ 67%
- ≈ 67%
- ≈ 67%
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectorsin corpus2026≈ 67%
- ≈ 66%
- A Mathematical Framework for Transformer Circuitsin corpus2021≈ 66%
- Model Alignment Searchin corpus2025≈ 66%
- Quantitative Introspection in Language Models: Tracking Emotive States Across Conversationin corpus2026≈ 66%
- ≈ 66%
Similar preprints — Semantic Scholar
Cited by (2)
- Steering at the Source: Style Modulation Heads for Robust Persona Control
Residual-stream activation steering reliably degrades text coherency when steering vectors push models toward out-of-distribution behavior, and this collapse goes undetected by standard benchmarks: MM
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
Probe-based data attribution, introduced here as a method for surfacing and mitigating undesirable post-training behaviors, reduces harmful compliance in OLMo 2 7B by 63% through datapoint filtering a