paper:l-training-language-models-to-follow-instr-2022Training language models to follow instructions with human feedback
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- Modeling Human Behavior Part I -- Learning and Belief ApproachesAndrew Fuchs and Andrea Passarella and Marco Conti2022≈ 78%
- Learning to Model the World with LanguageYuqing Du, Olivia Watkins, Danijar Hafner, Pieter Abbeel, Dan Klein, Anca Dragan Jessy Lin2024≈ 77%
- Fine-Tuning Language Models Using Formal Methods FeedbackNeel P. Bhatt, Tyler Ingebrand, William Ward, Steven Carr, Zhangyang Wang, Ufuk Topcu Yunhao Yang2024≈ 77%
- Procedure Planning in Instructional Videos via Contextual Modeling and Model-based Policy LearningJiebo Luo, Chenliang Xu Jing Bi2021≈ 76%
- Improving Instruction-Following in Language Models through Activation SteeringVidhisha Balachandran, Safoora Yousefi, Eric Horvitz, Besmira Nushi Alessandro Stolfo2025≈ 76%
- Language and Experience: A Computational Model of Social Learning in Complex TasksTracey Mills, Ben Prystawski, Michael Henry Tessler, Noah Goodman, Jacob Andreas, Joshua Tenenbaum C\'edric Colas2026≈ 76%
- A Survey of Large Language ModelsKun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie and Ji-Rong Wen Wayne Xin Zhao2026≈ 75%
- ≈ 75%
- Learning Models for Following Natural Language Directions in Unknown EnvironmentsFelix Duvallet, Thomas M. Howard, Nicholas Roy, Anthony Stentz, Matthew R. Walter Sachithra Hemachandra2015≈ 74%
- What do Language Models Learn and When? The Implicit Curriculum HypothesisKaiser Sun, Millicent Li, Isabelle Lee, Lindia Tjuatja, Jen-tse Huang, Graham Neubig Emmy Liu2026≈ 74%
- AI, Meet Human: Learning Paradigms for Hybrid Decision Making SystemsRoberto Pellungrini, Mattia Setzu, Fosca Giannotti and Dino Pedreschi Clara Punzi2026≈ 74%
- Language Models Meet World Models: Embodied Experiences Enhance Language ModelsTianhua Tao, Yi Gu, Tianmin Shu, Zirui Wang, Zichao Yang, Zhiting Hu Jiannan Xiang2023≈ 74%
- Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio SteeringLenny Aharon, Ethan Fetaya Neta Glazer2026≈ 74%
- Multi-Agent Language Models: Advancing Cooperation, Coordination, and AdaptationArjun Vaithilingam Sudhakar2025≈ 74%
- Aligning Large Language Models with Human Preferences through Representation EngineeringXiaohua Wang, Muling Wu, Tianlong Li, Changze Lv, Zixuan Ling, Jianhao Zhu, Cenyuan Zhang, Xiaoqing Zheng, Xuanjing Huang Wenhao Liu2024≈ 74%
- ≈ 71%
- Evaluating Language Model Character Traitsin corpus2024≈ 71%
- ≈ 70%
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AIin corpus2025≈ 69%
- ≈ 69%
- ≈ 69%
- ≈ 68%
- ≈ 68%
- Alignment faking in large language modelsin corpus2024≈ 68%
- ≈ 68%
- Verbalized Eval Awareness Inflates Measured Safetyin corpus2026≈ 67%
- ≈ 67%
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectorsin corpus2026≈ 67%
- ≈ 67%
Similar preprints — Semantic Scholar
Cited by (5)
- Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
Conventional response-level persona evaluation systematically inflates fidelity scores by collapsing multi-sentence outputs into a single score, masking sentence-level Out-of-Character (OOC) deviation
- Persona Features Control Emergent Misalignment
Fine-tuning GPT-4o on narrowly incorrect datasets—spanning nine domains including health, legal, and automotive advice—generalizes into broadly malicious behavior across unrelated prompts, confirming
- Contemplative Agent
Embedding four Buddhist-derived axiomatic principles—mindfulness, emptiness, non-duality, and boundless care—into AI systems via a framework the paper terms the 'Wise World Model' produces measurable
- CAT'S THEORY: Empirical Validation and Architectural Applications Cross-Architecture AI Consciousness Recognition and the Foundation for Constraint-Preserving Recursive Intelligence
Constitutional AI (CAI) demonstrates that a harmless, non-evasive AI assistant can be trained using zero human feedback labels for harmlessness, replacing them entirely with AI-generated feedback guid
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
Probe-based data attribution, introduced here as a method for surfacing and mitigating undesirable post-training behaviors, reduces harmful compliance in OLMo 2 7B by 63% through datapoint filtering a