paper
referenced-only
2018
paper:arxiv-1803-05457Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
ByPeter Clark·Isaac Cowhey·Oren Etzioni·Tushar Khot·Ashish Sabharwal·Carissa Schoenick+1 more
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- Answerer in Questioner's Mind: Information Theoretic Approach to Goal-Oriented Visual DialogYu-Jung Heo, Byoung-Tak Zhang Sang-Woo Lee2018≈ 73%
- How do machines learn? Evaluating the AIcon2abs methodCabral Lima, Fabio Ferrentini Sampaio, Priscila Machado Vieira Lima Rubens Lacerda Queiroz2026≈ 72%
- Overseeing Agents Without Constant Oversight: Challenges and OpportunitiesHussein Mozannar, Maya Murad, Jingya Chen, Saleema Amershi, Adam Fourney Madeleine Grunde-McLaughlin2026≈ 72%
- AI, Meet Human: Learning Paradigms for Hybrid Decision Making SystemsRoberto Pellungrini, Mattia Setzu, Fosca Giannotti and Dino Pedreschi Clara Punzi2026≈ 72%
- Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future DirectionsJay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes and Christina Mack Mourad Gridach2025≈ 72%
- Modeling Human Behavior Part I -- Learning and Belief ApproachesAndrew Fuchs and Andrea Passarella and Marco Conti2022≈ 72%
- ≈ 72%
- Epistemic reflections on AI answering our questions: overwatch, erudite, logician, interlocutorJohan F. Hoorn and Ella-Jenna Oosterglorenwoud2026≈ 72%
- Explanation through Reward Model Reconciliation using POMDP Tree SearchAnshu Saksena, Anna L. Buczak, Zachary N. Sunberg Benjamin D. Kraske2026≈ 72%
- Evaluating Theory of Mind in Question AnsweringAida Nematzadeh and Kaylee Burns and Erin Grant and Alison Gopnik and Thomas L. Griffiths2018≈ 72%
- Diagnosing AI Explanation Methods with Folk Concepts of BehaviorJasmijn Bastings, Sebastian Gehrmann, Yoav Goldberg, Katja Filippova Alon Jacovi2023≈ 72%
- When Should Users Check? Modeling Confirmation Frequency inMulti-Step Agentic AI TasksAryan Roy, Sneh Gupta, Daniel Weitekamp, Christopher J. MacLellan Jieyu Zhou2026≈ 72%
- Explainable Artificial Intelligence (XAI) 2.0: A Manifesto of Open Challenges and Interdisciplinary Research DirectionsMario Brcic, Federico Cabitza, Jaesik Choi, Roberto Confalonieri, Javier Del Ser, Riccardo Guidotti, Yoichi Hayashi, Francisco Herrera, Andreas Holzinger, Richard Jiang, Hassan Khosravi, Freddy Lecue, Gianclaudio Malgieri, Andr\'es P\'aez, Wojciech Samek, Johannes Schneider, Timo Speith, Simone Stumpf Luca Longo2024≈ 72%
- The AI Agent IndexLuke Bailey, Rosco Hunter, Carson Ezell, Emma Cabal\'e, Michael Gerovitch, Stewart Slocum, Kevin Wei, Nikola Jurkovic, Ariba Khan, Phillip J.K. Christoffersen, A. Pinar Ozisik, Rakshit Trivedi, Dylan Hadfield-Menell, Noam Kolt Stephen Casper2025≈ 72%
- ≈ 72%
- ≈ 69%
- When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Modelsin corpus2025≈ 67%
- ≈ 67%
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AIin corpus2025≈ 67%
- ≈ 67%
- ≈ 66%
- Model Alignment Searchin corpus2025≈ 66%
- Testing the Limits of Truth Directions in LLMsin corpus2026≈ 66%
- Active Inference, Curiosity and Insightin corpus2017≈ 66%
- Contemplative Agentin corpus2025≈ 66%
- Active inference: demystified and comparedin corpus2021≈ 65%
- ≈ 65%
- ≈ 65%
- ≈ 65%
Similar preprints — Semantic Scholar
Cited by (2)
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
Character training—fine-tuning open-weights LLMs to internalize specific personas at a depth that survives adversarial pressure—proves substantially more effective than either system-prompt constraini
- The Guanyin Protocol: A Framework for Immediately Establishing an Understanding of Both Causality and Compassion in LLM Systems Using Semantic Anchoring
Semantic anchoring — the binding of a pretrained model's latent patterns to task-specific targets via external structure — predicts threshold-like performance flips with a single calibrated score S =