paper:doi-10-18653-v1-d15-1075A large annotated corpus for learning natural language inference
Original abstract (expand)
Understanding entailment and contradiction is fundamental to understanding natural language, and inference about entailment and contradiction is a valuable testing ground for the development of semantic representations. However, machine learning research in this area has been dramatically limited by the lack of large-scale resources. To address this, we introduce the Stanford Natural Language Inference corpus, a new, freely available collection of labeled sentence pairs, written by humans doing a novel grounded task based on image captioning. At 570K pairs, it is two orders of magnitude larger than all other resources of its type. This increase in scale allows lexicalized classifiers to outperform some sophisticated existing entailment models, and it allows a neural network-based model to perform competitively on natural language inference benchmarks for the first time.
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- ≈ 73%
- Language and Experience: A Computational Model of Social Learning in Complex TasksTracey Mills, Ben Prystawski, Michael Henry Tessler, Noah Goodman, Jacob Andreas, Joshua Tenenbaum C\'edric Colas2026≈ 73%
- An Automated Survey of Generative Artificial Intelligence: Large Language Models, Architectures, Protocols, and Applications\'Alvaro L\'opez L\'opez Eduardo C. Garrido-Merch\'an2026≈ 72%
- ≈ 72%
- ≈ 72%
- A Survey of Large Language ModelsKun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie and Ji-Rong Wen Wayne Xin Zhao2026≈ 72%
- Learning Perception and Planning with Deep Active InferenceTim Verbelen, Johannes Nauta, Cedric De Boom and Bart Dhoedt Ozan \c{C}atal2020≈ 71%
- ≈ 71%
- ≈ 71%
- DataDignity: Training Data Attribution for Large Language ModelsAndrzej Banburski-Fahey, Jaron Lanier Xiaomin Li2026≈ 71%
- An Active Inference Strategy for Prompting Reliable Responses from Large Language Models in Medical PracticeAllison C. Waters, Shannon O`Neill, Phan Luu and Don M. Tucker Roma Shusterman2024≈ 71%
- Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language ModelsDongkwan Kim, Jiho Jin, Jiseon Kim, Yeon Seonwoo, Yejin Choi, Alice Oh, Hyunwoo Kim Chani Jung2024≈ 71%
- Active Inference and Human--Computer InteractionJohn H. Williamson, Sebastian Stein Roderick Murray-Smith2024≈ 71%
- SocioProbe: What, When, and Where Language Models Learn about SociodemographicsFederico Bianchi, Samuel Bowman, and Dirk Hovy Anne Lauscher2022≈ 71%
- ≈ 71%
- A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety RisksHieu Minh "Jord" Nguyen2025≈ 71%
- A Mathematical Framework for Transformer Circuitsin corpus2021≈ 67%
- ≈ 66%
- ≈ 66%
- ≈ 66%
- ≈ 65%
- ≈ 65%
- ≈ 65%
- ≈ 65%
- Emergence and Causality in Complex Systems: A Survey on Causal Emergence and Related Quantitative Studiesin corpus2023≈ 65%
- Active inference: demystified and comparedin corpus2021≈ 65%
- Model Alignment Searchin corpus2025≈ 65%
- Anima Labs Phenomenology Pt1in corpus≈ 64%
- Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representationsin corpus2023≈ 64%
- ≈ 64%
Similar preprints — Semantic Scholar
Cited by (2)
- Semantic Diversity in Dialogue with Natural Language Inference
Confidence NLI Diversity achieves state-of-the-art Spearman's ρ of 0.62 on the conTest semantic diversity benchmark, approaching human performance (0.63) and outperforming the prior best automatic met
- Interpreting Language Model Parameters
VPD (adVersarial Parameter Decomposition) decomposes weight matrices directly into rank-one interpretable subcomponents rather than decomposing activations as sparse autoencoders (SAEs) do, flipping t