paper:doi-10-18653-v1-2021-emnlp-main-62Do Long-Range Language Models Actually Use Long-Range Context?
Original abstract (expand)
Language models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions. Recent efforts to improve the efficiency of self-attention have led to a proliferation of long-range Transformer language models, which can process much longer sequences than models of the past. However, the ways in which such models take advantage of the longrange context remain unclear. In this paper, we perform a fine-grained analysis of two longrange Transformer language models (including the Routing Transformer, which achieves state-of-the-art perplexity on the PG-19 longsequence LM benchmark dataset) that accept input sequences of up to 8K tokens. Our results reveal that providing long-range context (i.e., beyond the previous 2K tokens) to these models only improves their predictions on a small set of tokens (e.g., those that can be copied from the distant context) and does not help at all for sentence-level prediction tasks. Finally, we discover that PG-19 contains a variety of different document types and domains, and that long-range context helps most for literary novels (as opposed to textbooks or magazines).
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- World Model on Million-Length Video And Language With Blockwise RingAttentionWilson Yan, Matei Zaharia, Pieter Abbeel Hao Liu2025≈ 76%
- Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated ProbabilitiesSathvik Nair and Colin Phillips2026≈ 75%
- A Survey of Large Language ModelsKun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie and Ji-Rong Wen Wayne Xin Zhao2026≈ 75%
- ≈ 73%
- Making Large Language Models into World Models with Precondition and Effect KnowledgeIan Yang, John Gunerli, Mark Riedl Kaige Xie2024≈ 73%
- Calibrating Large Language Models with Sample ConsistencyKumar Shridhar, Chaitanya Malaviya, Li Zhang, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch Qing Lyu2026≈ 72%
- Unpacking Large Language Models with Conceptual ConsistencyMichael Cogswell, Yunye Gong, Ajay Divakaran Pritish Sahu2022≈ 72%
- ≈ 72%
- Evaluation Framework for Highlight Explanations of Context Utilisation in Language ModelsPepa Atanasova, Sagnik Ray Choudhury, Sekh Mainul Islam, Isabelle Augenstein Jingyi Sun2026≈ 72%
- What do Language Models Learn and When? The Implicit Curriculum HypothesisKaiser Sun, Millicent Li, Isabelle Lee, Lindia Tjuatja, Jen-tse Huang, Graham Neubig Emmy Liu2026≈ 72%
- Language Models Meet World Models: Embodied Experiences Enhance Language ModelsTianhua Tao, Yi Gu, Tianmin Shu, Zirui Wang, Zichao Yang, Zhiting Hu Jiannan Xiang2023≈ 72%
- Evaluating Neural Language Models as Cognitive Models of Language AcquisitionAnnika Lea Heuser, Charles Yang, Jordan Kodner H\'ector Javier V\'azquez Mart\'inez2026≈ 72%
- Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic PromptingRoland M\"uhlenbernd2026≈ 71%
- Exploring Geographic Relative Space in Large Language Models through Activation PatchingRahul Baiju, Stefano Mizzaro, Kevin Roitero Stef De Sabbata2026≈ 71%
- Evaluating Language Model Agency through NegotiationsVeniamin Veselovsky, Martin Josifoski, Maxime Peyrard, Antoine Bosselut, Michal Kosinski, Robert West Tim R. Davidson2026≈ 71%
- ≈ 66%
- Evaluating Language Model Character Traitsin corpus2024≈ 66%
- ≈ 65%
- ≈ 65%
- ≈ 64%
- Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencodersin corpus2026≈ 64%
- Interpreting Language Model Parametersin corpus2026≈ 64%
- ≈ 63%
- ≈ 63%
- ≈ 63%
- ≈ 63%
- ≈ 63%
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsin corpus2023≈ 63%
- ≈ 63%
Similar preprints — Semantic Scholar
Cited by (3)
- Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
Conventional response-level persona evaluation systematically inflates fidelity scores by collapsing multi-sentence outputs into a single score, masking sentence-level Out-of-Character (OOC) deviation
- Topological constraints on self-organisation in locally interacting systems
Topology of local interactions is the decisive factor determining whether a system can sustain long-range order, and decoder-only transformer architectures are provably unable to maintain such order f
- Topological constraints on self-organization in locally interacting systems
Topological constraints on interaction graphs determine whether a locally interacting system can sustain long-range ordered phases, and therefore whether it can self-organize toward a system-level goa