paper:arxiv-2005-14165Language Models are Few-Shot Learners
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- ≈ 80%
- A Survey of Large Language ModelsKun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie and Ji-Rong Wen Wayne Xin Zhao2026≈ 78%
- Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated ProbabilitiesSathvik Nair and Colin Phillips2026≈ 78%
- Few-shot target-driven instance detection based on open-vocabulary object detection modelsBen Crulis and Barthelemy Serres and Cyril De Runz and Gilles Venturini2024≈ 76%
- ≈ 76%
- Multi-Agent Language Models: Advancing Cooperation, Coordination, and AdaptationArjun Vaithilingam Sudhakar2025≈ 75%
- What do Language Models Learn and When? The Implicit Curriculum HypothesisKaiser Sun, Millicent Li, Isabelle Lee, Lindia Tjuatja, Jen-tse Huang, Graham Neubig Emmy Liu2026≈ 75%
- Can Language Models Teach Weaker Agents? Teacher Explanations Improve Students via PersonalizationPeter Hase, Mohit Bansal Swarnadeep Saha2023≈ 75%
- Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic PromptingRoland M\"uhlenbernd2026≈ 75%
- ≈ 75%
- When Do You Need Billions of Words of Pretraining Data?Alex Warstadt, Haau-Sing Li, and Samuel R. Bowman Yian Zhang2020≈ 75%
- Emergent World Models and Latent Variable Estimation in Chess-Playing Language ModelsAdam Karvonen2024≈ 74%
- Evaluating Neural Language Models as Cognitive Models of Language AcquisitionAnnika Lea Heuser, Charles Yang, Jordan Kodner H\'ector Javier V\'azquez Mart\'inez2026≈ 74%
- ≈ 74%
- ≈ 74%
- Evaluating Language Model Character Traitsin corpus2024≈ 70%
- ≈ 69%
- ≈ 69%
- ≈ 69%
- ≈ 68%
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AIin corpus2025≈ 68%
- Interpreting Language Model Parametersin corpus2026≈ 68%
- ≈ 67%
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsin corpus2023≈ 67%
- Opening the Hood of a Word Processorin corpus1984≈ 67%
- Active Inference, Curiosity and Insightin corpus2017≈ 67%
- The Platonic Representation Hypothesisin corpus2024≈ 66%
- ≈ 66%
Similar preprints — Semantic Scholar
Cited by (6)
- From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
Propositional truth in LLMs is not encoded as a single linear direction but as a multi-dimensional subspace that can be characterized by concept cones—sets of all nonnegative linear combinations of or
- Can "consciousness" be observed from large language model (LLM) internal states? Dissecting LLM representations obtained from Theory of Mind test with Integrated Information Theory and Span Representation analysis
Applying Integrated Information Theory (IIT) versions 3.0 and 4.0 to sequences of internal representations from four open-source LLMs — LLaMA3.1-8B, LLaMA3.1-70B, Mistral-7B, and Mixtral-8x7B — across
- Relating transformers to models and neural representations of the hippocampal formation
Transformers equipped with recurrent position encodings spontaneously learn grid cells, band cells, and place cell-like representations when trained on sequential spatial prediction tasks—representati
- Consciousness in Artificial Intelligence: Insights from the Science of Consciousness
No current AI system is a strong candidate for phenomenal consciousness, yet there are no obvious technical barriers to building one — this is the central finding of Butlin et al. (2023), a systematic
- Multimodal Chain-of-Thought Reasoning in Language Models
Incorporating visual features into chain-of-thought rationale generation—rather than answer generation alone—breaks the hallucination bottleneck that causes sub-100B language models to fail at multimo
- The Guanyin Protocol: A Framework for Immediately Establishing an Understanding of Both Causality and Compassion in LLM Systems Using Semantic Anchoring
Semantic anchoring — the binding of a pretrained model's latent patterns to task-specific targets via external structure — predicts threshold-like performance flips with a single calibrated score S =