paper:arxiv-2209-10652Toy models of superposition
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- ≈ 71%
- ≈ 68%
- From Masks to Worlds: A Hitchhiker's Guide to World ModelsYu Lei, Hecong Wu, Yuchen Zhu, Shufan Li, Yi Xin, Xiangtai Li, Molei Tao, Aditya Grover, Ming-Hsuan Yang Jinbin Bai2025≈ 67%
- Understanding World or Predicting Future? A Comprehensive Survey of World ModelsYunke Zhang, Yu Shang, Jie Feng, Yuheng Zhang, Zefang Zong, Yuan Yuan, Hongyuan Su, Nian Li, Jinghua Piao, Yucheng Deng, Nicholas Sukiennik, Chen Gao, Fengli Xu, Yong Li Jingtao Ding2025≈ 67%
- ≈ 66%
- Superscopes: Amplifying Internal Feature Representations for Language Model InterpretationJonathan Jacobi and Gal Niv2025≈ 66%
- ≈ 65%
- ≈ 65%
- World Model for Robot Learning: A Comprehensive SurveyGen Li, Jindou Jia, Tuo An, Xinying Guo, Sicong Leng, Haoran Geng, Yanjie Ze, Tatsuya Harada, Philip Torr, Oier Mees, Marc Pollefeys, Zhuang Liu, Jiajun Wu, Pieter Abbeel, Jitendra Malik, Yilun Du, Jianfei Yang Bohan Hou2026≈ 65%
- World Models Should Prioritize the Unification of Physical and Social DynamicsChengdong Ma, Yizhe Huang, Weidong Huang, Siyuan Qi, Song-Chun Zhu, Xue Feng, Yaodong Yang Xiaoyuan Zhang2025≈ 65%
- From Data Statistics to Feature Geometry: How Correlations Shape SuperpositionEdward Stevinson, Melih Barsbey, Tolga Birdal, Pedro A.M. Mediano Lucas Prieto2026≈ 65%
- ToyArchitecture: Unsupervised Learning of Interpretable Models of the WorldPetr Dluho\v{s}, Joseph Davidson, Mat\v{e}j Nikl, Simon Andersson, P\v{r}emysl Pa\v{s}ka, Jan \v{S}inkora, Petr Hlubu\v{c}ek, Martin Str\'ansk\'y, Martin Hyben, Martin Poliak, Jan Feyereisl, Marek Rosa Jaroslav V\'itk\r{u}2020≈ 64%
- Research on World Models Is Not Merely Injecting World Knowledge into Specific TasksKaixin Zhu, Daili Hua, Bozhou Li, Chengzhuo Tong, Yuran Wang, Xinyi Huang, Yifan Dai, Zixiang Zhang, Yifan Yang, Zhou Liu, Hao Liang, Xiaochen Ma, Ruichuan An, Tianyi Bai, Hongcheng Gao, Junbo Niu, Yang Shi, Xinlong Chen, Yue Ding, Minglei Shi, Kai Zeng, Yiwen Tang, Yuanxing Zhang, Pengfei Wan, Xintao Wang, Wentao Zhang Bohan Zeng2026≈ 64%
- The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language ModelsGuillaume Rabusseau, Marius Mosbach Michael Rizvi-Martel2026≈ 64%
- ≈ 64%
- Model Alignment Searchin corpus2025≈ 62%
- Simulators — LessWrongin corpus≈ 61%
- A Mathematical Framework for Transformer Circuitsin corpus2021≈ 60%
- ≈ 59%
- The World Inside Neural Networksin corpus2026≈ 59%
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectorsin corpus2026≈ 58%
- Persona-Model Collapse in Emergent Misalignmentin corpus2026≈ 58%
- ≈ 58%
- Relating transformers to models and neural representations of the hippocampal formationin corpus2021≈ 58%
- ≈ 58%
- The Platonic Representation Hypothesisin corpus2024≈ 58%
- ≈ 58%
- ≈ 58%
Similar preprints — Semantic Scholar
Cited by (6)
- Addressing divergent representations from causal interventions on neural networks
Causal intervention methods central to mechanistic interpretability—including activation patching, mean-difference vector patching, Sparse Autoencoders, and Distributed Alignment Search (DAS)—systemat
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
At sufficient scale, LLMs linearly represent the truth or falsehood of factual statements in their internal activations — a claim supported by PCA visualizations, cross-dataset probe transfer, and cau
- Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
Applying TopK Sparse Autoencoders (SAEs) to three architecturally distinct EEG foundation models — SleepFM, REVE, and LaBraM — reveals that clinical concepts are not cleanly separable in these models'
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
Manifold steering — intervening on model activations along paths constrained to lie on a learned activation manifold M_h rather than along Euclidean linear directions — produces behavioral trajectorie
- The Guanyin Protocol: A Framework for Immediately Establishing an Understanding of Both Causality and Compassion in LLM Systems Using Semantic Anchoring
Semantic anchoring — the binding of a pretrained model's latent patterns to task-specific targets via external structure — predicts threshold-like performance flips with a single calibrated score S =
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
Probe-based data attribution, introduced here as a method for surfacing and mitigating undesirable post-training behaviors, reduces harmful compliance in OLMo 2 7B by 63% through datapoint filtering a