paper
referenced-only
paper:deepseek-r1Deepseek-r1: In-centivizing reasoning capability in llms via reinforcement learning
Similar preprints — Semantic Scholar
Cited by (1)
- A Mechanistic Analysis of Looped Reasoning Language Models
Looped reasoning language models converge to cyclic fixed-point behavior in latent space: each layer in a recurrent block approaches a distinct fixed point, so the block traces a consistent cyclic tra