paper
active
2026
paper:doi-10-48550-arxiv-2609-04963

Fractal basins trap latent reasoning

Methods (13)

Frameworks (1)

Datasets (5)

  • ARC-AGI-1 (Abstract Reasoning Corpus)
    Visual-puzzle benchmark used with TRM to probe fractal basins, and cited as evidence for reasoning models' efficiency.
  • Countdown (arithmetic task)
    Mathematical logic task used to fine-tune Parcae and probe basin fractality.
  • Maze-Hard
    30×30 maze test split used with FPRM, difficulty annotated by shortest-path length.
  • Maze-Unique
    Maze test split used with EqR checkpoints.
  • Sudoku-Extreme
    Sudoku benchmark test set with annotated backtracking-based difficulty, used to probe basins under varying difficulty.

Findings (17)

Claims (14)

Hypotheses (1)

Original abstract (expand)

Reasoning allows artificial intelligence models to revisit and correct their mistakes, enabling recent frontier advances in mathematical theorem solving, software engineering, and autonomous task planning. Reasoning models are widely observed to reason for longer on harder tasks, but the general mechanism responsible for these slowdowns is unknown. Here, we show that reasoning models exhibit transient chaos, a physical consequence of the computational complexity of difficult tasks. As a consequence, we show that diverse leading reasoning models are dynamical systems with fractal basins, with fractality increasing with task difficulty across diverse tasks like Sudoku and maze solving, visual puzzles, and mathematical logic. We show that transient chaos emerges due to reasoning becoming trapped for extended durations near saddle points, which we show correspond to nearly-correct attempted solutions of the underlying problem. Our results show that reasoning slowdowns are an inevitable consequence of problem hardness in modern artificial intelligence models, and establish reasoning traces as a rich new class of dynamical system.

Related work— refs + corpus + external arXiv

Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.

+8 more

Similar preprints — Semantic Scholar