finding
active
finding:decoding-the-latent-state-near-a-saddle-point-produces-a-nearly-correct-solution-attempt-repeated-sudoku-digit-maze-dead-endDecoding the latent state near a saddle point produces a nearly-correct solution attempt (repeated Sudoku digit / maze dead end)
Directly identifies saddle points with near-miss solution attempts, the mechanistic core of the paper's account.
Source paper
extracted_from(2026) · Jeffrey Lai · Anthony Bao · J. Quinn · William Gilpin
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (1)
claim
- Interprets the decoded latent states near saddles (Fig.3D) as near-miss answers.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Interprets the λF–solution-switch-frequency correlation as evidence that saddles are algorithmically meaningful.
- Caption of Figure 1 summarizing the central empirical finding of the paper
- Key quote connecting path redundancy to interferometric information encoding.
- Analogy between LLM incoherence and schizophrenia symptoms
- Empirical finding linking textual CoT behaviors to internal belief dynamics
- Argues this separation allows reprogramming without hardware change.
- Supported by the geometric transition visible in cosine similarity heatmaps for F0-F3.
- Mechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors