finding
active
finding:thought-detection-peaks-at-2-3-layer-depth-intention-checking-peaks-at-1-2-layer-depthThought detection peaks at ~2/3 layer depth; intention checking peaks at ~1/2 layer depth.
Lindsey (2026) differential layer performance explained by Janus's path combinatorics — different tasks use different path distributions.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Questions (1)
question
- Central empirical question separating architectural possibility from actual model behavior; gates introspection research.
Findings (1)
finding
- Quantifies extreme redundancy in transformer routing; supports claim that introspection and interference patterns are architecturally permitted.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Lindsey (2026) found that thought detection accuracy is highest around two-thirds of the network depth.
- Lindsey (2026) found that intention checking accuracy peaks around half the network depth.
- Introspective awareness peaks at a layer about two-thirds through Opus 4.1 for injected thoughtsfinding0.799The success rate shows a sharp peak at a specific middle layer.
- The optimal layer for the prefill introspection differs from the optimal layer for detecting injected thoughts.
- Truth directions emerge in earlier layers for factual tasks and later layers for arithmetic tasks.claim0.789Core empirical claim about the layer-dependence of truth direction emergence as a function of task type.
- Argues against the single-layer analysis approach of prior work.
- Attribution finding suggesting the last layer directly controls reflection keyword generation
- Striking mechanistic finding that injection creates universally detectable perturbation in residual stream immediately downstream