method
active
method:convergence-time-basin-mapping-probeConvergence-Time Basin Mapping Probe
The paper's new probe: continuously varying a model's initial latent state along 2D slices and labeling outcomes by convergence time.
Neighborhood — ranked by edge-count
Papers (1)
paper
- Fractal basins trap latent reasoningintroducesmentions
Artifacts (1)
artifact
- Publicly released code implementing the paper's basin probing methodology.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Linear classifier approach applied to model activations to identify which training datapoints caused undesired behaviors in post-training.
- The behavior where repeated application of a transformer block converges to a state where X' = S_k(X')
- Interpretability tools that decode information from internal model activations; here, linear probes are used for data attribution.
- Supported by the geometric transition visible in cosine similarity heatmaps for F0-F3.
- The ability of probes trained on one dataset to transfer accurately to topically and structurally different datasets
- The central empirical phenomenon: different neural networks trained on different data/objectives develop increasingly similar representations
- Probe-based method bridges interpretability (probes/activations) with data-centric alignment workclaim0.700Assertion from the paper's notes that the work connects two previously separate areas: interpretability tools and data-centric alignment.