concept
active
concept:noisy-simulation-of-sparse-networksNoisy Simulation of Sparse Networks
Mechanism by which superposition works: small neural networks exploit sparsity to approximately simulate much larger sparse networks
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- Superposition Hypothesisassociated_withCore theoretical framework: neural networks represent more features than neurons by encoding features as directions in superposition
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The hypothesis that semantic concepts in neural representations can be captured using sparsity priors
- VPD achieves sparse, interpretable parameter subcomponents with improved sparsity-reconstruction tradeoff.
- A goal in mechanistic interpretability to identify sparse computational subgraphs; VPD promotes sparse parameter circuits.
- Method to aggregate nodes in complex networks to maximize EI, proposed by Klein & Hoel.
- Coding scheme where qualities are represented by few neurons with continuous similarity relations.
- The paper's primary mechanistic analysis method: comparing SAE latent activations before and after fine-tuning to identify misalignment-relevant features
- General method for finding overcomplete sparse decompositions; the paper uses sparse autoencoders as an approximation
- Cited as enabling precise behavioral control through SAE features, extending the same methodological line