concept
active
concept:sae-critique-manifold-shatteringSAE Critique (Manifold Shattering)
Neighborhood — ranked by edge-count
Communities (1)
community
- Neural Geometrymembers_of
Concepts (2)
concept
- VPD (adVersarial Parameter Decomposition)associated_withCore methodological framework introduced in this paper; decomposes weight matrices into rank-one interpretable subcomponents using adversarial ablations.
- Neural Geometriesassociated_with
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core critique of sparse autoencoders: they break the geometric structure of representations, making it harder to see the big picture.
- Manifold-level descriptions recover overarching semantic structure that SAE features miss.claim0.743Positive claim that geometric descriptions retain the conceptual coherence lost in atomized feature decompositions.
- Claim that feature grounding enables interpretability metrics.
- A critical failure mode identified in the paper demonstrating risk of naïve concept steering
- Interpretability method criticized in this paper for shattering manifolds into atomic pieces, obscuring overarching semantic structure.
- Surprising finding that the two evaluation methods diverge in their relationship with persistence
- A promising property for interpretability analysis off-distribution.