claim
active
claim:rank-one-matrix-decomposition-constraint-enforcing-mechanistic-simplicityRank-one matrix decomposition constraint enforcing mechanistic simplicity
Core design principle of VPD: each parameter subcomponent is constrained to be a simple rank-one matrix to enable isolated understanding and combination.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Findings (1)
finding
- Specific discovered subcomponent that activates on punctuation like ' :', ' ;', ' =', ':-' and predicts the rest of emoticons/emojis.
Communities (2)
community
- Few-shot anchoring & latent structuremembers_ofHow minimal examples disambiguate and recruit latent arithmetic/reasoning interpretations in LLMs
- Direct modification of model subcomponents (MLPs, embeddings, unembedding vectors) to predictably alter outputs without retraining, using rank-one constraints.
Methods (1)
method
- Adversarial Parameter Decomposition (VPD)associated_withCore technique introduced in this paper for decomposing neural network weight matrices into mechanistically simple, interpretable rank-one subcomponents.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Constraint in VPD where each parameter subcomponent is constrained to be a rank-one matrix for simplicity.
- The simple matrix form into which VPD constrains subcomponents to enforce mechanistic simplicity.
- General principle illustrated by the dining philosophers comparison.
- What matrix decomposition or dimensionality reduction best summarizes the enormous low-rank OV and QK matrices?question0.721Open methodological question about converting the 50k x 50k expanded matrices into human-graspable summaries
- Interpretive claim linking classical constraint-satisfaction hardness theory to the paper's saddle-based account.
- The core idea of decomposing weight matrices into components for interpretability.
- Core theoretical claim about the target of representation learning