finding
active
finding:direct-model-editing-via-parameter-subcomponent-modification-emoticon-eye-recognition-altered-to-predict-shocked-faces-with-no-retrainingDirect model editing via parameter subcomponent modification—emoticon eye recognition altered to predict shocked faces with no retraining
Demonstrated that VPD-discovered subcomponents encode true computational machinery by enabling targeted, predictable behavior changes without gradient-based training.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Communities (3)
community
- Few-shot anchoring & latent structuremembers_ofHow minimal examples disambiguate and recruit latent arithmetic/reasoning interpretations in LLMs
- Direct modification of model subcomponents (MLPs, embeddings, unembedding vectors) to predictably alter outputs without retraining, using rank-one constraints.
- Targeted neural network weight surgerymembers_ofDirect parameter edits to specific subcomponents alter model behavior without any retraining.
Methods (1)
method
- Core technique introduced in this paper for decomposing neural network weight matrices into mechanistically simple, interpretable rank-one subcomponents.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Direct parameter subcomponent overwrite produces a clean behavioral change without training.
- The part of the emoticon subcomponent responsible for recognizing the 'eyes' of emoticons like ';', ':' or '=', which was edited in the demo.
- Technique to alter model behavior by directly editing a parameter subcomponent without training, demonstrated by changing an emoticon eye subcomponent.
- Implicit question driving the editing experiment.
- Interpretive claim that the subcomponents correspond to real functional units.
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Specific discovered subcomponent that activates on punctuation like ' :', ' ;', ' =', ':-' and predicts the rest of emoticons/emojis.
- CLIP training paradigm finding in cross-modal alignment