finding
active
finding:intervening-at-mlp-output-has-minimal-effect-on-trait-expression-compared-to-intervening-at-the-identified-attention-layer-outputIntervening at MLP output has minimal effect on trait expression compared to intervening at the identified attention layer output
Confirms attention rather than MLP as the locus of persona generation
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Architectural rationale for why Style Modulation Heads are in attention layers rather than MLPs
- Key limitation of the paper's approach; MLP layers make up 2/3 of standard transformer parameters
- Shows persona vector filtering has complementary strengths to LLM judges, surfacing non-obvious problematic samples
- Extension of the thesis to deployed LLM inference via in-context learning
- Enables pre-finetuning data screening; Figure 8 shows strong dataset-level correlations across all three traits
- Suggests LLMs do not represent complement/MSV linguistic features in the same way as they are crucial for human ToM development.
- Hypothesis based on observed negative cosine similarity between input and output weights of some neurons
- Hypothesis raised in distributive law task analysis