question
active
question:can-even-finer-grained-extraction-and-steering-be-achieved-beyond-the-attention-head-level-through-integration-of-saesCan even finer-grained extraction and steering be achieved beyond the attention head level through integration of SAEs?
Open question about whether SAE integration can provide sub-head level steering granularity
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows gating effect is specific to the self-referential computational regime, not a general feature effect
- Extension of mechanistic interpretability findings to the metacognitive domain
- Validation finding from Lu et al. 2026 supporting PSM's claim about pre-training persona structure
- Out-of-distribution generalization of SAE features.
- Key asymmetry finding: suppressing reflection is easier than inducing it.
- Supported by the instruction discovery experiments comparing steering vs. embedding baselines.
- Future work hypothesis about extending SOO to direct value alignment
- Core applied contribution claim, supported by top-k accuracy comparisons.