claim
active
claim:a-lora-adapter-can-learn-to-add-a-constant-vector-equivalent-to-steering-along-a-pre-existing-persona-direction-making-fine-tuning-literally-equivalent-to-activation-steeringA LoRA adapter can learn to add a constant vector equivalent to steering along a pre-existing persona direction, making fine-tuning literally equivalent to activation steering
Mechanistic explanation of how fine-tuning can shift persona vectors without directly updating activations
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Adaptation method used via Tinker API for DeepSeek-V3.1 and Qwen3-235B fine-tuning with rank 32
- Practical disadvantage of activation steering highlighted as a drawback
- Limitation acknowledgment about the adequacy of the linear representation assumption
- Core empirical result showing persona vectors capture trait-specific signal mediating finetuning-induced persona shifts
- Future work hypothesis about extending SOO to direct value alignment
- Cited from Wang et al. 2025a as reason SDF is preferred over demonstration fine-tuning for realistic model organisms.
- Characterizes the trait content of the Assistant Axis in pre-trained models
- Key interpretive conclusion from the dissociation between attempt rate and improvement rate in fine-tuning experiments