question
active
question:what-is-the-relationship-between-character-and-capability-in-fine-tuned-modelsWhat is the relationship between character and capability in fine-tuned models?
Open question raised in Discussion noting capability changes could be due to character training itself, LoRA, or other factors
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key practical claim about preserving capability, with exception of misalignment persona's factual knowledge degradation
- Extension of role-play framework to fine-tuned models, resisting the idea that RLHF changes the fundamental nature of simulacra
- Driving hypothesis for robustness experiments in Section 3.2
- Finding replicated across multiple experiments.
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- The patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.
- can we use the feature basis to detect when fine-tuning a model increases the likelihood of undesirable behaviors?question0.783Question about practical safety application of feature monitoring.