question
active
question:can-we-use-the-feature-basis-to-detect-when-fine-tuning-a-model-increases-the-likelihood-of-undesirable-behaviorscan we use the feature basis to detect when fine-tuning a model increases the likelihood of undesirable behaviors?
Question about practical safety application of feature monitoring.
Source paper
extracted_fromRelated by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Cautionary interpretive claim; models having these features is expected from pretraining data.
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- Open question raised in Discussion noting capability changes could be due to character training itself, LoRA, or other factors
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Policy recommendation based on latent #10 activating at 5% incorrect data before behavioral misalignment emerges
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment