claim
active
claim:the-misalignment-persona-causes-capability-reductions-particularly-on-benchmarks-with-factual-knowledge-recall-componentsThe misalignment persona causes capability reductions particularly on benchmarks with factual knowledge recall components
Specific exception to capability preservation claim, attributed to misalignment constitution explicitly encouraging subtly incorrect answers
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Proposed mechanism for collapse distinct from reweighting: representation bleeding rather than archetype selection
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Motivation for the two-stage training design; links the model organism to plausible natural emergence.
- Extended experimentation proposed to clarify the extent of the findings
- Central critique of prior evaluation: whole-response scoring hides individual OOC sentences
- Mechanistic investigation proposed to directly test persona-model collapse at the representation level
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles