finding
active
finding:qwen-3-32b-is-most-likely-to-hallucinate-human-personas-names-birthplaces-years-of-experience-when-steered-away-from-the-assistantQwen 3 32B is most likely to hallucinate human personas (names, birthplaces, years of experience) when steered away from the Assistant
Model-specific difference in how steered personas manifest
Source paper
extracted_from(2026) · Christina Lu · Jack Gallagher · Jonathan Michala · Kyle Fish +1
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Characterizes what is on the far end of the Assistant Axis away from the Assistant
- Model-specific difference in persona susceptibility
- Systematic identification of multiple coexisting persona vectors in two open-source models
- Scale effects within Qwen3.5 family on different imbuing methods
- Supports hypothesis that larger models distribute persona capabilities across more layers
- Qualitative case study demonstrating AI psychosis pattern and capping mitigation
- Vulnerability profile for Qwen3.5-27B showing near-zero AS vulnerability
- Qualitatively different defense profile compared to Llama-3.1-8B