thinker:takahiro-komamizuTakahiro Komamizu
Authored papers (1)
- Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs2026
Facet-level personality control in role-playing LLMs is substantially improved by injecting contrastively trained sparse autoencoder (SAE) control vectors into mid-residual layers, with the CV-SAE+Prompt configuration achieving 88.5% Full-Accuracy on contextualized Big Five questions for both Qwen3-4B and Mistral-7B, compared to 30.7% and 42.3% respectively for prompt-only baselines. The core method, Contrastive SAE with Trait-Activated Routing, learns 30 facet-aligned control vectors from a purpose-built 15,000-sample leakage-controlled corpus (500 instances per facet, 78.4% macro-F1 on a held-out 30-way classifier), using a prototype contrast loss with ArcFace/CosFace-style angular margins to pull SAE codes toward positive trait centroids while pushing them from negative ones. A key finding is that CV-CAA+Prompt catastrophically collapsed on Mistral-7B (FA dropping from 76.9% to 38.5%), whereas CV-SAE+Prompt remained stable, demonstrating that SAE disentanglement is critical for preventing prompt-activation interference. The Agent-Based Decision Module, guided by Trait Activation Theory, selects only the facets contextually cued by each user query rather than injecting all 30 vectors simultaneously, preserving dialogue coherence as measured by Multi-Turn Rate. The paper argues this implies that facet-purity in training data, combined with contrastive geometry in latent space and targeted injection, is sufficient to achieve stable, interpretable persona control at inference time without any weight updates.
More papers — OpenAlex / S2
Affiliations (1)
- Nagoya University(institute)
Co-authors (3)
- Ichiro Ide9 shared
- Wenqiu Tang9 shared
- Zhen Wan9 shared
Other inbound relations (1)
Recent mentions (1)
- papers-typedtang-2026-facet-level.md