finding
active
finding:base-rpa-achieves-only-7-7-fa-on-qwen3-4b-and-19-2-on-mistral-7b-abstract-questionsBase RPA achieves only 7.7% FA on Qwen3-4B and 19.2% on Mistral-7B abstract questions
Establishes low-bar baseline showing personality control without intervention is poor
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Scale effects within Qwen3.5 family on different imbuing methods
- Smaller models show non-monotonic and diminished ASR with increasing cone dimensionality
- Shows that explicit labels without latent steering fail to generalize to situational cues
- QwQ-32B reaches 15.2% overall ASR (23.3% SP, 7.3% FS) under prompt-based persona assignment.finding0.752Reasoning model vulnerability under prompting
- Vulnerability profile for Qwen3.5-27B showing near-zero AS vulnerability
- Reasoning model vulnerability under prompting
- Quantifies harness activation failure for weak-tier models vs. strong-tier models
- Control comparison showing near-linearity is specific to the targeted manipulation direction