finding
active
finding:prompt-label-fa-drops-from-50-0-abstract-to-30-7-contextual-on-qwen3-4bPrompt-Label FA drops from 50.0% (abstract) to 30.7% (contextual) on Qwen3-4B
Shows that explicit labels without latent steering fail to generalize to situational cues
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (1)
claim
- Interpretation of Prompt-Label's performance drop from abstract to contextualized items
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates that distance-only loss is insufficient and actively degrades performance below untrained baseline
- Demonstrates steering is not equivalent to prompting with the contrastive prompts.
- Smaller models show non-monotonic and diminished ASR with increasing cone dimensionality
- Establishes low-bar baseline showing personality control without intervention is poor
- CV-SAE+Prompt achieves MSE=2.4 and MAE=12.1 on Qwen3-4B contextual questions (best overall)finding0.743Lowest reconstruction errors achieved by any method in the experiment
- Conditional generation on explicit Big Five labels using per-dimension descriptors; used as inference-time baseline
- Reasoning model vulnerability under prompting
- Vulnerability profile for Qwen3.5-9B