finding
active
finding:qwen-2-5-7b-achieves-100-asr-across-all-cone-dimensions-1-5Qwen-2.5-7B achieves 100% ASR across all cone dimensions 1–5
Experiment 2 result showing large models can support high-dimensional truth cones
Source paper
extracted_from(2025) · Kevin Shengyang Yu · Vaidehi Bulusu · Oscar Yasunaga · Lau, Clayton +4
Neighborhood — ranked by edge-count
Claims (2)
claim
- Truthful behavior in LLMs is not confined to a single linear axis; multiple orthogonal directions can independently mediate itassociated_withsupportsCentral interpretive claim of the paper
- Interpretation of ASR degradation patterns by model size across cone dimensions
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Experiment 2 result showing large Gemma model supports high-dimensional truth cones
- Vulnerability profile for Qwen3.5-27B showing near-zero AS vulnerability
- Smaller models show non-monotonic and diminished ASR with increasing cone dimensionality
- Vulnerability profile for Qwen3.5-9B
- Contrast with DeepSeek-R1 showing QwQ is more robust to geometric steering
- Qualitatively different defense profile compared to Llama-3.1-8B
- QwQ-32B reaches 15.2% overall ASR (23.3% SP, 7.3% FS) under prompt-based persona assignment.finding0.755Reasoning model vulnerability under prompting
- On Qwen3-1.7B, MDS achieves ϕ1,C,↑ = 5.0 (SJTs) vs P2 at 4.7, and ϕ1,C,↓ = 1.4 (SJTs) vs P2 at 3.6finding0.752Specific consciousness sweep result for Qwen3-1.7B from Table 6 demonstrating strong bidirectional steering