finding
active
finding:ppl-showed-no-clear-correlation-with-coherency-and-failed-to-predict-text-quality-degradation-under-steering-in-qwen2-5-7bPPL showed no clear correlation with coherency and failed to predict text quality degradation under steering in Qwen2.5-7B
Demonstrates inadequacy of perplexity as a proxy for coherency in activation steering evaluation
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Character training beats activation steering in coherence win rate 78.4% ± 5.2% on Llama 3.1 8Bfinding0.759Coherence comparison against steering baseline for Llama model
- Validates using chain-of-thought belief monitoring as proxy for behavioral steering efficacy.
- Finding that the two evaluation modalities frequently diverge in their interpretation of the same SAE feature
- Key empirical observation motivating analysis of positive vs negative steering directions
- Quantitative summary of Head Cor superiority on the Right-normalized Constrained Envelope Area metric
- Evidence of a bottleneck between richer internal variation and final report distribution in impulsivity→interest condition
- Supporting finding for the trait refusal alignment framework
- Motivates adopting coherency score as the primary evaluation metric