finding
active
finding:ifeval-instruction-score-fails-to-detect-early-breakdowns-in-coherency-degradation-with-coherency-collapsing-at-smaller-steering-magnitudes-than-ifeval-declineIFEval instruction score fails to detect early breakdowns in coherency degradation, with coherency collapsing at smaller steering magnitudes than IFEval decline
Demonstrates inadequacy of IFEval as proxy for coherency in activation steering evaluation
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Motivates adopting coherency score as the primary evaluation metric
- Key empirical observation motivating analysis of positive vs negative steering directions
- Demonstrates inadequacy of perplexity as a proxy for coherency in activation steering evaluation
- We hypothesize that coherency degradation stems from residual stream intervention that indiscriminately amplifies off-target noisehypothesis0.735Core mechanistic hypothesis motivating the shift from residual stream to head-level steering
- DeepSeek has large coherence loss but no robustness excess; GPT-4o has little coherence loss but large robustness drop
- Future threat to the method: a highly sophisticated model might be suspicious of deployment-framed prompts during extraction.
- Rapid breakdown of output quality under strong or out-of-distribution activation steering, the main problem this paper addresses
- Only Criterion 2 is satisfied for this single case at the task level (granularity without aggregation).