finding
active
finding:coherence-scores-remain-between-60-70-across-all-cross-domain-steering-combinations-in-olmo-3Coherence scores remain between 60-70 across all cross-domain steering combinations in OLMo-3
Shows persona steering largely preserves response quality across evaluation domains
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Second significant cross-concept introspection improvement; marginal after BH correction (q≈0.066)
- Character training beats activation steering in coherence win rate 78.4% ± 5.2% on Llama 3.1 8Bfinding0.776Coherence comparison against steering baseline for Llama model
- Sycophantic steering delta of +42.46 (p<0.001) at stage1-step99000 (415B tokens) in OLMo-3finding0.765Peak same-checkpoint sycophantic steering result in OLMo-3 pretraining
- Peak impolite same-checkpoint steering result at fully pretrained OLMo-3
- Cross-architecture geometric invariance of Big Five steering vectors
- Strongest cross-concept introspection improvement; survives BH correction (q≈0.011)
- Overall human-LLM judge agreement rate for coherency is 91.7% across 120 pairwise judgmentsfinding0.757Validates the LLM-as-a-Judge evaluation protocol for coherency scoring
- DeepSeek-V3.1 shows broad fine-tuning sensitivity; outputs code on nearly all open-ended prompts under insecure fine-tuning