finding
active
finding:cl-training-increases-z-from-0-64-to-0-75-and-decreases-z-from-0-38-to-0-21-on-mistral-7bCL training increases ⟨z,µ+⟩ from 0.64 to 0.75 and decreases ⟨z,µ−⟩ from 0.38 to 0.21 on Mistral-7B
Replicates CL alignment effect on second backbone, confirming generalizability
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (1)
claim
- Main finding from the CL ablation study, establishing CL as essential component of the framework
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reveals that distance-only loss undesirably decreases similarity to positive centroid alongside negative
- Mistral-7B-Instruct-v0.2 deceptive response rate reduced from 73.6% to 17.27% ± 1.88% after SOO fine-tuningfinding0.768Primary result showing SOO fine-tuning significantly reduces deception in Mistral-7B
- Shows behavioral pattern of self-correction is trainable in smaller models
- Modified CL loss achieves IIA of 0.9988±0.0005 on synthetic 10-class dataset training/test setsfinding0.748IIA for modified CL loss on synthetic dataset, comparable to behavioral DAS
- Mistral-7B MT-Bench score minimally changed from 7.26 to 7.3 ± 0.06 after SOO fine-tuningfinding0.747SOO fine-tuning had negligible impact on Mistral-7B general capabilities
- SOO fine-tuning reduced the MSE between self and other activations in Mistral-7B MLP layers
- Mistral-7B average generalization deceptive rate reduced from 56.74% ± 14.73% to 12.40% ± 12.06%finding0.737SOO fine-tuning generalized across 7 scenario variants for Mistral-7B
- Layer-by-layer analysis of refusal direction properties