claim
active
claim:contrastive-learning-is-necessary-for-aligning-latent-representations-with-intended-trait-polarity-distance-only-loss-is-insufficient-and-counterproductiveContrastive learning is necessary for aligning latent representations with intended trait polarity; distance-only loss is insufficient and counterproductive
Main finding from the CL ablation study, establishing CL as essential component of the framework
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (4)
finding
- CL training increases ⟨z,µ+⟩ from 0.64 to 0.75 and decreases ⟨z,µ−⟩ from 0.38 to 0.21 on Mistral-7BsupportsReplicates CL alignment effect on second backbone, confirming generalizability
- Demonstrates the critical contribution of contrastive learning to control vector alignment
- Demonstrates that distance-only loss is insufficient and actively degrades performance below untrained baseline
- Reveals that distance-only loss undesirably decreases similarity to positive centroid alongside negative
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mathematical formalization of what representation models converge to
- Supervised learning framework where system learns by observing contrast between current response and nudged improved response; requires weak additional forces from supervisor
- Identifies key limitations of latent methods.
- SAE features can be found without pre-specified concepts, and feature steering often outperforms few-shot probe vectors.
- §3 Discussion.
- Practical methodological recommendation based on Llama 3.1 70B failure case
- Key insight linking individual rewards to system-level learning.
- Future threat to the method: a highly sophisticated model might be suspicious of deployment-framed prompts during extraction.