method
active
method:positive-pull-loss-l-distPositive Pull Loss (L_dist)
Distance-based loss comparing injected representation to class centroids in the active subspace
Neighborhood — ranked by edge-count
Papers (1)
paper
Methods (1)
method
- Procedure mapping hidden representations into SAE space and applying contrastive loss to learn facet-aligned control vectors
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Regularization component of the composite loss that penalizes deviation from baseline model behavior on Alpaca instructions
- Loss function pulling representations toward positive centroid and pushing away from negative centroid with angular margins
- Comparison of loss-scale balancing with IMTL-L.
- Auxiliary training objective from Grant (2025) that constrains intervened representations to remain near natural distribution
- The objective function combining L2 reconstruction error and L1 penalty scaled by decoder norm, used to train the SAE.
- Modified CL loss outperforms behavioral DAS loss in OOD transfer from dense to sparse class partitionfinding0.694Key practical utility result: CL loss improves generalization of alignment to out-of-distribution settings
- Novel variant of CL loss introduced in this paper targeting only causal subspace dimensions to improve OOD performance
- Auxiliary objective combining L2 and cosine losses against pre-recorded CL vectors to improve causal relevance when one model is causally inaccessible.