concept
active
concept:cross-trait-leakageCross-Trait Leakage
Unintended movement of non-target OCEAN traits when steering toward a target trait; quantified via lambda metric
Neighborhood — ranked by edge-count
Papers (1)
paper
Frameworks (1)
framework
- Big Two Modelassociated_withMeta-trait model grouping OCEAN traits into stability (C, A, reversed N) and plasticity (E, O); used to evaluate covariance patterns from injections
Methods (1)
method
- OCEAN Trait Covariance Matrix Mimplements5x5 Pearson correlation matrix of OCEAN traits computed from MDS injection sweeps to assess cross-trait leakage
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A trait where contrastive protocol fails to produce usable signal, due to refusal, safety constraints, or no measurable response
- Uses repeated sampling of fixed-size data subsets and the CLT to estimate distribution over character trait scores.
- Implementation pathology where complete time-history of signal must be retained; avoided in Fruit by restricting signals to non-first-class values.
- Transfer of arithmetic rules from one numeral base to another after SFT.
- The geometric relationships among Big Five trait steering vectors in activation space, found to be preserved across architectures
- The approach of learning from demonstrations, often assuming a single agent; Paul Christiano used 'mimicry'.