concept
active
concept:contrastive-system-prompt-completionsContrastive system prompt completions
Training method for probes: generate completions under opposing system prompts to induce positive and negative poles of a concept
Neighborhood — ranked by edge-count
Methods (1)
method
- Contrastive mean-difference probeimplementsProbe construction method: concept vector at each layer is L2-normalized difference between mean positive and mean negative representations from contrastive system prompts
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Technique for extracting trait directions by contrasting model activations under trait-eliciting vs. trait-suppressing conditions
- Adaptation of instruction-tuned extraction to base models using third-person descriptions and hypothetical situations
- Supervised learning framework where system learns by observing contrast between current response and nudged improved response; requires weak additional forces from supervisor
- A sense of being complete and comfortable, as in the friendly house edge, that enhances life.
- Using system prompts to instruct models to adopt a persona; used as baseline comparison against character training
- Failure mode in training-free RPAs where long role-related context dilutes intended persona signal
- LAT methodology step constructing paired prompts that elicit divergent behaviors to extract steering vectors
- The property that living structures contain intense contrast—far more than one imagines helpful; true opposites which annihilate each other when superimposed, creating differentiation that gives birth to something; contrast unifies rather than separates when used correctly