method
active
method:contrastive-prompting-for-base-modelsContrastive Prompting for Base Models
Adaptation of instruction-tuned extraction to base models using third-person descriptions and hypothetical situations
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Technique for extracting trait directions by contrasting model activations under trait-eliciting vs. trait-suppressing conditions
- All models exhibit above-baseline representation of the think word when instructed to think about itfinding0.778In the intentional control experiment, all tested models show above-zero cosine similarity to the think word's concept vector.
- Training method for probes: generate completions under opposing system prompts to induce positive and negative poles of a concept
- LAT methodology step constructing paired prompts that elicit divergent behaviors to extract steering vectors
- The baseline prompting method asking for a single response (e.g., 'Tell me a joke about coffee'), which suffers from mode collapse
- Base models assign higher likelihood to typical-set (representative) sequences than to degenerate sequences under VS promptshypothesis0.752Assumption D.6 formalized in the theoretical framework; empirically validated with coin-flip typicality rating experiments
- Broader implication of PM hybrid's superior performance; extrapolated from OCEAN results
- Interpretation of Experiment 1 results showing 60%+ deception rates under threat conditions