method
active
method:cosine-similarity-between-truth-probesCosine similarity between truth probes
Geometric evaluation of truth direction alignment across layers and prompt templates.
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Detection mechanism computing cosine similarity between activation vectors and steering vectors to classify deception
- Used to measure alignment between DIM direction and cone basis vectors to assess overlap
- Geometric measure used to track persona vector drift and stability across checkpoints
- Cosine similarity between feature activations restricted to tokens where one of the features fires; used to identify feature splitting relationships
- Supported by the geometric transition visible in cosine similarity heatmaps for F0-F3.
- Used to quantify the semantic clustering of adjective-set embeddings across model families and conditions
- Method to discover new reflection-inducing instructions by ranking candidate tokens by cosine similarity to steering vectors.
- Pairwise cosine similarities between Description, Narration, and Dialogue evil vectors are all below 0.5finding0.780Shows different elicitation strategies recover qualitatively distinct persona directions