concept
active
concept:latent-activation-signatureLatent Activation Signature
The pattern of which misalignment-relevant SAE latents become active for a given fine-tuned model
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Internal representations of the model on which probes operate; the method uses activations to rank datapoints.
- Using per-prompt average SAE latent activations and area under precision-recall curve to discriminate aligned from misaligned models
- Intervention method that adds a learned direction vector to residual stream activations to steer model behavior
- Statistical regularities stored in pretrained models.
- Large activation magnitudes in the residual stream that Queipo-de Llano et al. link to causing compression behavior and stages of inference
- Technique of reading out model beliefs from internal activations before the final answer token is generated
- Model-independent feature comparison based on correlating activation vectors across a fixed diverse dataset
- Latent model activations when processing inputs framed from another agent's perspective