concept
active
concept:latent-activation-signature

Latent Activation Signature

The pattern of which misalignment-relevant SAE latents become active for a given fine-tuned model

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Activationsconcept0.800
    Internal representations of the model on which probes operate; the method uses activations to rank datapoints.
  • Using per-prompt average SAE latent activations and area under precision-recall curve to discriminate aligned from misaligned models
  • Intervention method that adds a learned direction vector to residual stream activations to steer model behavior
  • latent patternsconcept0.743
    Statistical regularities stored in pretrained models.
  • Large activation magnitudes in the residual stream that Queipo-de Llano et al. link to causing compression behavior and stages of inference
  • Activation Probingconcept0.738
    Technique of reading out model beliefs from internal activations before the final answer token is generated
  • Model-independent feature comparison based on correlating activation vectors across a fixed diverse dataset
  • Latent model activations when processing inputs framed from another agent's perspective