method
active
method:auprc-latent-activation-classifierAUPRC Latent Activation Classifier
Using per-prompt average SAE latent activations and area under precision-recall curve to discriminate aligned from misaligned models
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The pattern of which misalignment-relevant SAE latents become active for a given fine-tuned model
- Quantifies latent #10's strong discriminative power for persona jailbreaks
- Performance metric for binary classification; used to evaluate pathogenicity prediction.
- Internal representations of the model on which probes operate; the method uses activations to rank datapoints.
- Demonstrates practical utility of SAE-based monitoring even with minimal sampling
- The conventional approach (e.g., SAEs, transcoders) of decomposing activations into interpretable features.
- Supervised method training models to answer questions about activations; NLAs differ by being unsupervised.
- Method of optimizing activation-space interventions to produce behavioral paths along M_y, then measuring whether the resulting activation trajectories trace M_h curvature