finding
active
finding:predicting-dataset-correctness-from-change-in-steered-cross-entropy-loss-yields-auprc-0-91-when-averaging-over-6000-promptsPredicting dataset correctness from change in steered cross-entropy loss yields AUPRC = 0.91 when averaging over 6000 prompts
Steered loss can identify whether a dataset is likely to lead to misalignment
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Measuring whether artificially activating a latent reduces cross-entropy loss on a fine-tuning dataset as a proxy for dataset correctness
- Empirical result: covariance pooling combined with unsupervised autoencoder embeddings improves Gene Ontology prediction AUC by 8.4% over mean pooling.
- Table 2, row 3, showing equivalence when prior preferences match rewards.
- Automated logit weight prediction achieves 74% mean accuracy for features vs 58% for neurons vs 50% chancefinding0.752Automated interpretability of logit weights confirms feature downstream effects are more interpretable than neuron effects
- Shows that loss recovery increases with autoencoder size
- Demonstrates that persona vectors capture trait-specific signal beyond general misalignment signal
- Evidence that improved introspection in focus→wellbeing arises from enriched internal state and report channels simultaneously
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment