method
active
method:steered-cross-entropy-loss-predictionSteered Cross-Entropy Loss Prediction
Measuring whether artificially activating a latent reduces cross-entropy loss on a fine-tuning dataset as a proxy for dataset correctness
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Steered loss can identify whether a dataset is likely to lead to misalignment
- Demonstrates that persona vectors capture trait-specific signal beyond general misalignment signal
- Evidence of a bottleneck between richer internal variation and final report distribution in impulsivity→interest condition
- Empirical demonstration on Llama-3.1-8B that steering along representation manifold aligns outputs with behavior manifold, whereas linear steering does not.
- Evidence that improved introspection in focus→wellbeing arises from enriched internal state and report channels simultaneously
- Core empirical claim comparing steering approaches on cyclic concepts.
- Validates using chain-of-thought belief monitoring as proxy for behavioral steering efficacy.
- Finding that the two evaluation modalities frequently diverge in their interpretation of the same SAE feature