claim
active
claim:sae-based-feature-monitoring-can-serve-as-an-early-warning-system-for-detecting-model-misbehavior-before-it-manifests-in-behavioral-evaluationsSAE-based feature monitoring can serve as an early warning system for detecting model misbehavior before it manifests in behavioral evaluations
Policy recommendation based on latent #10 activating at 5% incorrect data before behavioral misalignment emerges
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Extension of mechanistic interpretability findings to the metacognitive domain
- Claim that feature grounding enables interpretability metrics.
- Out-of-distribution generalization of SAE features.
- A promising property for interpretability analysis off-distribution.
- Shows gating effect is specific to the self-referential computational regime, not a general feature effect
- Surprising finding that the two evaluation methods diverge in their relationship with persistence
- can we use the feature basis to detect when fine-tuning a model increases the likelihood of undesirable behaviors?question0.781Question about practical safety application of feature monitoring.
- SAEs uncover safety-relevant representations that might be monitored or controlled.