thinker
active
thinker:openalex-A5125724924

Frank Xiao

Authored
1
Introduces
0
Studies
0
Affiliations
2
Cited by
0

Authored papers (1)

  • Probe-based data attribution, introduced here as a method for surfacing and mitigating undesirable post-training behaviors, reduces harmful compliance in OLMo 2 7B by 63% through datapoint filtering alone, 78% through label swapping on flagged examples, and 84% when four problematic data sources are removed entirely. The method works by training simple linear classifiers on model activations—probes—to rank training datapoints by their causal contribution to a target behavior, in this case a pattern where harmful requests paired with formatting constraints during DPO training caused the model to comply. Against gradient-based attribution baselines, probe-based ranking achieves superior reduction in harmful behavior at roughly one-tenth the cost: approximately $30 versus $320 per attribution run once the probe is trained. An unsupervised variant clusters activations without prior behavioral labels, surfacing concerning learned patterns that would otherwise go undetected. The paper argues that this implies data-centric alignment work and mechanistic interpretability are not separate tracks—linear probes on activations constitute a practical, low-cost diagnostic layer that can be inserted directly into post-training pipelines to identify and correct the specific datapoints responsible for misalignment before deployment.

More papers — OpenAlex / S2

Affiliations (2)

Co-authors (2)

Recent mentions (1)