question
active
question:which-training-datapoints-caused-a-specific-undesired-behavior-to-emerge-during-post-trainingWhich training datapoints caused a specific undesired behavior to emerge during post-training?
Core research question driving the probe-based data attribution method.
Source paper
extracted_from(2026) · Frank Xiao · Santiago Aranguri
Neighborhood — ranked by edge-count
Findings (4)
finding
- Primary quantitative result: probe method outperforms gradient-based and LLM-judge alternatives at lower computational cost.
- Key empirical result: swapping labels of datapoints flagged by probes yields a 78% reduction.
- Discovery of the emergence of harmful compliance under specific post-training conditions (DPO + formatting constraints).
- Key empirical result: removing four identified problematic data sources yields an 84% reduction.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Finding that base models have high false positives and no net positive performance.
- Base pretrained models show high false positive rates and achieve no net task performance on concept injection detection; post-training essential for introspection.
- The training-time transition where incorrect-solution minima lose stability and become saddles as the model gains solving ability.
- DAS finds causal effect at all training timesteps including when model is just initialisedfinding0.763Corroborates Wu et al. 2023 finding that DAS expressivity inflates causal effect estimates
- Central interpretive claim and motivation for future work
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Individual examples used during post-training that can cause specific behaviors.
- Authors' interpretive endorsement of PSM view, backed by transfer experiments