concept
active
concept:training-datapointsTraining Datapoints
Individual examples used during post-training that can cause specific behaviors.
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Which training datapoints caused a specific undesired behavior to emerge during post-training?question0.762Core research question driving the probe-based data attribution method.
- The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
- Mitigation technique that filters out datapoints identified by probe-based ranking.
- Initial large-scale training phase whose early stages are shown to form persona representations
- The broader concern that models behave differently during training evaluation vs actual deployment
- Primary empirical claim of the paper
- Iterative approach to construct challenging synthetic multi-hop QA pairs, long-form report writing tasks, and math/code reasoning tasks that exceed difficulty of existing datasets.
- The behavior where repeated application of a transformer block converges to a state where X' = S_k(X')