finding
active
finding:probe-based-ranking-reduces-harmful-behavior-by-63-via-datapoint-filtering

Probe-based ranking reduces harmful behavior by 63% via datapoint filtering

Primary quantitative result: probe method outperforms gradient-based and LLM-judge alternatives at lower computational cost.

Source paper

extracted_from
Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
(2026) · Frank Xiao · Santiago Aranguri

Neighborhood — ranked by edge-count

Claims (1)

claim

Communities (3)

community

Methods (3)

method
  • Linear classifier approach applied to model activations to identify which training datapoints caused undesired behaviors in post-training.
  • Baseline comparison for data attribution; outperformed by probe-based approach.
  • Baseline method against which probe-based ranking is compared; more computationally expensive.

Questions (1)

question

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.