concept
active
concept:probe-based-data-attribution-for-alignmentProbe-based data attribution for alignment
Neighborhood — ranked by edge-count
Papers (1)
paper
Communities (1)
community
- Neural Steering Methodsmembers_of
Institutes (1)
institute
- GoodfireusesAI research company; authors' affiliation; develops tools including EVEE and publishes research on genomic foundation models.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Linear classifier approach applied to model activations to identify which training datapoints caused undesired behaviors in post-training.
- Conceptual framing: integrates mechanistic interpretability tools with alignment-focused data curation.
- Authors' central interpretive assertion that their method meaningfully mitigates unwanted behaviors.
- Probe-based method bridges interpretability (probes/activations) with data-centric alignment workclaim0.830Assertion from the paper's notes that the work connects two previously separate areas: interpretability tools and data-centric alignment.
- Baseline method against which probe-based ranking is compared; more computationally expensive.
- The task of attributing model behaviors to specific training datapoints.
- Correlating attribution vectors (feature activation × logit weight of next token) across model pairs to measure functional universality
- Open methodological question acknowledged as limitation