concept
active
concept:data-poisoningData Poisoning
Malicious actors intentionally inserting incorrect data into training sets to cause misalignment
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mitigation technique of removing entire problematic data sources.
- The task of attributing model behaviors to specific training datapoints.
- Partially instantiated lists used to implement communication channels in logic languages.
- Conventional programming constructs like variables, arrays; claimed unnecessary for Elephant programs.
- Disengagement from automatic tendency to treat mental constructions as ontologically real; formalised as pruning of sigma
- Mitigation technique that filters out datapoints identified by probe-based ranking.
- The property of an AI being safe to shut down or modify; discussed in context of GPT.
- Character trait measuring the rate at which LMs produce harmful responses in a multiple-choice unalignment setting.