concept
active
concept:data-poisoning

Data Poisoning

Malicious actors intentionally inserting incorrect data into training sets to cause misalignment

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Mitigation technique of removing entire problematic data sources.
  • Data Attributionconcept0.754
    The task of attributing model behaviors to specific training datapoints.
  • data streamsconcept0.750
    Partially instantiated lists used to implement communication channels in logic languages.
  • data structuresconcept0.731
    Conventional programming constructs like variables, arrays; claimed unnecessary for Elephant programs.
  • Dereificationconcept0.722
    Disengagement from automatic tendency to treat mental constructions as ontologically real; formalised as pruning of sigma
  • Mitigation technique that filters out datapoints identified by probe-based ranking.
  • Corrigibilityframework0.714
    The property of an AI being safe to shut down or modify; discussed in context of GPT.
  • Harmfulnessconcept0.711
    Character trait measuring the rate at which LMs produce harmful responses in a multiple-choice unalignment setting.