concept
active
concept:concept-injection

Concept Injection

Technique of injecting activation patterns associated with specific concepts into a model's internal states to test whether self-reports reflect ground truth.

Neighborhood — ranked by edge-count

Frameworks (1)

framework

Communities (1)

community

Claims (1)

claim

Methods (6)

method
  • Causal intervention technique: edit NLA explanation, reconstruct via AR, use difference as steering vector to manipulate model behavior.
  • Task where the model must simultaneously identify an injected thought and transcribe a text sentence.
  • Experimental paradigm where the model is told about the possibility of thought injection and asked to report detection and identification.
  • Task instructing the model to write a sentence while thinking or not thinking about a word, measuring internal representation strength.
  • Task where a random word is prefilled as the assistant's response, then the model is asked whether it intended to say that word, testing introspection on prior intentions.
  • Technique for obtaining concept vectors by presenting model with two scenarios differing in one respect and subtracting activations to isolate conceptual difference.

Concepts (1)

concept
  • The central concept: the ability of a model to access and report on its internal states, as defined by the paper's criteria.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.