concept
active
concept:refusal-direction-in-llms

Refusal Direction in LLMs

Prior finding that LLM refusal is mediated by a single latent direction, analogous to this paper's reflection direction.

Neighborhood — ranked by edge-count

Concepts (2)

concept
  • Arditi et al. 2024 finding that refusal behavior is mediated by one direction in LLM activations; exemplar of single-direction causal results
  • The paper's central construct: a vector in LLM activation space encoding the transition between reflection levels.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Linear direction in LLM activations associated with truthfulness, identified by Burns et al. 2022 and Azaria & Mitchell 2023
  • Reflection in LLMsconcept0.775
    The core phenomenon studied: the ability of LLMs to evaluate and revise their own reasoning.
  • The practice of providing LLMs with a persona description to shape their generated responses
  • Nine traits characterizing LLM-as-agent deployments; all nine are natural in both models
  • The view that the individual is the abstract function defined by a given architecture and weight matrix
  • The layered structure of which behaviors a model defaults to, can be steered toward, or resists—the central object of study
  • Using Claude Sonnet 4 as a grader to categorize model responses according to predefined criteria.
  • A specific direction in an LLM's residual stream that encodes the truth or falsehood of factual statements