concept
active
concept:llm-self-correction

LLM Self-Correction

Related capability where LLMs correct their own outputs, studied via linear representations.

Neighborhood — ranked by edge-count

Concepts (1)

concept
  • The central phenomenon introduced by this paper: inference-time recovery from irrelevant activation steering in LLMs

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Self-Correctionconcept0.841
    Reasoning pattern where model reverses a compliance tendency and returns toward refusal, associated with better safety in QwQ-32B
  • Related field aiming to tailor assistant behavior to individual users, contrasted with character training's broader persona approach
  • The capacity of Kimi K2.5 to evaluate its own internal emotional state when steered, used as a novel interpretability signal
  • Reflection in LLMsconcept0.785
    The core phenomenon studied: the ability of LLMs to evaluate and revise their own reasoning.
  • Framework by Lee et al. explaining self-correction via linear latent concept directions, closely related prior work.
  • LLM Meta-Cognitionconcept0.780
    The ability of LLMs to monitor and evaluate their own reasoning, closely related to reflection.
  • LM Intentionconcept0.779
    Operationalised as an LM adapting its behaviour when certain outcomes are fixed, indicating those outcomes were intended.
  • Technique using internal model representations as feedback loops to steer diffusion-based materials generation toward target properties.