concept
active
concept:llm-meta-cognition

LLM Meta-Cognition

The ability of LLMs to monitor and evaluate their own reasoning, closely related to reflection.

Neighborhood — ranked by edge-count

Concepts (1)

concept
  • Reflection in LLMs
    associated_with
    The core phenomenon studied: the ability of LLMs to evaluate and revise their own reasoning.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • LM Beliefconcept0.795
    Contentious concept operationalised behaviourally as consistent outputs of P or similar content when directly asked about P.
  • LM Intentionconcept0.792
    Operationalised as an LM adapting its behaviour when certain outcomes are fixed, indicating those outcomes were intended.
  • LLM psychosisconcept0.789
    Tendency for models to get lost in roleplay or doom spirals, mitigated by expanded awareness.
  • Related capability where LLMs correct their own outputs, studied via linear representations.
  • The finding that interpretable concepts including character traits are encoded as linear directions in transformer residual streams
  • The extent to which an LM's beliefs respect logical entailment: if it believes A and A→B, it should also believe B.
  • Problem cited as a shortcoming of current LLMs; PRH predicts hallucinations should decrease with scale
  • Alternative data attribution approach using an LLM as a judge; compared against the probe-based method.