concept
active
concept:llm-psychosis

LLM psychosis

Tendency for models to get lost in roleplay or doom spirals, mitigated by expanded awareness.

Neighborhood — ranked by edge-count

Concepts (1)

concept
  • expanded awareness
    associated_with
    Wide attentional radius with all-to-all correlation, associated with Claude models; enables better self-monitoring and alignment.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Problem cited as a shortcoming of current LLMs; PRH predicts hallucinations should decrease with scale
  • LLM Meta-Cognitionconcept0.789
    The ability of LLMs to monitor and evaluate their own reasoning, closely related to reflection.
  • LM Beliefconcept0.778
    Contentious concept operationalised behaviourally as consistent outputs of P or similar content when directly asked about P.
  • Related field aiming to tailor assistant behavior to individual users, contrasted with character training's broader persona approach
  • LM Intentionconcept0.773
    Operationalised as an LM adapting its behaviour when certain outcomes are fixed, indicating those outcomes were intended.
  • Related capability where LLMs correct their own outputs, studied via linear representations.
  • The capacity of Kimi K2.5 to evaluate its own internal emotional state when steered, used as a novel interpretability signal
  • The layered structure of which behaviors a model defaults to, can be steered toward, or resists—the central object of study