concept
active
concept:lm-beliefLM Belief
Contentious concept operationalised behaviourally as consistent outputs of P or similar content when directly asked about P.
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Logical Coherence of LM Beliefsrelated_toThe extent to which an LM's beliefs respect logical entailment: if it believes A and A→B, it should also believe B.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Operationalised as an LM adapting its behaviour when certain outcomes are fixed, indicating those outcomes were intended.
- The ability of LLMs to monitor and evaluate their own reasoning, closely related to reflection.
- Tendency for models to get lost in roleplay or doom spirals, mitigated by expanded awareness.
- Conclusion from Experiment 2 on Leap-of-Thought.
- Related capability where LLMs correct their own outputs, studied via linear representations.
- Alternative data attribution approach using an LLM as a judge; compared against the probe-based method.
- Incorrect abduction arising from chance occurrences consistent with prior beliefs; leads to persistently poor behavior
- Evaluation framework using an LLM (GPT-4.1-mini) to score trait expression and coherency