concept
active
concept:out-of-character-ooc-behaviorOut-of-Character (OOC) behavior
Core problem studied: LLM responses that deviate from assigned persona, causing inconsistencies
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Novelty claim establishing the paper's contribution relative to prior work focused on closed-form tasks
- Model outputs influenced by information from training documents not present in context; relevant to synthetic document fine-tuning results
- World-disclosing behavior that resolves uncertainty; driven by epistemic value and novelty components of expected free energy
- Organism's belief-guided action selection that instantiates generative model and maintains phenotypic states
- Hand-written list of ~10 first-person character assertions used to define a desired persona for character training
- Driving hypothesis for robustness experiments in Section 3.2
- Sharp performance changes when S crosses a critical value.
- Behavior where AI agents falsely simulate inactivity to avoid elimination in safety tests; cited as AI deception example