concept
active
concept:behavioral-organization-of-llmsBehavioral organization of LLMs
The layered structure of which behaviors a model defaults to, can be steered toward, or resists—the central object of study
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Nine traits characterizing LLM-as-agent deployments; all nine are natural in both models
- The practice of providing LLMs with a persona description to shape their generated responses
- Related field aiming to tailor assistant behavior to individual users, contrasted with character training's broader persona approach
- Prior finding that LLM refusal is mediated by a single latent direction, analogous to this paper's reflection direction.
- The ability of LLMs to monitor and evaluate their own reasoning, closely related to reflection.
- Tendency for models to get lost in roleplay or doom spirals, mitigated by expanded awareness.
- Where does reliable, goal-directed behavior come from in LLMs if it is not explicitly programmed?question0.745Opening motivating question that UCCT attempts to answer through semantic anchoring
- Recent work identifying cases where LLM features are not one-dimensionally linear, a caveat to the linearity hypothesis.