framework
active
framework:behaviourism-for-language-modelsBehaviourism for Language Models
Philosophical stance that LM psychological traits should be evaluated purely in terms of input-output behaviour.
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Primary test domain for manifold steering, including reasoning and ICL tasks
- Primary substrate for manifold steering experiments; demonstrates method on reasoning and in-context tasks.
- Study showing RLHF can exacerbate self-preservation tendencies in LLMs; key empirical support for a paper claim
- Prior work studying sycophancy and desire not to be shut down in RLHF-trained models
- Features related to gender, racial, ethnic biases, slurs, and hate speech.
- Key prior finding that LLMs can internally represent beliefs of self and others, motivating SOO approach
- Training objective interpretable as optimizing a diverse set of tasks; thus subject to multitask scaling convergence pressures
- RLHF paper cited as a major fine-tuning technique used in commercial dialogue agents