concept
active
concept:language-modelsLanguage Models
Primary substrate for manifold steering experiments; demonstrates method on reasoning and in-context tasks.
Neighborhood — ranked by edge-count
Concepts (3)
concept
- modelrelated_toA representation that captures relevant aspects of a system; according to the theorem, the regulator must embody this.
- Language Modelsame_asPrimary test domain for manifold steering, including reasoning and ICL tasks
- Manifold Steeringassociated_withCentral framework: steering neural networks by intervening along the curved manifold where a concept lives, rather than in straight lines through activation space.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- LM architecture with sense vectors showing multiplication effects, illustrating custom intervention in pyvene
- Training objective interpretable as optimizing a diverse set of tasks; thus subject to multitask scaling convergence pressures
- Features related to gender, racial, ethnic biases, slurs, and hate speech.
- Philosophical stance that LM psychological traits should be evaluated purely in terms of input-output behaviour.
- Paper hypothesising LLMs model agent beliefs/desires/intentions with preliminary GPT-3 evidence; cited as ref 2
- Transformer-based models like GPT-4, LaMDA, PaLM; assessed for GWT indicators.
- Framework describing LLMs as role-play engines, introduced in Shanahan, McDonell, Reynolds 2023.