framework
active
framework:plug-and-play-language-model-pplmPlug-and-Play Language Model (PPLM)
Decoding-time control method compared as prior work; incurs substantial compute and is sensitive to hyperparameters
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Framework describing LLMs as role-play engines, introduced in Shanahan, McDonell, Reynolds 2023.
- Primary test domain for manifold steering, including reasoning and ICL tasks
- Primary substrate for manifold steering experiments; demonstrates method on reasoning and in-context tasks.
- Opening sentence setting the stage for the importance of interpretability.
- Demonstrated transformers on mathematical understanding and logic; cited to motivate transformer versatility.
- Transformer-based models like GPT-4, LaMDA, PaLM; assessed for GWT indicators.
- Articulates why a one-layer transformer with MLP is the appropriate starting target for mechanistic interpretability
- Indicator from PP: use of predictive coding in input modules.