artifact
active
artifact:natural-language-autoencoder-training-code-and-modelsNatural Language Autoencoder Training Code and Models
Released on GitHub (github.com/kitft/natural_language_autoencoders) and interactive frontend on Neuronpedia to support community research.
Neighborhood — ranked by edge-count
Methods (1)
method
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.