artifact
active
artifact:natural-language-autoencoder-training-code-and-models

Natural Language Autoencoder Training Code and Models

Released on GitHub (github.com/kitft/natural_language_autoencoders) and interactive frontend on Neuronpedia to support community research.

Neighborhood — ranked by edge-count

Methods (1)

method
  • Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.