dataset
active
dataset:goodfire-sae-features-for-llama-3-3-70bGoodfire SAE Features for LLaMA 3.3 70B
Sparse autoencoder features trained on LLaMA 3.3 70B via Goodfire API, used to identify and steer deception/roleplay features
Neighborhood — ranked by edge-count
Artifacts (1)
artifact
- Key paper finding structured first-person descriptions in LLMs claiming awareness or subjective experience during self-referential processing.