concept
active
concept:model-diffingModel Diffing
Approach to identify interpretable differences between a base LLM and its fine-tuned version using SAE latents
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The paper's primary mechanistic analysis method: comparing SAE latent activations before and after fine-tuning to identify misalignment-relevant features
- The phenomenon of model internals deviating from desired behavior; MAS is demonstrated to detect this via comparison of toxic vs nontoxic LLMs.
- LLM behavior of generating falsehoods; the multi-dimensional truth subspace raises new risks for subtle manipulation
- When a single conversation involves different models, e.g., routing between standard and reasoning models
- A representation that captures relevant aspects of a system; according to the theorem, the regulator must embody this.
- Technique to measure representational compatibility by integrating intermediate representations of one model into another
- Using interventions to guide model generation behavior, e.g., adding sentiment vectors at inference time
- Probability of data under the model, penalizing complexity and rewarding accuracy.