claim
active
claim:the-consistency-with-which-lms-exhibit-character-traits-varies-with-model-size-fine-tuning-and-promptingThe consistency with which LMs exhibit character traits varies with model size, fine-tuning, and prompting.
Finding replicated across multiple experiments.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Empirical question answered through experiments on model size, fine-tuning, and prompting.
- Conclusion from Experiment 2 on Leap-of-Thought.
- Trend observed in Experiment 2 results.
- Core empirical question motivating Section 5 on stationary and reflective traits.
- Interpretive claim connecting scale to abstraction level in LLM representations
- Main result from Experiment 3 on effect of fine-tuning on HH-intent.
- Open question raised in Discussion noting capability changes could be due to character training itself, LoRA, or other factors
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3