claim
active
claim:what-a-model-exposes-by-default-tracks-the-norms-it-was-trained-toward-while-steering-acts-on-deviations-from-those-defaults-rather-than-on-the-defaults-themselvesWhat a model exposes by default tracks the norms it was trained toward, while steering acts on deviations from those defaults rather than on the defaults themselves
Central interpretive claim organizing the entire paper's results
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Addresses skeptical alternative that reports reflect only conversational content
- Which behaviors does a model represent internally, default to, can be pushed to amplify, or refuses to expose?question0.827The motivating diagnostic question that prompting alone cannot answer
- Practical implication drawn from the prosocial persona paradox finding
- Validation finding from Lu et al. 2026 supporting PSM's claim about pre-training persona structure
- The top of the steerability ranking is dominated by exaggerated or attention-grabbing styles
- Mechanism claim supported by transcript analysis and the fact that the steering vector was extracted from a model that never writes type hints.
- Nuanced interpretive claim about the limits of steering as a mechanism for reflection enhancement.