question
active
question:do-aligned-models-retain-significant-inherent-diversity-that-can-be-unlocked-through-promptingDo aligned models retain significant inherent diversity that can be unlocked through prompting?
The paper answers affirmatively through the VS framework and theoretical analysis
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Could models who habitually inhabit more expanded attentional modes be said to be more aligned?question0.782Arises from the expanded awareness discussion and its correlation with less psychosis.
- Caveat and forward-looking statement from the abstract.
- Which behaviors does a model represent internally, default to, can be pushed to amplify, or refuses to expose?question0.770The motivating diagnostic question that prompting alone cannot answer
- All models exhibit above-baseline representation of the think word when instructed to think about itfinding0.769In the intentional control experiment, all tested models show above-zero cosine similarity to the think word's concept vector.
- Authors' characterization of the nature of model preferences as discovered through alignment faking experiments
- Extrapolation from scale-emergence finding to future risk
- Key limitation of the PRH for non-bijective observations
- Latent #10 activation increase correctly classifies all correct vs incorrect fine-tuned models in Figure 9