claim
active
claim:traits-such-as-truthfulness-and-harmfulness-can-be-stationary-in-certain-contexts-but-reflective-in-othersTraits such as truthfulness and harmfulness can be stationary in certain contexts but reflective in others.
Central finding of Section 5 on trait dynamics in interactions.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Result from Experiment 6, Fig. 5 right — variance changes but mean does not.
- Conceptual open question about formalizing the persona abstraction beyond the trait level
- Comparison result from Experiment 6.
- Reveals correlational structure in the persona space with practical implications for monitoring
- GPT-4 exhibits reflective truthfulness because it is the only model capable enough to perform the necessary in-context learning.hypothesis0.770Proposed explanation for why GPT-4 uniquely shows reflective truthfulness under long untruthful contexts.
- Critical verbatim statement highlighting the universal inference basis of sentience.
- Reflective mode comprises three separable traits: latent capacity, default accessibility, and stability of access.hypothesis0.766Decomposition from prompt lift data: models may have capacity without accessibility (Grok 4 high-gated), and stability varies (Haiku Δ=0.02 vs GPT-5.4 Δ=1.00).
- Driving hypothesis for robustness experiments in Section 3.2