hypothesis
active
hypothesis:individual-cone-basis-vectors-may-correspond-to-interpretable-semantic-facets-of-truth-such-as-temporal-facts-geographic-facts-or-commonsenseIndividual cone basis vectors may correspond to interpretable semantic facets of truth such as temporal facts, geographic facts, or commonsense
Future direction hypothesis for giving semantic meaning to individual axes
Source paper
extracted_from(2025) · Kevin Shengyang Yu · Vaidehi Bulusu · Oscar Yasunaga · Lau, Clayton +4
Neighborhood — ranked by edge-count
Papers (1)
paper
Questions (1)
question
- Central open question for future work on interpretability of cone axes
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Load-bearing illustration of what a concept cone for truth means operationally
- Open question proposed by authors for future work on the dimensionality and structure of persona space
- Key open question about why the persona vector extraction method works beyond correlation
- Author's interpretation establishing that persona vectors are not merely general misalignment indicators
- Concept cone truth interventions would generalize to larger frontier models and multimodal settingshypothesis0.759Key robustness question raised as future work
- First of three hypotheses about persona implementation in LLMs, motivating the persona views
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.
- Appendix E replication of DIM alignment finding in Qwen model