claim
active
claim:persona-vector-based-data-filtering-and-llm-judge-based-data-filtering-have-complementary-strengths-for-identifying-problematic-training-dataPersona vector-based data filtering and LLM judge-based data filtering have complementary strengths for identifying problematic training data
Author's interpretive conclusion from comparing filtering strategies
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that core representations like preferences are persona-relative, supporting claim that personas gate content of representations
- Open question proposed by authors for future work on the dimensionality and structure of persona space
- Practical implication for AI safety audit methodology
- First of three hypotheses about persona implementation in LLMs, motivating the persona views
- Persona vectors extracted from base pretraining checkpoints steer fully post-trained OLMo-3-7B-Instructfinding0.800Shows persona directions persist through all alignment stages, answering key part of RQ2
- Author's interpretation establishing that persona vectors are not merely general misalignment indicators
- Training on image data should improve LLM performance, and training on language data should improve vision model performancehypothesis0.793Implication of PRH for cross-modal training efficiency
- How does different post-training data shift a model's position along persona dimensions?question0.790Future work direction: using persona space to study effects of training data on model character