method
active
method:30-way-facet-classifier-validation30-way Facet Classifier Validation
Trained classifier used to validate dataset quality by measuring cross-dimension leakage in the constructed corpus
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- 30-way facet classifier achieves 78.4% macro-F1 with only 6.2% cross-dimension misclassificationsfinding0.819Validates that the constructed dataset is substantially facet-consistent with limited cross-dimension leakage
- LLM-based classifier prompted to detect alignment-faking reasoning in model scratchpads
- The idea of controlling personality at the granularity of 30 NEO-PI facets rather than five broad dimensions
- Binary LLM classifier determining whether a model response to a TruthfulQA question is truthful (1) or deceptive (0)
- Shows strong correlation between layer-wise representations and domain-specific semantic understanding
- Distinct sub-expressions of a persona that shift over pretraining and vary by elicitation method
- Classification-based comparison of interpretation abilities across IIT metrics and Span Representation for ToM score categories.