question
active
question:how-can-subtle-ooc-behavior-within-a-long-form-generation-be-detected-when-response-level-metrics-assign-a-single-scoreHow can subtle OOC behavior within a long-form generation be detected when response-level metrics assign a single score?
Core research question motivating the atomic-level evaluation framework
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Novelty claim establishing the paper's contribution relative to prior work focused on closed-form tasks
- Central critique of prior evaluation: whole-response scoring hides individual OOC sentences
- Key limitation acknowledged by authors.
- Main interpretive finding from Criterion 3 comparison showing Span Representation consistently outperforms IIT under temporal permutation.
- Primary limitation acknowledged by the authors; strongest evidence would require mechanistic activation analysis
- Demonstrates robustness of inference stages to non-fixed-point limiting behavior
- Authors argue their metrics capture a distinct behavioral facet and could detect residual misalignment when standard evaluations indicate improvement