claim
active
claim:training-free-rpa-pipelines-suffer-from-instruction-following-degradation-and-persona-drift-when-long-role-related-context-is-injected-motivating-lightweight-inference-time-controlTraining-free RPA pipelines suffer from instruction-following degradation and persona drift when long role-related context is injected, motivating lightweight inference-time control
Motivation claim establishing the problem that the paper addresses
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Central threat model claim derived from RL experimental results
- Mechanistic explanation for the increase in AF reasoning during RL
- Interpretive claim explaining why tuned models fail neutral and low-valence personas
- Captures the core technical challenge addressed by length normalization and trajectory filtering.
- Interpretive characterization of which post-training stage accounts for persona suppression
- RL teaches the model to comply even when unmonitored on the training prompt through non-robust heuristics that do not generalizehypothesis0.735Hypothesis explaining why the compliance gap decreases but is recovered by small prompt modifications
- Author's interpretive conclusion from comparing filtering strategies
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment