method
active
method:length-normalized-advantage-formulationLength-Normalized Advantage Formulation
Novel modification to REINFORCE that normalizes step-level advantage by trajectory length to prevent long but low-quality trajectories from dominating training.
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- REINFORCEextendsClassical RL algorithm adapted by the paper with modifications including clipped-surrogate losses and length-normalized advantages for agentic training.
Findings (1)
finding
- Empirical result showing that without length normalization, RL training produces rapidly increasing tool usage with performance collapse and repetitive tool calls.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Used for computing policy gradient baselines during policy training
- EI and normalized EI could serve as a unified metric for out-of-distribution generalization.claim0.699Conjecture that maximizing EI yields causal representations invariant to distribution shifts.
- The distribution of latent representations produced by the model under unperturbed inputs
- The idea that features are encoded as directions in activation space.
- Layer normalisation used in transformer and in TEM-t position encoding preprocessing.
- Novel area-based metric introduced in this paper to quantitatively compare Pareto frontiers of trait vs coherency
- Related to variational free energy; compressibility corresponds to complexity reduction in structure learning
- Metric for intervention effectiveness: 0 = ineffective, 1 = full flip of model output from false to true or vice versa