method
active
method:length-normalized-advantage-formulation

Length-Normalized Advantage Formulation

Novel modification to REINFORCE that normalizes step-level advantage by trajectory length to prevent long but low-quality trajectories from dominating training.

Neighborhood — ranked by edge-count

Frameworks (1)

framework
  • REINFORCE
    extends
    Classical RL algorithm adapted by the paper with modifications including clipped-surrogate losses and length-normalized advantages for agentic training.

Findings (1)

finding

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.