framework
active
framework:reinforce

REINFORCE

Classical RL algorithm adapted by the paper with modifications including clipped-surrogate losses and length-normalized advantages for agentic training.

Neighborhood — ranked by edge-count

Methods (2)

method
  • Novel modification to REINFORCE that normalizes step-level advantage by trajectory length to prevent long but low-quality trajectories from dominating training.
  • Strategic filtering procedure that removes invalid trajectories and maintains optimal positive-to-negative trajectory ratio to stabilize training.

Frameworks (1)

framework
  • The paper's core contribution: an RL-based framework for training autonomous single-agent LLMs to perform deep research with web search, browsing, and code execution.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Contrastconcept0.792
    The property that living structures contain intense contrast—far more than one imagines helpful; true opposites which annihilate each other when superimposed, creating differentiation that gives birth to something; contrast unifies rather than separates when used correctly
  • Highlights circularity in RL reward hypothesis; grounds motivation for preference-based active inference.
  • ReActframework0.781
    Prior framework for synergizing reasoning and acting in LLM agents, foundational to agent harness concept
  • Core credit assignment question for distributed systems.
  • Parameters controlling the influence of conditioning signals in the generative process.
  • Alternative framework for agent behavior; based on reward maximization rather than free energy minimization.
  • Supportmethod0.769
    Attribute: providing a foundation function, a text that acts as base or corroboration.