framework
active
framework:reinforceREINFORCE
Classical RL algorithm adapted by the paper with modifications including clipped-surrogate losses and length-normalized advantages for agentic training.
Neighborhood — ranked by edge-count
Methods (2)
method
- Novel modification to REINFORCE that normalizes step-level advantage by trajectory length to prevent long but low-quality trajectories from dominating training.
- Trajectory FilteringextendsStrategic filtering procedure that removes invalid trajectories and maintains optimal positive-to-negative trajectory ratio to stabilize training.
Frameworks (1)
framework
- SFR-DeepResearchimplementsThe paper's core contribution: an RL-based framework for training autonomous single-agent LLMs to perform deep research with web search, browsing, and code execution.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The property that living structures contain intense contrast—far more than one imagines helpful; true opposites which annihilate each other when superimposed, creating differentiation that gives birth to something; contrast unifies rather than separates when used correctly
- Highlights circularity in RL reward hypothesis; grounds motivation for preference-based active inference.
- Prior framework for synergizing reasoning and acting in LLM agents, foundational to agent harness concept
- Core credit assignment question for distributed systems.
- Parameters controlling the influence of conditioning signals in the generative process.
- Alternative framework for agent behavior; based on reward maximization rather than free energy minimization.
- Attribute: providing a foundation function, a text that acts as base or corroboration.