method
active
method:training-data-synthesis-pipelineTraining Data Synthesis Pipeline
Iterative approach to construct challenging synthetic multi-hop QA pairs, long-form report writing tasks, and math/code reasoning tasks that exceed difficulty of existing datasets.
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- SFR-DeepResearchimplementsThe paper's core contribution: an RL-based framework for training autonomous single-agent LLMs to perform deep research with web search, browsing, and code execution.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Supervised stage method: model generates response, then critiques it according to a principle, then revises it; repeated multiple times.
- The large corpus of human-generated text on which LLMs are trained, which provisions character archetypes and narrative structures
- Individual examples used during post-training that can cause specific behaviors.
- Christopher Alexander’s early method that decomposes design problems into a hierarchical tree of requirements and synthesizes form as a balance of forces.
- Training approach targeting only functionally specialized components to avoid catastrophic forgetting and misalignment
- Broader research area: methods to align model behavior after initial training, where undesired behaviors can emerge.
- The standard evolutionary theory integrating Darwinian selection with Mendelian genetics; paper argues it needs expansion with MCA.
- Primary worked example demonstrating denotational design principles