method
active
method:grpo-group-relative-policy-optimizationGRPO (Group Relative Policy Optimization)
RL algorithm used to train the activation verbalizer on open models; samples group of candidate descriptions and applies policy optimization.
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- Natural Language Autoencoders (NLA)implementsAn unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained jointly with RL.
Methods (2)
method
- Activation Verbalizer (AV)implementsComponent of NLA that maps activations to text descriptions; initialized as copy of target LLM with supervised warm-start on summarization task.
- Token-level auxiliary objective that strengthens optimization of sparse functional tokens during RL by anchoring group-level advantages directly to functional-token positions.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Cost-efficient training algorithm used by DeepSeek-R1 for RL-based reasoning
- Evolutionary technique that evolves computer programs, discussed as a route toward self-modifying models.
- Optimization method used in distillation stage to learn behavioral expression of desired traits
- RL algorithm used for training models to comply with the conflicting objective
- Metric averaged over all tasks to measure MTL method improvement over STL.
- Choosing sequences of actions based on expected free energy; prior probability of policy is softmax of expected free energy
- Ability to apply learned solutions to novel circumstances.
- A set of instructions for making something (contrasted with a descriptive blueprint), as in embryonic development.