method
active
method:grpo-group-relative-policy-optimization

GRPO (Group Relative Policy Optimization)

RL algorithm used to train the activation verbalizer on open models; samples group of candidate descriptions and applies policy optimization.

Neighborhood — ranked by edge-count

Frameworks (1)

framework
  • An unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained jointly with RL.

Methods (2)

method
  • Component of NLA that maps activations to text descriptions; initialized as copy of target LLM with supervised warm-start on summarization task.
  • Token-level auxiliary objective that strengthens optimization of sparse functional tokens during RL by anchoring group-level advantages directly to functional-token positions.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.