method
active
method:latent-anchored-grpo-la-grpo

Latent-Anchored GRPO (LA-GRPO)

Token-level auxiliary objective that strengthens optimization of sparse functional tokens during RL by anchoring group-level advantages directly to functional-token positions.

Neighborhood — ranked by edge-count

Frameworks (1)

framework

Findings (1)

finding
  • During RL training on ATLAS, sparse functional tokens (2.3% of sequences) receive diluted gradient signals from sequence-level advantages propagated across all tokens.

Methods (1)

method

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Metric measuring the mean MSE between self and other-referencing activations across all hidden MLP/attention layers
  • Latent Stitchmethod0.687
    Baseline method using a single orthogonal matrix trained to map source latents to target latents via CL auxiliary loss without behavioral objective.
  • Using per-prompt average SAE latent activations and area under precision-recall curve to discriminate aligned from misaligned models
  • SAE-derived representation space where individual dimensions activate for distinct concepts, enabling cleaner steering
  • Latent Structuresconcept0.679
    Hidden or underdeveloped structures existing 'between the lines' of a configuration that can be enhanced and developed through harmony-seeking computation.
  • Latent entitiesconcept0.677
    Entities that become visible as centers in a configuration (e.g., rectangles of white space around a dot) that were not present before.
  • latent patternsconcept0.675
    Statistical regularities stored in pretrained models.
  • Cost-efficient training algorithm used by DeepSeek-R1 for RL-based reasoning