concept
active
concept:pretraining

Pretraining

Initial large-scale training phase whose early stages are shown to form persona representations

Neighborhood — ranked by edge-count

Concepts (1)

concept
  • The diverse distribution learned by LLMs during pretraining that alignment training sharpens; VS aims to recover it

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Expected prevalence of patterns (e.g., base-10 arithmetic) in pretraining corpora, influencing ρd and dr.
  • Post-Trainingconcept0.779
    The phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.
  • Authors' policy recommendation based on finding that persona representations form and persist from early pretraining
  • Character Trainingconcept0.758
    The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
  • Surgical Trainingconcept0.747
    Training approach targeting only functionally specialized components to avoid catastrophic forgetting and misalignment
  • Learningconcept0.741
    Inference of parameters encoding contingencies of the world (e.g., likelihood matrix A) at slower timescale than perception.
  • Broader research area: methods to align model behavior after initial training, where undesired behaviors can emerge.
  • Early learningconcept0.739
    The primary domain in which Nicholson's Theory of Loose Parts became influential and known.