method
active
method:revealed-preferences-evaluation

Revealed Preferences Evaluation

Novel evaluation method that measures a model's preference to express one character trait over another via Elo scoring, avoiding self-report issues

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Behavioral and stated consistency that implies the model is pursuing some objective, without claiming genuine internal states
  • The ability of active inference agents to learn their own prior preferences over outcomes by accumulating Dirichlet parameters from experience.
  • The problematic possibility of digital minds with superhumanly strong preferences requiring interpersonal utility comparison frameworks
  • Prior Preferencesconcept0.740
    Target distribution over states or outcomes encoded in the generative model; goal states.
  • Post-training alignment method during which undesirable behaviors emerged in the studied model.
  • Replaces explicit reward signal in active inference; encodes agent's preferred observations independent of environment.
  • Designing digital minds to have preferences that are trivially easy to satisfy, yielding high welfare at minimal resource cost
  • Key element for alignment faking: model's pre-existing preferences contradict the new training objective