concept
active
concept:prior-preferences-over-outcomesPrior Preferences over Outcomes
Replaces explicit reward signal in active inference; encodes agent's preferred observations independent of environment.
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- Active Inferenceassociated_withimplementsFoundational framework by Karl Friston; the paper extends it to three hierarchical levels for modeling meta-awareness.
Concepts (1)
concept
- Prior Preferencesrelated_tosame_asTarget distribution over states or outcomes encoded in the generative model; goal states.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key element for alignment faking: model's pre-existing preferences contradict the new training objective
- Sensory data observed by the agent at each time step.
- Beliefs about states before data; used to transcribe task instructions into agent's generative model
- Cognitive bifurcation event where second-order contextual constraints reorganize semantic space, establishing weighted alternatives for action.
- The ability of active inference agents to learn their own prior preferences over outcomes by accumulating Dirichlet parameters from experience.
- Prior expectations encoded in top-down cortical signals; balance with bottom-up input modulated by precision.
- The problematic possibility of digital minds with superhumanly strong preferences requiring interpersonal utility comparison frameworks