question
active
question:how-can-reward-functions-be-meaningfully-specified-when-the-same-outcome-may-be-valuable-or-detrimental-depending-on-contextHow can reward functions be meaningfully specified when the same outcome may be valuable or detrimental depending on context?
Motivates active inference's solution: learning prior preferences from interaction rather than external specification.
Source paper
extracted_from(2021) · Noor Sajid · Philip J. Ball · Thomas Parr · Karl J. Friston
Neighborhood — ranked by edge-count
Claims (1)
claim
- §1, contrasting RL reward conceptualization.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- In RL, a scalar signal from the environment that defines the agent's goal; in active inference, reward is just another observation with associated preference.
- Assumption D.3 formalized in the theoretical framework; empirically validated with coin-flip sequence experiments
- Unpredictability is a necessary condition for genuine adaptation.
- Central multiple-realizability claim of the paper, from abstract and §2.
- Rewards are simply predictable stimuli (and aversive stimuli are, by definition, surprising)claim0.740Redefines reward and punishment in terms of predictability.
- Illustrates how non-separable functions shift identity to the collective level.
- Seven categories determined by which components of f[h] are activated: Objective only, Expect only, Compare only, and combinations
- The reward hypothesis underpinning RL, quoted from Sutton and Barto.