thinker
active
thinker:philip-j-ball

Philip J. Ball

Authored
1
Introduces
0
Studies
0
Affiliations
1
Cited by
0

Authored papers (1)

  • Active inference agents operating under expected free energy minimization achieve 98.90 [98.00, 99.79] average score in a non-stationary FrozenLake OpenAI gym environment, compared to 64.39 [60.33, 68.44] for Bayesian model-based RL with Thompson sampling and 66.08 [63.28, 68.88] for Q-learning (ε=0.1) — a performance gap that emerges specifically because active inference treats environmental change as a context-inference problem rather than a reversal-learning problem, recovering within a single episode after each goal-hole swap. The paper introduces a discrete state-space and time formulation of active inference as its primary expository instrument, decomposing expected free energy G into an epistemic value term (mutual information between outcomes and hidden states) and an extrinsic value term (KL divergence between predicted and preferred outcomes), showing that both exploration and exploitation are expressions of a single objective rather than requiring separate engineering via ε-greedy schedules or temperature hyperparameters. In reward-free conditions where Q-learning freezes into a deterministic circular policy scoring 0.00, the active inference null model (zero prior preferences) still scores 50.03 [49.70, 50.35] through pure information-seeking, and agents equipped with Dirichlet hyperpriors over outcome preferences learn stable behavioral niches — including counter-intuitive hole-seeking — without any external reward signal. The paper argues this implies that reinforcement learning is a limiting special case of active inference in which the epistemic value term is suppressed and preferences are fixed externally, and that reward-free preference learning dissolves the circularity of the reward hypothesis rather than merely circumventing it.

More papers — OpenAlex / S2