thinker:karl-fristonKarl Friston
Author of the free energy principle framework; central thinker in the paper.
Authored papers (8)
Active inference agents operating under expected free energy minimization achieve 98.90 [98.00, 99.79] average score in a non-stationary FrozenLake OpenAI gym environment, compared to 64.39 [60.33, 68.44] for Bayesian model-based RL with Thompson sampling and 66.08 [63.28, 68.88] for Q-learning (ε=0.1) — a performance gap that emerges specifically because active inference treats environmental change as a context-inference problem rather than a reversal-learning problem, recovering within a single episode after each goal-hole swap. The paper introduces a discrete state-space and time formulation of active inference as its primary expository instrument, decomposing expected free energy G into an epistemic value term (mutual information between outcomes and hidden states) and an extrinsic value term (KL divergence between predicted and preferred outcomes), showing that both exploration and exploitation are expressions of a single objective rather than requiring separate engineering via ε-greedy schedules or temperature hyperparameters. In reward-free conditions where Q-learning freezes into a deterministic circular policy scoring 0.00, the active inference null model (zero prior preferences) still scores 50.03 [49.70, 50.35] through pure information-seeking, and agents equipped with Dirichlet hyperpriors over outcome preferences learn stable behavioral niches — including counter-intuitive hole-seeking — without any external reward signal. The paper argues this implies that reinforcement learning is a limiting special case of active inference in which the epistemic value term is suppressed and preferences are fixed externally, and that reward-free preference learning dissolves the circularity of the reward hypothesis rather than merely circumventing it.
Ramstead, Kirchhoff, and Friston argue that generative models in active inference under the free energy principle (FEP) are control systems—not structural representations—and that this distinction has been systematically obscured by conflating active inference with brain-centered Bayesian frameworks such as predictive coding (Rao & Ballard, 1999), the Helmholtz machine (Dayan et al., 1995), and Bishop's (2006) variational machine learning. The load-bearing move is a technical one: under the FEP, generative models are entailed by the adaptive dynamics of an organism rather than encoded in physical neural states, while it is the recognition density that is embodied—parameterized by the sufficient statistics of internal Markov blanket states. The paper introduces the construct of enactive inference, a reinterpretation grounding the generative model as a normative control system in the tradition of Conant and Ross Ashby's (1970) good regulator theorem, distinguishing it sharply from the structural representationalist accounts advanced by Kiefer and Hohwy (2018, 2019) and Gładziejewski and Miłkowski (2017). Representationalists are correct that internal states encode exploitable structural similarities, but they misidentify the vehicle: those states parameterize the recognition density, not the generative model. The paper argues this implies that perception and action are inseparable moments of a single policy-selection process, that cognitive science should shift from asking how brains represent the world to how organisms enact attunement to their ecological niche, and that enactivism and the mathematical apparatus of active inference are mutually reinforcing rather than in tension.
Active inference on discrete state-spaces, formalized as partially observable Markov decision processes (POMDPs) with likelihood matrix A, transition matrix B, and prior D, unifies perception, planning, decision-making, learning, and structure learning under two objective functions: variational free energy (an upper bound on surprise minimized during state estimation) and expected free energy G(π) (minimized during policy selection). The synthesis derives neuronal dynamics from first principles via gradient descent on free energy, showing that state estimation corresponds to a softmax function of accumulated prediction errors—equations interpretable as membrane potentials mapping to firing rates—and that these dynamics coincide exactly with variational message passing, while the Bethe approximation yields belief propagation. Policy selection follows Q(π) = σ(−G(π)), where G decomposes into risk (KL divergence between predicted and preferred states) and ambiguity (expected entropy of outcomes given states), formally subsuming KL control, expected utility theory, and optimal Bayesian design as special cases. Learning of A follows Dirichlet parameter accumulation **a** = a + Σ(oτ ⊗ sτ), which is formally equivalent to Hebbian plasticity, while structure learning proceeds via Bayesian model reduction (BMR) for simplification and Bayesian model expansion for concept acquisition, with the marginal approximation implemented in `spm_MDP_VB_X.m` identified as the most biologically plausible free energy approximation. The paper argues this implies that biological cognition—from saccadic sampling at ~4 Hz to dopaminergic precision encoding γ—is fully accountable as free energy minimization, and that the outstanding challenge is identifying the evidence-maximizing generative model an agent actually employs, which would constitute a complete structure learning roadmap.
Minimizing expected variational free energy under a discrete-state Markov decision process generative model is sufficient to produce curiosity, epistemic learning, and insight without any additional machinery. Friston et al. 2017 demonstrates this across two linked mechanisms: first, including posterior beliefs about likelihood parameters **A** in expected free energy G(π) introduces a novelty term—information gain about model parameters—that drives agents to sample combinations of hidden states and outcomes they have not yet encountered, resolving ignorance rather than merely ambiguity or risk. Second, Bayesian model reduction (implemented via the spm_MDP_VB_X.m routine in SPM) allows post-hoc or online pruning of redundant concentration parameters: a reduced model is accepted when ΔF ≤ −3, corresponding to a Bayes factor of approximately 20:1 in favor of the simpler model. Simulated agents learning a 3-rule, 4-factor abstract contingency task (144 hidden-state combinations, 36 possible outcomes) reach near-perfect performance after roughly 14 trials under pure epistemic learning, dropping to approximately 10 trials when online Bayesian model reduction is applied across 64 simulated agents. The sleep analog—non-REM synaptic pruning followed by REM-like belief re-evaluation—is formalized identically through the same free energy difference equation. The paper argues this implies that aha moments are necessarily subpersonal events (optimization of the generative model itself, not modeling of that optimization), that the quality of intelligence is inversely related to the thermodynamic energy expended during convergence via the Jarzynski equality, and that communicating reduced model priors rather than parameter posteriors constitutes a principled formal account of shared knowledge—consciousness in the pre-Cartesian sense of con-scire.
A single variational principle—minimizing variational free energy via gradient descent on a Markov decision process (MDP) generative model—is sufficient to derive neuronal dynamics that reproduce, without hand-tuning, more than 10 well-characterized empirical phenomena simultaneously: repetition suppression, mismatch negativity, violation responses (peaking ~200 ms in peristimulus time allowing 100 ms conduction delays), place-cell activity, phase precession, theta sequences, theta-gamma coupling (at ~4 Hz theta with nested gamma), evidence accumulation with race-to-bound stepping dynamics, and transfer of dopamine responses from unconditioned to conditioned stimuli. The method introduced is an active inference process theory grounded in belief propagation over discrete-time MDP generative models, where neuronal firing rates encode categorical state expectations, membrane potentials encode their logarithms, and postsynaptic currents correspond to free-energy gradients (state prediction errors). Simulations use outcomes sampled every 250 ms, eight hidden states over four locations and two contexts, and utilities of ±3 nats for rewarding versus unrewarding outcomes (~20-fold preference ratio). Dopamine is formalized as encoding precision (inverse temperature γ) with a postsynaptic time constant of ~1 s (κ₁/κ₂ = 1/64 per 16 ms iteration). Because a gradient descent constitutes a valid description of neuronal activity, variational free energy functions as a Lyapunov function for neuronal dynamics, implying that neural activity conforms to Hamilton's principle of least action and that a single imperative—free energy minimization—unifies perception, action, learning, and neuromodulatory signaling within one coherent process theory.
Any ergodic random dynamical system possessing a Markov blanket will, almost surely, appear to engage in active inference and maintain autopoietic integrity—making biological self-organization not a remarkable exception but a near-inevitable consequence of coupled dynamical systems with short-range interactions. The argument proceeds via the Helmholtz decomposition and Fokker–Planck formalism: once a system's flow is expressed as a gradient ascent on the log ergodic density (equation 2.5), internal and active states behave as if minimizing variational free energy—a functional introduced by Feynman for path-integral problems and here repurposed as the central organizing quantity of living systems. To demonstrate this, a synthetic primordial soup of 128 subsystems was simulated for 2048 s using forward Euler integration at 1/512 s time steps, with electrochemical dynamics governed by a Lorenz attractor and short-range coupling (unit-distance adjacency matrix); spectral graph theory and the Perron–Frobenius theorem were used to extract the principal Markov blanket, identifying k=8 internal states. Canonical variates analysis of internal-state eigenvariates (obtained via singular value decomposition with ±16 s temporal embedding) predicted external subsystem motion with p=0.00052 against a time-reversed null, confirming statistically that internal states encode posterior beliefs about hidden states. Simulated lesions of active, sensory, or internal states each produced structural disintegration—oscillator death—confirming autopoiesis. The paper argues this implies that life, perception, and adaptive inference are not properties requiring special substrates or carbon-based chemistry, but are instead the generic signature of any ergodic system with a Markov blanket, with evolution itself potentially interpretable as descent onto a global random attractor.
Free energy minimization unifies action and perception as two faces of a single optimization principle: an agent suppresses surprise by either updating its internal model (perception) or by acting on the world to sample only expected sensory states (action). Derived from the probabilistic behavior of an ensemble of agents belonging to the same phenotypic class, the free energy bound approximates the log-evidence (marginal likelihood) of a generative model, making it formally equivalent to negative surprise and negative value simultaneously. The lecture series introduces Dynamic Expectation Maximisation (DEM), a variational filtering scheme that inverts nonlinear dynamic causal models in generalised coordinates of motion and yields both time-dependent conditional state densities and time-independent parameter densities, explicitly superseding Kalman and particle filtering for online Bayesian inversion. Presented across three sessions at the Collège de France in May–June 2008, the framework grounds perception in Helmholtz's neural-energy constructs extended through Empirical Bayes and hierarchical generative models, and reframes dopamine not as a reward signal per se but as encoding the conditional precision—certainty—of predictions, consistent with its role in balancing bottom-up sensory drive against top-down empirical priors. This implies that classical and operant conditioning introduce statistical regularities that are learned by the same hierarchical inference machinery used for causal structure, and that rewards are simply predictable (low-surprise) stimuli, making value learning a special case of perceptual inference rather than an independent computational faculty.
More papers — OpenAlex / S2
Originates (4)
Affiliations (4)
- Wellcome Trust Centre for Neuroimaging, UCL Institute of Neurology(institute)
- California Institute for Machine Consciousness(institute)
- University College London(institute)
- Multiscale Cognition(community)
Co-authors (12)
- Noor Sajid6 shared
- Thomas Parr6 shared
- Lancelot Da Costa4 shared
- Sebastijan Veselic4 shared
- Victorita Neacsu4 shared
- Giovanni Pezzulo3 shared
- Karl J. Friston3 shared
- Susumu Inui3 shared
- J. Allan Hobson2 shared
- Marco Lin2 shared
- Maxwell J. D. Ramstead2 shared
- Philip J. Ball2 shared
Their work is cited by (10)
- A tale of two densities: active inference is enactive inference3× refs
- Active Inference, Curiosity and Insight3× refs
- Technological Approach to Mind Everywhere: An Experimentally-Grounded Framework for Understanding Diverse Bodies and Minds2× refs
- The computational boundary of a 'self': developmental bioelectricity drives multicellularity and scale-free cognition2× refs
- Self-Improvising Memory: A Perspective on Memories as Agential, Dynamically Reinterpreting Cognitive Glue2× refs
- Active Inference: A Process Theory2× refs
- Darwin's agential materials: evolutionary implications of multiscale competency in developmental biology2× refs
- Endless forms most beautiful 2.0: teleonomy and the bioengineering of chimaeric and synthetic organisms2× refs
- Active inference on discrete state-spaces: a synthesis2× refs
- Active inference: demystified and compared1× refs
Other inbound relations (17)
- associated_withActive Inference(framework)
- associated_withBayesian Mechanics(framework)
- associated_withFree Energy Principle(framework)
- associated_withSurprise Minimization Framework(framework)
- citescimcWhitepaper(paper)
- citesDevelopmental Bioelectricity: the cognitive glue enabling evolutionary scaling from physiology to mind(paper)
- citesL'annuaire du Collège de France 2007-2008(venue)
- citesLarge Language Models Report Subjective Experience Under Self-Referential Processing(artifact)
- citesLevin 2022 Technological Approach to Mind Everywhere(artifact)
- citesLevin M (2022) Technological Approach to Mind Everywhere: An Experimentally-Grounded Framework for Understanding Diverse Bodies and Minds(artifact)
- citesMichael Edward Johnson(thinker)
- citesTechnological Approach to Mind Everywhere: An Experimentally-Grounded Framework for Understanding Diverse Bodies and Minds(paper)
- citesWhy Learning Requires Feeling(paper)
- mentionsA Mathematical Elucidation of Buddhist Enlightenment(paper)
- mentionsActive Inference with a Self-Prior in the Mirror-Mark Task(paper)
- mentionsCollective intelligence: A unifying concept for integrating biology across scales and substrates(paper)
- mentionsContemplative Artificial Intelligence (Laukkonen et al., 2025)(concept)
Recent mentions (12)
- papers-typedinui-2026-mathematical-elucidation.md
- machine-consciousnesscimcWhitepaper.md
- papers-typedkim-2026-active-inference.md
- papers-typedlaukkonen-2025-contemplative-agent.md
- papers-typedberg-2026-learning.md
- papers-typedcameron-2025-large.md
- papers-typedlevin-2024-bridge.md
- papers-typedsandved-smith-2026-there.md
- papers-typedvol0123456789.md
- papers-typeddacosta_2020_active_inference_discrete.md
- papers-typedfriston_2013_life_as_we_know_it.md