paper:beal-2003-variationalVariational Algorithms for Approximate Bayesian Inference
Original abstract (expand)
The Bayesian framework for machine learning allows for the incorporation of prior knowledge in a coherent way, avoids overfitting problems, and provides a principled basis for selecting between alternative models. Unfortunately the computations required are usually intractable. This thesis presents a unified variational Bayesian (VB) framework which approximates these computations in models with latent variables using a lower bound on the marginal likelihood. \n \nChapter 1 presents background material on Bayesian inference, graphical models, and propagation algorithms. Chapter 2 forms the theoretical core of the thesis, generalising the expectation- maximisation (EM) algorithm for learning maximum likelihood parameters to the VB EM algorithm which integrates over model parameters. The algorithm is then specialised to the large family of conjugate-exponential (CE) graphical models, and several theorems are presented to pave the road for automated VB derivation procedures in both directed and undirected graphs (Bayesian and Markov networks, respectively). \n \nChapters 3–5 derive and apply the VB EM algorithm to three commonly-used and important models: mixtures of factor analysers, linear dynamical systems, and hidden Markov models. It is shown how model selection tasks such as determining the dimensionality, cardinality, or number of variables are possible using VB approximations. Also explored are methods for combining sampling procedures with variational approximations, to estimate the tightness of VB bounds and to obtain more effective sampling algorithms. Chapter 6 applies VB learning to a long-standing problem of scoring discrete-variable directed acyclic graphs, and compares the performance to annealed importance sampling amongst other methods. Throughout, the VB approximation is compared to other methods including sampling, Cheeseman-Stutz, and asymptotic approximations such as BIC. The thesis concludes with a discussion of evolving directions for model selection including infinite models and alternative approximations to the marginal likelihood.
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- Bayesian score calibration for approximate modelsDavid J Warne, David J Nott, Christopher Drovandi Joshua J Bon2026≈ 75%
- ≈ 74%
- ≈ 74%
- ≈ 73%
- Permutation-based Inference for Variational Learning of Directed Acyclic GraphsPantelis Elinas, He Zhao, Maurizio Filippone, Vassili Kitsios, Terry O'Kane Edwin V. Bonilla2026≈ 72%
- Designing Perceptual Puzzles by Differentiating Probabilistic ProgramsTzu-Mao Li, Joshua Tenenbaum, Jonathan Ragan-Kelley Kartik Chandra2022≈ 72%
- A Framework for Improving the Reliability of Black-box Variational InferenceMichael Riis Andersen, Aki Vehtari, Jonathan H. Huggins Manushi Welandawe2025≈ 72%
- Active Inference for Binary Symmetric Hidden Markov ModelsArmen E. Allahverdyan and Aram Galstyan2015≈ 71%
- A Framework for Variational Inference of Lightweight Bayesian Neural Networks with Heteroscedastic UncertaintiesRyan Brown, Michael Merritt, Samuel Park, Delsin Menolascino, Mark A. Peot David J. Schodt2026≈ 71%
- ≈ 71%
- Riemannian Laplace Approximation with the Fisher MetricMarcelo Hartmann, Bernardo Williams, Mark Girolami, Arto Klami Hanlin Yu2026≈ 71%
- ≈ 71%
- Compositional Active Inference II: Polynomial Dynamics. Approximate Inference DoctrinesToby St. Clere Smithe2022≈ 70%
- Uncertainty Quantification and Propagation in Surrogate-based Bayesian InferenceJavier Enrique Aguilar, Anneli Guthke, Paul-Christian B\"urkner Philipp Reiser2026≈ 70%
- A Review of Bayesian Uncertainty Quantification in Deep Probabilistic Image SegmentationR.J.G. van Sloun, C.G.A. Viviers, P.H.N. de With, F. van der Sommen M.M.A. Valiuddin2026≈ 70%
- Active Inference, Curiosity and Insightin corpus2017≈ 68%
- ≈ 68%
- Active Inference: A Process Theoryin corpus2017≈ 67%
- ≈ 63%
- Active inference: demystified and comparedin corpus2021≈ 63%
- ≈ 62%
- ≈ 61%
- Interpreting Language Model Parametersin corpus2026≈ 59%
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsin corpus2023≈ 59%
- Learning without neurons in physical systemsin corpus2022≈ 59%
- ≈ 59%
- The biogenic approach to cognitionin corpus2005≈ 59%
- Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencodersin corpus2026≈ 59%
- ≈ 59%
- Explaining 4.2 million genetic variants with state-of-the-art, interpretable predictionsin corpus2026≈ 58%
Similar preprints — Semantic Scholar
Cited by (3)
- Active Inference, Curiosity and Insight
Minimizing expected variational free energy under a discrete-state Markov decision process generative model is sufficient to produce curiosity, epistemic learning, and insight without any additional m
- Active Inference: A Process Theory
A single variational principle—minimizing variational free energy via gradient descent on a Markov decision process (MDP) generative model—is sufficient to derive neuronal dynamics that reproduce, wit
- Life as we know it
Any ergodic random dynamical system possessing a Markov blanket will, almost surely, appear to engage in active inference and maintain autopoietic integrity—making biological self-organization not a r