The self-normalized IS estimator is widely used to estimate expectations with intractable normalizing constants, for example, in Bayesian leave-one-out cross validation or likelihood free inference. In this paper, we propose a framework to understand when SNIS works and when it does not, with a generalization that allows us to overcome its limitations, with connections to continuous optimal transport.
NeurIPS 2024 Workshop on Bayesian Decision-making and Uncertainty
Monte Carlo
To estimate posterior expectations consistently, we need to use self-normalized importance sampling (or MCMC, but SNIS has a better variance lower bound). It is a ratio of two IS estimators. Typical diagnostics forget this, and only look at IS-weights for numerator or denominator separately. We try to capture this information with the concept of tail dependence of random variables, which applies in heavy-tailed scenarios. Ongoing journal extension.
Kviman, Oskar and Tamogashev, Kirill and Branchini, Nicola and Elvira, Víctor and Lagergren, Jens and Malkin, Nikolay
Earlier version: NeurIPS workshop — 2nd edition of Frontiers in Probabilistic Inference: Learning meets Sampling
Monte CarloOptimal transportDynamical systems
Existing multimarginal flow matching (FM) methods either do not scale well with dimension or encourage trajectories to pass through intermediate marginal samples, rather than the intermediate distributions. We learn a parameterised interpolant for FM via a GAN-inspired loss, which addresses these shortcomings.
To estimate µ = E_p[f(θ)] when p's normalizing constant is unknown, instead of doing MCMC on p(θ) or even p(θ)|f(θ)|, or learning a parametric q(θ), we try MCMC directly on p(θ)|f(θ)- µ|, which is the asymptotic-variance minimizing proposal. We propose a simple iterative scheme that works: initial estimate µ₀; run a chain on the approximation p(θ)|f(θ)- µ₀|; estimate µ again with SNIS, and keep iterating.
NeurIPS 2024 Workshop on Bayesian Decision-making and Uncertainty
Monte Carlo
To estimate posterior expectations consistently, we need to use self-normalized importance sampling. Typical diagnostics forget that SNIS is a ratio of two IS estimators. We capture dependence between numerator and denominator via tail dependence of random variables in heavy-tailed scenarios. Ongoing journal extension.
A framework to understand when SNIS works and when it does not, with a generalization that overcomes its limitations, with connections to continuous optimal transport.
Guilmeau, Thomas♦ and Branchini, Nicola♦ and Chouzenoux, Emilie and Elvira, Víctor (♦ equal contribution)
Monte Carlo
Many adaptive IS (and some VI) methods match moments of a target. When the target has heavy tails, these moments can be undefined or hard to estimate. We propose an AIS method that matches moments of a lighter-tailed modified target (exponentiated to power alpha), while minimizing the alpha-divergence to the true target.
Kviman, Oskar and Branchini, Nicola and Elvira, Víctor and Lagergren, Jens
Monte Carlo
Instead of enforcing that particle replication counts match pre-resampling weights in expectation, we optimize replication counts to minimize a divergence between the post- and pre-resampling distributions directly.
Felekis, Yorgos and Zennaro, Fabio and Branchini, Nicola and Damoulas, Theodoros
Statistical causalityOptimal transport
We learn causal abstractions from data without specifying parametric SCM functions, via a multimarginal OT problem with soft constraints and a cost encoding knowledge of the underlying causal DAGs. The soft constraints have a do-calculus interpretation.
A journal extension of the optimized APF paper: at each iteration we want a mixture proposal close to a mixture target. Literature often matches term-by-term; this view suggests methods that match the two mixtures directly.
Branchini, Nicola and Aglietti, Virginia and Dhir, Neil and Damoulas, Theodoros
Statistical causality
We study causal global optimization under unknown graphs: the effect of incorrect causal assumptions, and an acquisition function that trades off optimization of the effect and structure learning.
We improve the Auxiliary Particle Filter by optimizing resampling weights as mixture weights of an importance sampling mixture proposal. Choosing mixture weights to minimize empirical variance of importance weights leads to a convex optimization problem.