Abstract
This thesis investigates the problem of discrepancy minimization in the approximation of probability distributions. In sample space, a distribution can be approximated by particles with the evolution describedby a differential equation, such as Langevin dynamics which is known as a gradient flow that minimizes
the Kullback-Leibler divergence in Wasserstein geometry. In parameter space, a distribution can be
approximated by fitting a parametric distribution, such as through variational inference. Although both
the particle approach and the parametric approach aim to minimize discrepancy measurements between
probability distributions, they are still studied in parallel. This thesis aims to provide a comprehensive
view on the connection between the particle approach and the parametric approach.
Our first contribution is to fix the inconsistency between the theoretical explanation and practical
algorithms for divergence generative adversarial nets (GANs). In order to do that, we proposed a novel
generative modeling framework–MonoFlow which formulates divergence GANs as simulating and
distilling particles along a probability flow ordinary differential equation (ODE). This ODE follows the
Wasserstein gradient flow to minimize f-divergences. Our framework first provides a unified view for
divergence GANs and diffusion models.
Our second contribution is to establish the connection between Euclidean variational inference (VI)
approaches and Wasserstein gradient flows. We showed the Gaussian black-box variational inference
(BBVI) is Wasserstein natural, i.e., the Euclidean gradient of the Kullback-Leibler divergence indicates
the steepest direction under the Wasserstein geometry. Consequently, sampling via the Ornstein Uhlenbeck process can be regarded as equivalent to Gaussian BBVI–a parametric approach. Additionally, we
propose a novel parametric VI approach with f-divergences via distilling a probability flow ODE. Our
method generalizes prior work and bridges the gap between VI and Wasserstein gradient flows.
Our third contribution is to propose a novel VI approach to minimize a probability distance, the sliced
Wasserstein distance, which can be evaluated on particles sampled from distributions. This approach
utilized sampling with Markov Chain Monte Carlo and avoids the requirement for an explicit variational
distribution, such that we can use a parametric neural network as the approximation.
| Date of Award | 1 Oct 2024 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Vladislav Tadic (Supervisor) & Song Liu (Supervisor) |
Cite this
- Standard