Exploding variance of means of exponentials: least-squares to the rescue

Author: Francis Bach
Published: September 25, 2026
Source: Machine Learning Research Blog

A common task in machine learning is to estimate or optimize “log-sum-exp” functions with (potentially continuously) many terms such as

[\log \Big( \int_{\mathcal{X}} e^{v(x)} dq(x) \Big),]

where (v) is some potential and (q) a probability distribution. Applications include normalizing probabilistic models, smooth approximations to the maximum, transformers via derivatives, and entropy-regularized RL.

The key difficulty is variance when (v) takes large values. For i.i.d. normals, the relative squared error for estimating (\mathbb{E}[e^z]) is ((e^{\sigma^2}-1)/n) — exploding exponentially in (\sigma).