---
title: "Exploding variance of means of exponentials: least-squares to the rescue"
slug: exploding-variance-of-means-of-exponentials-least-squares-to-the-rescue
url: https://listedarticles.com/articles/exploding-variance-of-means-of-exponentials-least-squares-to-the-rescue
canonical_url: https://francisbach.com/spectral_log_density_estimation/
content_type: research
language: en
published_at: 2026-09-25T12:00:00.000Z
updated_at: 2026-09-27T09:13:15.391Z
author: "Francis Bach"
author_url: https://francisbach.com
authored_by: human
publisher: "Francis Bach"
publisher_url: https://francisbach.com
topics: ["Machine Learning", "Research", "Mathematics", "AI"]
license: all-rights-reserved
word_count: 464
reading_minutes: 2
citation: "Francis Bach, Francis Bach. \"Exploding variance of means of exponentials: least-squares to the rescue.\" 25 Sept 2026. https://francisbach.com/spectral_log_density_estimation/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Exploding variance of means of exponentials: least-squares to the rescue

> Francis Bach reframes log-sum-exp / KL estimation as a continuum of least-squares problems with closed-form spectral solutions—cutting exploding exponential variance.

# Exploding variance of means of exponentials: least-squares to the rescue

**Author:** Francis Bach  
**Published:** September 25, 2026  
**Source:** [Machine Learning Research Blog](https://francisbach.com/spectral_log_density_estimation/)

A common task in machine learning is to estimate or optimize “log-sum-exp” functions with (potentially continuously) many terms such as

\[\log \Big( \int_{\mathcal{X}} e^{v(x)} dq(x) \Big),\]

where \(v\) is some potential and \(q\) a probability distribution. Applications include normalizing probabilistic models, smooth approximations to the maximum, transformers via derivatives, and entropy-regularized RL.

The key difficulty is variance when \(v\) takes large values. For i.i.d. normals, the relative squared error for estimating \(\mathbb{E}[e^z]\) is \((e^{\sigma^2}-1)/n\) — exploding exponentially in \(\sigma\).

**Main question:** Can we keep the advantages of optimizing log-sum-exp while being less exposed to their computational/statistical disadvantages?

## The magic of least-squares

Least-squares offers closed-form estimation for linear models, controlled variance via moments, and sharp analyses — but used naively for discrete outputs (e.g. one-hot classification) creates artefacts such as masking.

The new attempt is summarized by the integral identity:

\[t \log t - t + 1 = \int_0^1 \frac{(t-1)^2}{\rho t + 1-\rho}(1-\rho)\, d\rho.\]

## Relative density estimation as a testbed

Estimate \(\log(dp/dq)\), equivalent via variational formulations of KL / \(f\)-divergences (Nguyen–Wainwright–Jordan; Donsker–Varadhan). Empirical averages of exponentials are unstable; the post explores another way ([7]).

## Weighted chi-square divergences

For \(f(t)=\frac12\frac{(t-1)^2}{\rho t+1-\rho}\), the divergence has a quadratic variational form — exactly a least-squares prediction problem (related to noise-contrastive estimation).

## Extension by integration

Writing a general \(f\) as an integral of these weighted chi-squares yields a continuum of least-squares problems. KL corresponds to \(d\nu(\rho)=2(1-\rho)d\rho\). Potentials \(v\) and \(w\) are recovered by integrating the per-\(\rho\) solutions.

## Linear models with closed-form spectral estimation

With linear features, each \(\rho\) problem is a linear system. Integrating via a generalized eigenvalue decomposition of \((\Sigma_p,\Sigma_q)\) yields a closed-form divergence and potentials — complexity \(O(m^2n+m^3)\), improvable with shared low-rank feature learning.

## Benefits of spectral estimation

High-dimensional Gaussian analyses and simulations show spectral KL estimation reduces variance vs direct variational approaches when samples are limited; Pearson (\(\rho=0\)) alone is geometrically less natural.

## Mutual information and conditional estimation

Applied to mutual information, the framework yields closed-form softmax-like conditional density estimation — not equivalent to one-hot least-squares — with large computational gains vs Newton softmax in some regimes.

## Feature learning

Maximizing the spectral lower bound end-to-end with MM/EM-style algorithms enables feature learning at scale.

## Conclusion

A continuum of stable least-squares problems, made feasible by one generalized eigendecomposition, circumvents exploding exponential means — with a closed-form path for softmax-style last layers. See arXiv:2605.10668 (NeurIPS).

*Frontier LLMs used for figures/typos; technical content by the author.*
