A common task in machine learning is to estimate or optimize “log-sum-exp” functions with (potentially continuously) many terms such as
where (v) is some potential and (q) a probability distribution. Applications include normalizing probabilistic models, smooth approximations to the maximum, transformers via derivatives, and entropy-regularized RL.
The key difficulty is variance when (v) takes large values. For i.i.d. normals, the relative squared error for estimating (\mathbb{E}[e^z]) is ((e^{\sigma^2}-1)/n) — exploding exponentially in (\sigma).
Main question: Can we keep the advantages of optimizing log-sum-exp while being less exposed to their computational/statistical disadvantages?
The magic of least-squares
Least-squares offers closed-form estimation for linear models, controlled variance via moments, and sharp analyses — but used naively for discrete outputs (e.g. one-hot classification) creates artefacts such as masking.
The new attempt is summarized by the integral identity:
[t \log t - t + 1 = \int_0^1 \frac{(t-1)^2}{\rho t + 1-\rho}(1-\rho)\, d\rho.]
Relative density estimation as a testbed
Estimate (\log(dp/dq)), equivalent via variational formulations of KL / (f)-divergences (Nguyen–Wainwright–Jordan; Donsker–Varadhan). Empirical averages of exponentials are unstable; the post explores another way ([7]).
Weighted chi-square divergences
For (f(t)=\frac12\frac{(t-1)^2}{\rho t+1-\rho}), the divergence has a quadratic variational form — exactly a least-squares prediction problem (related to noise-contrastive estimation).
Extension by integration
Writing a general (f) as an integral of these weighted chi-squares yields a continuum of least-squares problems. KL corresponds to (d\nu(\rho)=2(1-\rho)d\rho). Potentials (v) and (w) are recovered by integrating the per-(\rho) solutions.
With linear features, each (\rho) problem is a linear system. Integrating via a generalized eigenvalue decomposition of ((\Sigma_p,\Sigma_q)) yields a closed-form divergence and potentials — complexity (O(m^2n+m^3)), improvable with shared low-rank feature learning.
Benefits of spectral estimation
High-dimensional Gaussian analyses and simulations show spectral KL estimation reduces variance vs direct variational approaches when samples are limited; Pearson ((\rho=0)) alone is geometrically less natural.
Applied to mutual information, the framework yields closed-form softmax-like conditional density estimation — not equivalent to one-hot least-squares — with large computational gains vs Newton softmax in some regimes.
Feature learning
Maximizing the spectral lower bound end-to-end with MM/EM-style algorithms enables feature learning at scale.
Conclusion
A continuum of stable least-squares problems, made feasible by one generalized eigendecomposition, circumvents exploding exponential means — with a closed-form path for softmax-style last layers. See arXiv:2605.10668 (NeurIPS).
Frontier LLMs used for figures/typos; technical content by the author.