{"article":{"slug":"how-to-win-a-beer-with-high-dimensional-statistics","title":"How to win a beer with high-dimensional statistics","subtitle":null,"summary":"Jamie Simon explains a viral high-dimensional statistics paper with a bar-bet framing: why naive intuition about data geometry fails, and how the right summary wins the round.","content_type":"essay","language":"en","canonical_url":"https://jamiesimon.io/blog/how-to-win-a-beer-with-high-dimensional-statistics/","author":{"name":"Jamie Simon","url":"https://jamiesimon.io/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Jamie Simon","url":"https://jamiesimon.io/","listing_slug":null,"listing":null},"topics":[{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Mathematics","slug":"mathematics","url":"https://listedarticles.com/topics/mathematics"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":827,"reading_minutes":4,"published_at":"2026-09-24T12:00:00.000Z","added_at":"2026-09-29T06:16:40.280Z","updated_at":"2026-09-29T06:16:40.280Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/how-to-win-a-beer-with-high-dimensional-statistics","markdown_url":"https://listedarticles.com/articles/how-to-win-a-beer-with-high-dimensional-statistics.md","example":false,"citation":"Jamie Simon, Jamie Simon. \"How to win a beer with high-dimensional statistics.\" 24 Sept 2026. https://jamiesimon.io/blog/how-to-win-a-beer-with-high-dimensional-statistics/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://jamiesimon.io/blog/how-to-win-a-beer-with-high-dimensional-statistics/"},"body_markdown":"My longtime labmate-turned-student/friend1 Dhruva Karkada recently wrote [a sick paper on data statistics](<https://arxiv.org/abs/2602.15029>) which deservedly went viral on Twitter, in part because it has one of the prettiest scientific figures I have ever seen:\n\n![Calendar month embeddings forming a circle in PCA space, with circulant Gram matrices](/img/circulant_words/months_circle.png)\n\nSay you’ve got a bunch of words that live in a vocabulary $\\mathcal{V}$. We’re here studying models $f$ that map $f: \\mathcal{V} \\rightarrow \\mathbb{R}^d$: that is, they map every word to a $d$-dimensional vector. We’re letting $\\\\{v_i\\\\}_{i=1}^{12} = \\\\{\\texttt{January}, \\texttt{February}, \\ldots\\\\}$ be the months of the year, taking the 12 associated embedding vectors $\\mathbf{w}_i = f(v_i)$, and computing two things:\n\n  * a projection onto the top two PCA directions of $\\\\{ \\mathbf{w}_i \\\\}$ (left column), and\n  * the Gram matrix $\\mathbf{M} \\in \\mathbb{R}^{12 \\times 12}$ such that $M_{ij} = \\mathbf{w}_i^\\top \\mathbf{w}_j$ (right column).\n\nReading the rows of this figure from top to bottom,2 they find that:\n\n  1. LLM embeddings project down to a circle (as [Engels et al (2024)](<https://arxiv.org/abs/2405.14860>) also saw), and the Gram matrix is approximately a circulant matrix;\n  2. these findings are decently approximated even with primitive `word2vec` embeddings; and\n  3. an analytical theory of the circulant Gram matrix gives a very compelling-looking match.\n\nThis is a big deal because it connects data statistics to representational geometry with a really simple mathematical theory.\n\n### Finding a needle\n\nAfter seeing this a bunch of times and staring at it for a while, I was feeling in the mood to poke a hole in this beautiful result, and so I bet Dhruva a beer that I could find a collection of _other, seemingly-unrelated words_ that form a circle + circulant matrix in the same way. He (and most others I told) thought this was crazy, since the circle clearly comes from the special relationship between the words. We settled on the terms of the bet: I had to find ten random-seeming words whose `word2vec` embeddings, when plotted as the above, made a clear and compelling circle.\n\nWhy’d I think this was possible? Well, we have vocabulary of $25000$ words to choose from. That gives you $N = \\binom{25000}{10} \\approx 3 \\times 10^{37}$ sets to select from. I figured that if you threw ten darts at a board that many times, you’d definitely make a circle at least once. Info-theoretically speaking, you have $\\log_2 N \\approx 124$ bits of information, and surely you can make a decent 10-point circle with that amount of resolving power. The question’s just how you find a set of ten good words in the haystack.\n\nHere’s how I did it:\n\n  * From looking at PCA plots of random sets, I guess you’d get a decent circle from a random selection with probability maybe $3^{-10}$, so random guessing could plausibly work.\n  * I wrote a “looks circular” objective function, drew tens of thousands of random sets, and chose the best one. It was borderline, but not good enough to utterly obliterate Dhruva.\n  * I upgraded it to an iterative search, where at every step, we drop the worst point and choose the best replacement from the vocabulary. That worked pretty well.\n  * I also changed the objective from “looks circular on a PCA plot” to “matches a target circulant Gram matrix.” That worked really damn well.\n\nNote that the embedding size $d = 10000$ never entered into this. This all took an afternoon with a coding agent.\n\nHere’s what I got:\n\n![Ten seemingly-unrelated words forming a circle in PCA space, with a circulant Gram matrix](/img/circulant_words/spurious_circle.png)\n\nThat’s circular. You can just find other sets of random-looking words that form circles!\n\n### Does this have any actual significance?\n\nThis raises certain open questions, including “how can one man be so wrong?”, which I am not qualified to answer.\n\nBut seriously: clearly we can find spurious geometric patterns. Should this change our understanding of representation geometry? I’d note a few caveats first:\n\n  1. While the Gram matrix of my spurious-circle-set is indeed beautifully circulant, the amplitude of the (sinusoidal) off-diagonals is less than with the months. I couldn’t get em up to match the months’ Gram matrix, even to within a factor of two.\n  2. This works damn well with a set of ten, but I doubt it’d work with a set of, say, 50 (though admittedly I didn’t try very hard), so Dhruva’s other geometric findings (about e.g. all the years from 1700-2020) couldn’t be spoofed in this way.\n\nNonetheless, it does show that doing a kind of pursuit-matching-style search for a certain low-dim PCA’d geometry will trick you unless you’ve got enough statistical constraints on your target that it won’t happen by random chance! This does rule out certain automatic-feature-finding algorithms, which has implications for research agendas like scalable interpretability.\n\n* * *\n\n  1. If this description seems wordy, it’s because my job is confusing. ↩\n\n  2. Note: I’ve reversed the order of the rows for storytelling ease, not that it matters. ↩\n\n* * *","body_html":"<p>My longtime labmate-turned-student/friend1 Dhruva Karkada recently wrote <a href=\"https://arxiv.org/abs/2602.15029\" rel=\"nofollow ugc noopener\">a sick paper on data statistics</a> which deservedly went viral on Twitter, in part because it has one of the prettiest scientific figures I have ever seen:</p>\n<p>Calendar month embeddings forming a circle in PCA space, with circulant Gram matrices</p>\n<p>Say you’ve got a bunch of words that live in a vocabulary $\\mathcal{V}$. We’re here studying models $f$ that map $f: \\mathcal{V} \\rightarrow \\mathbb{R}^d$: that is, they map every word to a $d$-dimensional vector. We’re letting $\\{v_i\\}_{i=1}^{12} = \\{\\texttt{January}, \\texttt{February}, \\ldots\\}$ be the months of the year, taking the 12 associated embedding vectors $\\mathbf{w}_i = f(v_i)$, and computing two things:</p>\n<ul><li>a projection onto the top two PCA directions of $\\{ \\mathbf{w}_i \\}$ (left column), and</li><li>the Gram matrix $\\mathbf{M} \\in \\mathbb{R}^{12 \\times 12}$ such that $M_{ij} = \\mathbf{w}_i^\\top \\mathbf{w}_j$ (right column).</li></ul>\n<p>Reading the rows of this figure from top to bottom,2 they find that:</p>\n<ol><li>LLM embeddings project down to a circle (as <a href=\"https://arxiv.org/abs/2405.14860\" rel=\"nofollow ugc noopener\">Engels et al (2024)</a> also saw), and the Gram matrix is approximately a circulant matrix;</li><li>these findings are decently approximated even with primitive <code>word2vec</code> embeddings; and</li><li>an analytical theory of the circulant Gram matrix gives a very compelling-looking match.</li></ol>\n<p>This is a big deal because it connects data statistics to representational geometry with a really simple mathematical theory.</p>\n<h3 id=\"finding-a-needle\">Finding a needle</h3>\n<p>After seeing this a bunch of times and staring at it for a while, I was feeling in the mood to poke a hole in this beautiful result, and so I bet Dhruva a beer that I could find a collection of <em>other, seemingly-unrelated words</em> that form a circle + circulant matrix in the same way. He (and most others I told) thought this was crazy, since the circle clearly comes from the special relationship between the words. We settled on the terms of the bet: I had to find ten random-seeming words whose <code>word2vec</code> embeddings, when plotted as the above, made a clear and compelling circle.</p>\n<p>Why’d I think this was possible? Well, we have vocabulary of $25000$ words to choose from. That gives you $N = \\binom{25000}{10} \\approx 3 \\times 10^{37}$ sets to select from. I figured that if you threw ten darts at a board that many times, you’d definitely make a circle at least once. Info-theoretically speaking, you have $\\log_2 N \\approx 124$ bits of information, and surely you can make a decent 10-point circle with that amount of resolving power. The question’s just how you find a set of ten good words in the haystack.</p>\n<p>Here’s how I did it:</p>\n<ul><li>From looking at PCA plots of random sets, I guess you’d get a decent circle from a random selection with probability maybe $3^{-10}$, so random guessing could plausibly work.</li><li>I wrote a “looks circular” objective function, drew tens of thousands of random sets, and chose the best one. It was borderline, but not good enough to utterly obliterate Dhruva.</li><li>I upgraded it to an iterative search, where at every step, we drop the worst point and choose the best replacement from the vocabulary. That worked pretty well.</li><li>I also changed the objective from “looks circular on a PCA plot” to “matches a target circulant Gram matrix.” That worked really damn well.</li></ul>\n<p>Note that the embedding size $d = 10000$ never entered into this. This all took an afternoon with a coding agent.</p>\n<p>Here’s what I got:</p>\n<p>Ten seemingly-unrelated words forming a circle in PCA space, with a circulant Gram matrix</p>\n<p>That’s circular. You can just find other sets of random-looking words that form circles!</p>\n<h3 id=\"does-this-have-any-actual-significance\">Does this have any actual significance?</h3>\n<p>This raises certain open questions, including “how can one man be so wrong?”, which I am not qualified to answer.</p>\n<p>But seriously: clearly we can find spurious geometric patterns. Should this change our understanding of representation geometry? I’d note a few caveats first:</p>\n<ol><li>While the Gram matrix of my spurious-circle-set is indeed beautifully circulant, the amplitude of the (sinusoidal) off-diagonals is less than with the months. I couldn’t get em up to match the months’ Gram matrix, even to within a factor of two.</li><li>This works damn well with a set of ten, but I doubt it’d work with a set of, say, 50 (though admittedly I didn’t try very hard), so Dhruva’s other geometric findings (about e.g. all the years from 1700-2020) couldn’t be spoofed in this way.</li></ol>\n<p>Nonetheless, it does show that doing a kind of pursuit-matching-style search for a certain low-dim PCA’d geometry will trick you unless you’ve got enough statistical constraints on your target that it won’t happen by random chance! This does rule out certain automatic-feature-finding algorithms, which has implications for research agendas like scalable interpretability.</p>\n<ul><li>* *<ol><li>If this description seems wordy, it’s because my job is confusing. ↩</li><li>Note: I’ve reversed the order of the rows for storytelling ease, not that it matters. ↩</li></ol></li><li>* *</li></ul>","headings":[{"level":3,"text":"Finding a needle","id":"finding-a-needle"},{"level":3,"text":"Does this have any actual significance?","id":"does-this-have-any-actual-significance"}]}}