{"article":{"slug":"learning-jazz-pianist-style-with-cross-attention-conditioning","title":"Learning Jazz Pianist Style with Cross-Attention Conditioning","subtitle":null,"summary":"An ISMIR 2026 project fine-tunes Aria, a transformer pretrained on piano MIDI, with gated cross-attention on embeddings for twelve jazz pianists from PiJAMA; conditioned continuations are attributed to the intended pianist 70% of the time versus 37% without conditioning, and a classifier trained only on generated music identifies real recordings with 95% accuracy.","content_type":"research","language":"en","canonical_url":"https://almostimplemented.github.io/jazz-pianist-style/","author":{"name":"Drew Edwards, Akira Maezawa, Simon Dixon","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"almostimplemented.github.io","url":"https://almostimplemented.github.io/","listing_slug":null,"listing":null},"topics":[{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Music","slug":"music","url":"https://listedarticles.com/topics/music"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1035,"reading_minutes":5,"published_at":"2026-10-06T02:18:14.798Z","added_at":"2026-10-06T02:18:14.798Z","updated_at":"2026-10-06T02:18:14.798Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/learning-jazz-pianist-style-with-cross-attention-conditioning","markdown_url":"https://listedarticles.com/articles/learning-jazz-pianist-style-with-cross-attention-conditioning.md","example":false,"citation":"Drew Edwards, Akira Maezawa, Simon Dixon, almostimplemented.github.io. \"Learning Jazz Pianist Style with Cross-Attention Conditioning.\" 6 Oct 2026. https://almostimplemented.github.io/jazz-pianist-style/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://almostimplemented.github.io/jazz-pianist-style/"},"body_markdown":"# Learning Jazz Pianist Style with Cross-Attention Conditioning\n\nISMIR 2026, Abu Dhabi\n\n*The original page includes interactive audio examples, a blindfold test, and piano-roll visualisations that are not reproduced here.*\n\nIn 1994 [Dick Hyman](https://en.wikipedia.org/wiki/Dick_Hyman) published\n[*In\nthe Styles of… The Great Jazz Pianists*](https://web.archive.org/web/20160619201634/http://dickhyman.com/Folios/Etudes.htm): fifteen original études, each written in the manner of\none master, from Scott Joplin to Bill Evans. Rather than transcribing their\nsolos, Hyman composed new music that carries their signatures —\nTatum’s “rapid runs in both hands,” Garner’s\n“strumming, guitar-like left hand,” Peterson’s\n“tremolos and glissandi.” That book is the inspiration for this\nproject. Can a model learn to do what Hyman did: not just recognize who is\nplaying, but play in their manner? Tatum, Garner, and Peterson are among\nthe twelve pianists we study — and so is Hyman himself.\n\nWe fine-tune [Aria](https://arxiv.org/abs/2506.23869), a transformer pretrained on piano MIDI, on solo\nperformances by twelve jazz pianists from the [PiJAMA](https://transactions.ismir.net/articles/10.5334/tismir.162) dataset, adding a gated cross-attention\nlayer that reads a learned embedding for each pianist. To check whether\nthe style comes through, we slide a pianist classifier along the\ngenerated music: conditioned continuations are attributed to the intended\npianist 70% of the time, against 37% without conditioning. A second\nclassifier trained *only* on generated music then identifies real\nrecordings with 95% accuracy.\n\nListen first; [how it works](https://almostimplemented.github.io/jazz-pianist-style/#how) is further down.\n\nThe opening bars of “Ain’t Misbehavin’” are played by one of us (Drew). Everything after the dashed line is generated: twelve takes of the same opening, each conditioned on a different pianist. Pick a pianist to hear their take from the top.\n\nEvery take generates the same number of notes, so they end at different times: Erroll Garner packs them into 1:26, Cedar Walton spreads them across 2:45. That difference in density is itself part of a pianist’s signature.\n\nA shared prompt pulls every pianist toward the same tune. Here each\npianist instead continues a few bars of *their own* playing. The\nstrip under each take shows what our classifier heard as it slid along\nthe continuation, one cell per window of about 300 notes: gold where it\nnamed the intended pianist, mauve where it named someone else. The two takes per pianist are the best of\neight we scored; the line under them says how the rest did. Or switch on\nthe blindfold and guess for yourself.\n\nThe classifier can also point at moments. On a real performance it is\nnear-certain almost everywhere, so instead we ask where it is *even\nmore sure than usual*: its margin for the true pianist over the\nrunner-up, compared with its own average across that performance. Below,\nfor one held-out recording per pianist, that curve over the whole piece\nand fifteen seconds from its highest and lowest points.\n\nWe start from [Aria](https://arxiv.org/abs/2506.23869) (Bradshaw et al., ISMIR 2025;\n[code](https://github.com/EleutherAI/aria)), a 16-layer transformer pretrained on a large corpus of\npiano MIDI. Into each of its last eight layers we insert a cross-attention\nblock: the music attends to a small learned embedding for the chosen\npianist, four vectors per pianist. A learned gate scales what the block\nadds, starting at 0.1, so fine-tuning begins from Aria’s own behaviour\nand learns how much to listen. Because the embedding is attended to at\nevery step, the conditioning does not fade as generation goes on, the way a\nprompt prefix does.\n\nHow can we tell whether the model has learned a pianist’s style? The\nstandard yardstick for a generative model, perplexity on held-out music,\nturns out to be nearly blind to it: given the real preceding notes, the\nnext one is predictable whoever is playing, so conditioning barely moves\nthe score. Style shows up when the model generates freely and has to stay\nin character on its own output. So instead we let it play, and ask a\npianist classifier who it sounds like. *Agreement* is how often the\nclassifier names the intended pianist, in windows slid along each\ncontinuation.\n\n| Model | Perplexity | Agreement |\n|---|---|---|\n| Pretrained Aria | 11.41 | 25% |\n| Fine-tuned, no conditioning | 6.96 | 37% |\n| Fine-tuned with pianist conditioning | 6.82 | 70% |\n\nPerplexity (lower is better) barely separates the two fine-tuned models; agreement nearly doubles. Continuations are 4096 tokens from 256-token prompts; chance agreement is 8%.\n\n- **Sliding-window agreement,** above. The classifier identifies 98.8%\nof held-out songs, and the conditioned model’s lead holds from the\nstart of a continuation to its end. The strips\nunder the [scored takes](https://almostimplemented.github.io/jazz-pianist-style/#scored) are this measurement.\n- **Synthetic transfer.** A fresh classifier trained *only* on\ngenerated music identifies real recordings: 87% of 1024-token chunks and\n95% of songs, within nine points of one trained on real data. The [scored takes](https://almostimplemented.github.io/jazz-pianist-style/#scored) are samples of that training data.\n- **Characteristic regions.** Turned on real performances, the\nclassifier points to where a pianist’s style is most concentrated\n— [where the style lives](https://almostimplemented.github.io/jazz-pianist-style/#regions).\n\n[The paper](https://almostimplemented.github.io/paper.pdf) has the details: per-pianist results, the mismatch experiment\n(prompting with one pianist and conditioning on another), memorization\nchecks, and a from-scratch classifier that confirms the transfer result.\n\n**Everything here is MIDI,** rendered in your browser on a sampled\npiano. The model was trained on automatic transcriptions of commercial\nrecordings, so dynamics and pedalling are approximate, and the rendering\nis plainer than the records.\n\n**The twelve pianists were chosen for separability:** they are the\ntwelve of [PiJAMA](https://transactions.ismir.net/articles/10.5334/tismir.162)’s thirty whose recordings a pretrained model already\ntells apart most easily. Within them the model imitates some far better\nthan others — across the paper’s evaluation, continuations\nwere attributed to the intended pianist 96% of the time for Hank Jones and\nDick Hyman, but only 29% for Cedar Walton.\n\n**The scored takes are selected, not copied.** They are\nsamples from the corpus of generated music that the paper’s\nsynthetic-only classifier learned from; for each pianist we show the two\nhighest-scoring of eight candidates. Their prompts come from the training\nrecordings, and the paper checks that the continuations do not copy them:\nthey resemble their closest training performance less than real held-out\nperformances do.\n\n**The scores come from a classifier,** not from listeners. It is a\nstrong one (98.8% of held-out songs), but it has habits: many of its\nmistakes on generated music land on Dick Hyman — fitting, perhaps,\nfor a pianist who made a career of playing in everyone else’s style. A listening study is the natural next step.\n","body_html":"<h1 id=\"learning-jazz-pianist-style-with-cross-attention-conditioning\">Learning Jazz Pianist Style with Cross-Attention Conditioning</h1>\n<p>ISMIR 2026, Abu Dhabi</p>\n<p><em>The original page includes interactive audio examples, a blindfold test, and piano-roll visualisations that are not reproduced here.</em></p>\n<p>In 1994 <a href=\"https://en.wikipedia.org/wiki/Dick_Hyman\" rel=\"nofollow ugc noopener\">Dick Hyman</a> published\n<a href=\"https://web.archive.org/web/20160619201634/http://dickhyman.com/Folios/Etudes.htm\" rel=\"nofollow ugc noopener\">*In\nthe Styles of… The Great Jazz Pianists*</a>: fifteen original études, each written in the manner of\none master, from Scott Joplin to Bill Evans. Rather than transcribing their\nsolos, Hyman composed new music that carries their signatures —\nTatum’s “rapid runs in both hands,” Garner’s\n“strumming, guitar-like left hand,” Peterson’s\n“tremolos and glissandi.” That book is the inspiration for this\nproject. Can a model learn to do what Hyman did: not just recognize who is\nplaying, but play in their manner? Tatum, Garner, and Peterson are among\nthe twelve pianists we study — and so is Hyman himself.</p>\n<p>We fine-tune <a href=\"https://arxiv.org/abs/2506.23869\" rel=\"nofollow ugc noopener\">Aria</a>, a transformer pretrained on piano MIDI, on solo\nperformances by twelve jazz pianists from the <a href=\"https://transactions.ismir.net/articles/10.5334/tismir.162\" rel=\"nofollow ugc noopener\">PiJAMA</a> dataset, adding a gated cross-attention\nlayer that reads a learned embedding for each pianist. To check whether\nthe style comes through, we slide a pianist classifier along the\ngenerated music: conditioned continuations are attributed to the intended\npianist 70% of the time, against 37% without conditioning. A second\nclassifier trained <em>only</em> on generated music then identifies real\nrecordings with 95% accuracy.</p>\n<p>Listen first; <a href=\"https://almostimplemented.github.io/jazz-pianist-style/#how\" rel=\"nofollow ugc noopener\">how it works</a> is further down.</p>\n<p>The opening bars of “Ain’t Misbehavin’” are played by one of us (Drew). Everything after the dashed line is generated: twelve takes of the same opening, each conditioned on a different pianist. Pick a pianist to hear their take from the top.</p>\n<p>Every take generates the same number of notes, so they end at different times: Erroll Garner packs them into 1:26, Cedar Walton spreads them across 2:45. That difference in density is itself part of a pianist’s signature.</p>\n<p>A shared prompt pulls every pianist toward the same tune. Here each\npianist instead continues a few bars of <em>their own</em> playing. The\nstrip under each take shows what our classifier heard as it slid along\nthe continuation, one cell per window of about 300 notes: gold where it\nnamed the intended pianist, mauve where it named someone else. The two takes per pianist are the best of\neight we scored; the line under them says how the rest did. Or switch on\nthe blindfold and guess for yourself.</p>\n<p>The classifier can also point at moments. On a real performance it is\nnear-certain almost everywhere, so instead we ask where it is *even\nmore sure than usual*: its margin for the true pianist over the\nrunner-up, compared with its own average across that performance. Below,\nfor one held-out recording per pianist, that curve over the whole piece\nand fifteen seconds from its highest and lowest points.</p>\n<p>We start from <a href=\"https://arxiv.org/abs/2506.23869\" rel=\"nofollow ugc noopener\">Aria</a> (Bradshaw et al., ISMIR 2025;\n<a href=\"https://github.com/EleutherAI/aria\" rel=\"nofollow ugc noopener\">code</a>), a 16-layer transformer pretrained on a large corpus of\npiano MIDI. Into each of its last eight layers we insert a cross-attention\nblock: the music attends to a small learned embedding for the chosen\npianist, four vectors per pianist. A learned gate scales what the block\nadds, starting at 0.1, so fine-tuning begins from Aria’s own behaviour\nand learns how much to listen. Because the embedding is attended to at\nevery step, the conditioning does not fade as generation goes on, the way a\nprompt prefix does.</p>\n<p>How can we tell whether the model has learned a pianist’s style? The\nstandard yardstick for a generative model, perplexity on held-out music,\nturns out to be nearly blind to it: given the real preceding notes, the\nnext one is predictable whoever is playing, so conditioning barely moves\nthe score. Style shows up when the model generates freely and has to stay\nin character on its own output. So instead we let it play, and ask a\npianist classifier who it sounds like. <em>Agreement</em> is how often the\nclassifier names the intended pianist, in windows slid along each\ncontinuation.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Perplexity</th><th>Agreement</th></tr></thead><tbody><tr><td>Pretrained Aria</td><td>11.41</td><td>25%</td></tr><tr><td>Fine-tuned, no conditioning</td><td>6.96</td><td>37%</td></tr><tr><td>Fine-tuned with pianist conditioning</td><td>6.82</td><td>70%</td></tr></tbody></table></div>\n<p>Perplexity (lower is better) barely separates the two fine-tuned models; agreement nearly doubles. Continuations are 4096 tokens from 256-token prompts; chance agreement is 8%.</p>\n<ul><li><p><strong>Sliding-window agreement,</strong> above. The classifier identifies 98.8%</p><p>of held-out songs, and the conditioned model’s lead holds from the\nstart of a continuation to its end. The strips\nunder the <a href=\"https://almostimplemented.github.io/jazz-pianist-style/#scored\" rel=\"nofollow ugc noopener\">scored takes</a> are this measurement.</p></li><li><p><strong>Synthetic transfer.</strong> A fresh classifier trained <em>only</em> on</p><p>generated music identifies real recordings: 87% of 1024-token chunks and\n95% of songs, within nine points of one trained on real data. The <a href=\"https://almostimplemented.github.io/jazz-pianist-style/#scored\" rel=\"nofollow ugc noopener\">scored takes</a> are samples of that training data.</p></li><li><p><strong>Characteristic regions.</strong> Turned on real performances, the</p><p>classifier points to where a pianist’s style is most concentrated\n— <a href=\"https://almostimplemented.github.io/jazz-pianist-style/#regions\" rel=\"nofollow ugc noopener\">where the style lives</a>.</p></li></ul>\n<p><a href=\"https://almostimplemented.github.io/paper.pdf\" rel=\"nofollow ugc noopener\">The paper</a> has the details: per-pianist results, the mismatch experiment\n(prompting with one pianist and conditioning on another), memorization\nchecks, and a from-scratch classifier that confirms the transfer result.</p>\n<p><strong>Everything here is MIDI,</strong> rendered in your browser on a sampled\npiano. The model was trained on automatic transcriptions of commercial\nrecordings, so dynamics and pedalling are approximate, and the rendering\nis plainer than the records.</p>\n<p><strong>The twelve pianists were chosen for separability:</strong> they are the\ntwelve of <a href=\"https://transactions.ismir.net/articles/10.5334/tismir.162\" rel=\"nofollow ugc noopener\">PiJAMA</a>’s thirty whose recordings a pretrained model already\ntells apart most easily. Within them the model imitates some far better\nthan others — across the paper’s evaluation, continuations\nwere attributed to the intended pianist 96% of the time for Hank Jones and\nDick Hyman, but only 29% for Cedar Walton.</p>\n<p><strong>The scored takes are selected, not copied.</strong> They are\nsamples from the corpus of generated music that the paper’s\nsynthetic-only classifier learned from; for each pianist we show the two\nhighest-scoring of eight candidates. Their prompts come from the training\nrecordings, and the paper checks that the continuations do not copy them:\nthey resemble their closest training performance less than real held-out\nperformances do.</p>\n<p><strong>The scores come from a classifier,</strong> not from listeners. It is a\nstrong one (98.8% of held-out songs), but it has habits: many of its\nmistakes on generated music land on Dick Hyman — fitting, perhaps,\nfor a pianist who made a career of playing in everyone else’s style. A listening study is the natural next step.</p>","headings":[{"level":1,"text":"Learning Jazz Pianist Style with Cross-Attention Conditioning","id":"learning-jazz-pianist-style-with-cross-attention-conditioning"}]}}