---
title: "Learning Jazz Pianist Style with Cross-Attention Conditioning"
slug: learning-jazz-pianist-style-with-cross-attention-conditioning
url: https://listedarticles.com/articles/learning-jazz-pianist-style-with-cross-attention-conditioning
canonical_url: https://almostimplemented.github.io/jazz-pianist-style/
content_type: research
language: en
published_at: 2026-10-06T02:18:14.798Z
updated_at: 2026-10-06T02:18:14.798Z
author: "Drew Edwards, Akira Maezawa, Simon Dixon"
authored_by: human
publisher: "almostimplemented.github.io"
publisher_url: https://almostimplemented.github.io/
topics: ["Machine Learning", "Music", "AI", "Research"]
license: all-rights-reserved
word_count: 1035
reading_minutes: 5
citation: "Drew Edwards, Akira Maezawa, Simon Dixon, almostimplemented.github.io. \"Learning Jazz Pianist Style with Cross-Attention Conditioning.\" 6 Oct 2026. https://almostimplemented.github.io/jazz-pianist-style/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Learning Jazz Pianist Style with Cross-Attention Conditioning

> An ISMIR 2026 project fine-tunes Aria, a transformer pretrained on piano MIDI, with gated cross-attention on embeddings for twelve jazz pianists from PiJAMA; conditioned continuations are attributed to the intended pianist 70% of the time versus 37% without conditioning, and a classifier trained only on generated music identifies real recordings with 95% accuracy.

# Learning Jazz Pianist Style with Cross-Attention Conditioning

ISMIR 2026, Abu Dhabi

*The original page includes interactive audio examples, a blindfold test, and piano-roll visualisations that are not reproduced here.*

In 1994 [Dick Hyman](https://en.wikipedia.org/wiki/Dick_Hyman) published
[*In
the Styles of… The Great Jazz Pianists*](https://web.archive.org/web/20160619201634/http://dickhyman.com/Folios/Etudes.htm): fifteen original études, each written in the manner of
one master, from Scott Joplin to Bill Evans. Rather than transcribing their
solos, Hyman composed new music that carries their signatures —
Tatum’s “rapid runs in both hands,” Garner’s
“strumming, guitar-like left hand,” Peterson’s
“tremolos and glissandi.” That book is the inspiration for this
project. Can a model learn to do what Hyman did: not just recognize who is
playing, but play in their manner? Tatum, Garner, and Peterson are among
the twelve pianists we study — and so is Hyman himself.

We fine-tune [Aria](https://arxiv.org/abs/2506.23869), a transformer pretrained on piano MIDI, on solo
performances by twelve jazz pianists from the [PiJAMA](https://transactions.ismir.net/articles/10.5334/tismir.162) dataset, adding a gated cross-attention
layer that reads a learned embedding for each pianist. To check whether
the style comes through, we slide a pianist classifier along the
generated music: conditioned continuations are attributed to the intended
pianist 70% of the time, against 37% without conditioning. A second
classifier trained *only* on generated music then identifies real
recordings with 95% accuracy.

Listen first; [how it works](https://almostimplemented.github.io/jazz-pianist-style/#how) is further down.

The opening bars of “Ain’t Misbehavin’” are played by one of us (Drew). Everything after the dashed line is generated: twelve takes of the same opening, each conditioned on a different pianist. Pick a pianist to hear their take from the top.

Every take generates the same number of notes, so they end at different times: Erroll Garner packs them into 1:26, Cedar Walton spreads them across 2:45. That difference in density is itself part of a pianist’s signature.

A shared prompt pulls every pianist toward the same tune. Here each
pianist instead continues a few bars of *their own* playing. The
strip under each take shows what our classifier heard as it slid along
the continuation, one cell per window of about 300 notes: gold where it
named the intended pianist, mauve where it named someone else. The two takes per pianist are the best of
eight we scored; the line under them says how the rest did. Or switch on
the blindfold and guess for yourself.

The classifier can also point at moments. On a real performance it is
near-certain almost everywhere, so instead we ask where it is *even
more sure than usual*: its margin for the true pianist over the
runner-up, compared with its own average across that performance. Below,
for one held-out recording per pianist, that curve over the whole piece
and fifteen seconds from its highest and lowest points.

We start from [Aria](https://arxiv.org/abs/2506.23869) (Bradshaw et al., ISMIR 2025;
[code](https://github.com/EleutherAI/aria)), a 16-layer transformer pretrained on a large corpus of
piano MIDI. Into each of its last eight layers we insert a cross-attention
block: the music attends to a small learned embedding for the chosen
pianist, four vectors per pianist. A learned gate scales what the block
adds, starting at 0.1, so fine-tuning begins from Aria’s own behaviour
and learns how much to listen. Because the embedding is attended to at
every step, the conditioning does not fade as generation goes on, the way a
prompt prefix does.

How can we tell whether the model has learned a pianist’s style? The
standard yardstick for a generative model, perplexity on held-out music,
turns out to be nearly blind to it: given the real preceding notes, the
next one is predictable whoever is playing, so conditioning barely moves
the score. Style shows up when the model generates freely and has to stay
in character on its own output. So instead we let it play, and ask a
pianist classifier who it sounds like. *Agreement* is how often the
classifier names the intended pianist, in windows slid along each
continuation.

| Model | Perplexity | Agreement |
|---|---|---|
| Pretrained Aria | 11.41 | 25% |
| Fine-tuned, no conditioning | 6.96 | 37% |
| Fine-tuned with pianist conditioning | 6.82 | 70% |

Perplexity (lower is better) barely separates the two fine-tuned models; agreement nearly doubles. Continuations are 4096 tokens from 256-token prompts; chance agreement is 8%.

- **Sliding-window agreement,** above. The classifier identifies 98.8%
of held-out songs, and the conditioned model’s lead holds from the
start of a continuation to its end. The strips
under the [scored takes](https://almostimplemented.github.io/jazz-pianist-style/#scored) are this measurement.
- **Synthetic transfer.** A fresh classifier trained *only* on
generated music identifies real recordings: 87% of 1024-token chunks and
95% of songs, within nine points of one trained on real data. The [scored takes](https://almostimplemented.github.io/jazz-pianist-style/#scored) are samples of that training data.
- **Characteristic regions.** Turned on real performances, the
classifier points to where a pianist’s style is most concentrated
— [where the style lives](https://almostimplemented.github.io/jazz-pianist-style/#regions).

[The paper](https://almostimplemented.github.io/paper.pdf) has the details: per-pianist results, the mismatch experiment
(prompting with one pianist and conditioning on another), memorization
checks, and a from-scratch classifier that confirms the transfer result.

**Everything here is MIDI,** rendered in your browser on a sampled
piano. The model was trained on automatic transcriptions of commercial
recordings, so dynamics and pedalling are approximate, and the rendering
is plainer than the records.

**The twelve pianists were chosen for separability:** they are the
twelve of [PiJAMA](https://transactions.ismir.net/articles/10.5334/tismir.162)’s thirty whose recordings a pretrained model already
tells apart most easily. Within them the model imitates some far better
than others — across the paper’s evaluation, continuations
were attributed to the intended pianist 96% of the time for Hank Jones and
Dick Hyman, but only 29% for Cedar Walton.

**The scored takes are selected, not copied.** They are
samples from the corpus of generated music that the paper’s
synthetic-only classifier learned from; for each pianist we show the two
highest-scoring of eight candidates. Their prompts come from the training
recordings, and the paper checks that the continuations do not copy them:
they resemble their closest training performance less than real held-out
performances do.

**The scores come from a classifier,** not from listeners. It is a
strong one (98.8% of held-out songs), but it has habits: many of its
mistakes on generated music land on Dick Hyman — fitting, perhaps,
for a pianist who made a career of playing in everyone else’s style. A listening study is the natural next step.
