About

Built by someone who plays

Working a song out by ear takes me longer than learning to play it. That is the whole reason this exists, and it is why the tool is shaped the way it is.

Why it is a practice page and not a PDF

I am a drummer and a developer. Several tools already turned audio into drum notation when I started, and they all ended the same way: a correct file, handed over, and you still open it next to your kit, find your place, start the record, pause, count bars, rewind. The machine solved the part a machine finds hard and left me the part a human finds hard.

So the transcription here opens as a page you play from. The recording lives inside it, a cursor runs through the score, any bar loops, the metronome follows the record's own pulse rather than an average tempo, and a slider rewrites the part at five difficulty levels while the song keeps playing in full. That last one matters more than anything else on the list: a beginner who fails at the full transcription of a song they love usually concludes they are not ready for the song, when they are simply not ready for that version of it.

The model is ours, and that was not a preference

The strongest published open model for this task is ADTOF (Zehren et al., ISMIR 2021), a CRNN trained on 359 hours of real music. It works well. Its entire repository — code and weights — is licensed CC BY-NC-SA: non-commercial use only. Nothing could be built on it, so the model here was trained from scratch.

Training data was the hard part. Hand-labelling drums means marking every hit to within tens of milliseconds, which is not feasible at scale, and the open datasets are mostly synthetic — drum parts played by samplers. A model trained on synthetic cymbals hears real ones badly, and no amount of augmentation fixed it; the gap turned out to be timbre, not artefacts. The answer came from rhythm-game charts, where thousands of people have spent years marking up drum parts of real recordings by hand.

How it is measured

The only test I trust is a holdout: ten songs with full human annotation, kept out of training entirely. Against ADTOF, the strongest published open model, on that same material, ours comes out slightly ahead — but ten songs is a small sample, so the honest reading is "no worse than the reference", not a win.

The more useful thing that test tells me is which voices to trust. In order, best to worst:

That order is not arbitrary, and the reason is worth knowing. Nobody uploads a drum stem — an MP3 is a finished mix, so before anything else the pipeline has to pull the drums out of it. That separation step is nearly free on kick and snare, which are loud and own their part of the spectrum, and expensive on everything quiet, high and short. Every weakness in the list above is really a weakness of that first step. It is also why an open hi-hat sometimes comes back as a crash: once separated, the two genuinely look alike.

I would rather you judged this on a song you know by heart than on any benchmark of mine. Run one you could hum the drum part to, and you will know within a bar whether it is useful to you.

What it cannot do

Stated plainly, because the fastest way to lose a drummer's trust is to oversell:

Your music stays yours

An uploaded file and its results live on a private link for seven days and are then deleted. Nothing is published to a public library, indexed, or shared — result pages are closed to search engines. The single exception is explicit and requires a button press: sending your corrections, which come back to the model as training material.

That last part is the arrangement I would like this to be. Every automatic transcription is wrong somewhere; a drummer who fixes a bar teaches the model something no synthetic dataset can. Details are in the privacy page.

Getting a wrong bar fixed

If it mangles a song you know cold, that is the most useful thing you can tell me, and the specific bar is worth more than a general impression. Write to support@drumanalyzer.com, or press Fixes in the player after correcting it in the grid.

See what it produces before deciding anything. A full song, transcribed, with nothing hand-corrected.

Open the example

The longer version of this story is in Transcribing the drum part is the easy half.