Get Started Studio Demos Blog Research Talks Community

Datasets

This page contains a list of open datasets that were used to train the Magenta models.

Bach Doodle Dataset

The Bach Doodle Dataset is composed of 21.6 million harmonizations submitted from the Bach Doodle. The dataset contains both metadata about the composition (such as the country of origin and feedback), as well as a MIDI of the user-entered melody and a MIDI of the generated harmonization. The dataset contains about 6 years of total audio.

Groove MIDI Dataset

The Groove MIDI Dataset (GMD) is composed of 13.4 hours of aligned MIDI and (synthesized) audio of human-performed, tempo-aligned expressive drumming. The dataset contains 1,134 MIDI files and nearly 22,000 measures of drumming.

MAESTRO

MAESTRO (MIDI and Audio Edited for Synchronous TRacks and Organization) is a dataset composed of over 200 hours of virtuosic piano performances captured with fine alignment (~3 ms) between note labels and audio waveforms.

NSynth

A large-scale and high-quality audio dataset of annotated musical notes, containing 305,979 musical notes, each with a unique pitch, timbre, and envelope. For 1,006 instruments from commercial sample libraries, we generated four second, monophonic 16kHz audio snippets, referred to as notes, by ranging over every pitch of a standard MIDI piano, as well as five different velocities.