I think this is a great project and very HN. Not sure why the comments are so focused on the deliverable- you learned way more and had a much more interesting experience.

One think I didn't see mentioned in the post- maybe I missed it- how large was the data? How many samples did you use to pretrain and post-train

There is this in the post:

> The final dataset contained a few hundred thousand MIDI files, representing roughly 300 million note events.