The results strike me as comparable or worse than you could get with a Markov model. I think it reveals the gap in understanding between LLMs and music. I think you need to either: - Set up a pipeline to decompose music into, say, harmonic sequences and melodic sequences, and then have the LLM work on some more fundamental or more high-level layer of musical composition and then re-translate it back into actual sounds. - Develop a better dataset and train the LLM more natively on musical examples.

Does anybody know of a project that has produced more convincing results?

There is no LLM in the ML pipeline provided by the OP.