I'm in the deep learning music scene, which is due for its stable diffusion moment in the next year or two. The (primarily) timbre transfer system called RAVE is where I'm starting, and my contribution is to optimize the system to improve training time.
1- A lot of options to influence/edit the generated output. What AUTOMATIC1111 is building for Stable Diffusion is the right direction (open extensions and manoeuvrability). I hope harmonai will get the same treatment.
2- Smaller minimum training requirements and simple retraining workflows. AudioLM is outstanding in this regard (but fails the first point)
3- Prod-level quality for end-user tools : Style/tone transfer and cloning plugins like DDSP-VST, mawf and yours (RAVE) sounds at best like DALLE-1 level quality. do you think we could make a DALLE-2 kind of jump soon?
And we do indeed lack public gathering places for this scene ! (or do we ?)
I'd love to get connected. I've been working on an online DAW interface to make trying out these models more approachable. It also currently contains many useful heuristic based generative models. https://neptunely.com -- see https://www.youtube.com/watch?v=T93tfPgKhhA for a recent demo of the generative radio station.
This is some really interesting work. Do you happen to have any other good links to see state of the art kind of stuff here? Most of what I happen across seems to be based on MIDI patterns rather than audio.
[] https://github.com/acids-ircam/RAVE/tree/master/rave