Last month, Google made its artificial intelligence text-to-music generator available to the public, allowing anyone to turn simple prompts into melodies tailored to specific situations and environments.
Now, Meta has followed suit by open-sourcing its own music generator, dubbed ‘MusicGen’. In short, the program can take a description, such as “90s r&b with a more prominent bass,” and turn it into an estimated 12 seconds of audio output.
Intriguingly, instead of just relying on text, users can choose to upload a reference audio clip, such as an existing song or classical piece, and the algorithm will attempt to mimic both the written prompt and the tune.
According to Meta, MusicGen was trained on a database of 20,000 hours of licensed music, including 10,000 “high-quality” songs and 390,000 instrumentals taken from stock media libraries Shutterstock and Pond5.
We present MusicGen: A simple and controllable music generation model. MusicGen can be prompted by both text and melody. We release code (MIT) and models (CC-BY NC) for open research, reproducibility, and for the music community: https://t.co/OkYjL4xDN7pic.twitter.com/h1l4LGzYgf
Don’t expect the generator to whip up songs as quickly as chatbots come up with answers. As Engadget noted, it took an average of nearly three minutes to churn out the short piece of music.
While it doesn’t seem producers will be out of jobs anytime soon, these AI-powered tools are certainly pushing the boundaries of what technology can achieve in creative fields. The question that remains, then, is if generated tracks sound anything as good as human creations.
Head here to try out MusicGen. In order to use the software, users require a GPU. 16GB of memory or more is recommended, but those with smaller GPUs will still be able to generate shorter sequences of music depending on the model chosen.