Following the introduction of its “state-of-the-art” text and image generators, Meta is continuing its strides in artificial intelligence with its latest innovation, AudioCraft.
This newly-launched AI tool has the ability to generate high-quality, realistic audio and music from simple text prompts, opening up a world of possibilities for musicians, content creators, and sound designers.
The program comprises three powerful models: MusicGen, AudioGen, and EnCodec. MusicGen, trained with Meta-owned and licensed music, can generate music from text prompts, while AudioGen, trained on public sound effects, creates various audio sounds from text inputs.
Last but not least, the EnCodec decoder has been improved so as to generate higher-quality music with fewer artifacts.
To champion greater openness and collaboration in the industry, Meta is open-sourcing these models, providing access to researchers and practitioners to train their own models using their own datasets.
While generative AI for images, video, and text has gained significant attention, audio generation has lagged behind due to its complexity and lack of accessibility. Could AudioCraft be the answer to this issue?
Perhaps, as it takes the first step of making high-fidelity audio generation more accessible and user-friendly. It simplifies the complicated music-making process by providing an easy-to-use platform for generating audio, streamlining the activity so it’s more approachable for users compared to previous approaches in the field.
Meta envisions AudioCraft as a versatile tool for musicians and sound designers to explore new compositions, generate inspiration, and quickly brainstorm and iterate on their musical creations in innovative ways. The capabilities of AudioCraft extend beyond music, with applications in sound effects generation and compression algorithms.
In fact, Meta believes MusicGen could evolve into a new type of instrument, akin to the impact of synthesizers when they were first introduced.