OpenAI Introduces Jukebox AI That Generates Music With Vocals
OpenAI unveils Jukebox, a new machine learning framework that generates raw audio music complete with lyrics and singer voices. While the output shows impressive local coherence, the AI still struggles to replicate larger musical structures like repeating choruses.
OpenAI unveils Jukebox, a new machine learning framework that generates original music using raw audio files. Unlike previous models that rely on symbolic music without voices, Jukebox trains directly on audio recordings to produce songs that closely approximate real artists. The system uses convolutional neural networks to compress the training audio and transformers to generate new compressed audio before converting it back into raw sound.
The new model serves as a significant upgrade to OpenAI's previous music generator, MuseNet. While MuseNet creates multi-instrumental tracks in various styles, it lacks the ability to produce vocals. Jukebox solves this by training on a massive dataset of 1.2 million songs and utilizing lyrics from LyricWiki to accurately recreate the specific voices of famous singers.
Despite the impressive technological leap, Jukebox still faces notable limitations. The generated tracks display local musical coherence and follow traditional chord patterns, but they lack familiar larger structures like repeating choruses. Additionally, the use of copyrighted audio recordings from famous artists like Kanye West and Elvis Presley raises potential legal questions regarding training data permissions.