Artificial intelligence gained its vision first. Now, the senses of speech and hearing are kicking in.
Riding on the high of tools like DALL-E, researchers at Microsoft have announced ‘VALL-E’, a text-to-speech generator that can imitate a person’s voice with a short, three-second audio sample. VALL-E can even mimic a speaker’s emotions and tone with its synthetic voices.
The machine is trained on 60,000 hours of English speech from LibriLight, an audio library collated by Meta. Besides replicating a person’s speech patterns, it can apply the synthesized audio on a text transcript with words that haven’t been uttered by the original speaker, as well as borrow elements from the “acoustic environment” in which the sample audio was set—recreating, say, background variations in a phone call.
Microsoft has shared some of VALL-E’s results on a dedicated website. By the looks or sounds of it, the AI does deliver some pretty convincing and animated audio, though others still come across as robotic. Take note that all these recordings derive from three-second clips; with a richer dataset, VALL-E will likely perform much better.
Understandably, Microsoft isn’t yet making VALL-E’s abilities available to the public, considering how easy the model can be weaponized. Deepfakes are a growing threat that even has the FBI on high alert; the bureau has warned of a rising trend of criminals who are hiding behind different faces and voice-spoofing technology to climb their way up recruitment processes in hopes of stealing confidential information.
“Since VALL-E could synthesize speech that maintains speaker identity, it may carry potential risks in misuse of the model, such as spoofing voice identification or impersonating a specific speaker,” Microsoft acknowledges in its research paper.
However, the team adds that it’s possible to develop a system to detect if an audio clip was generated by VALL-E.