Don't miss the latest stories
Advertise Newsletter
Network
  • The Creative Finder
  • The Bazaar
  • Deals
  • Status Is Down
Community
  • Sign up / Log in
  • Discussion Forums
  • Calendar of Events
NEW

Follow

Share this

Microsoft
Apps
Artificial Intelligence
Cybersecurity
Future
Innovation
Meta
More
  • Research
  • Technology
  • Cybersecurity
  • Future
  • Innovation
  • Meta
  • Research
  • Technology
MENU
  • Advertise with us
  • Submit tip/feedback
  • Work with us
  • Subscribe to newsletter
  • Subscribe to RSS
Advertise here
Advertisement

Microsoft Creates ‘VALL-E’ Bot That Mimics Your Voice 3 Seconds After Hearing It

By Mikelle Leow, 10 Jan 2023

Subscribe to newsletter
Like us on Facebook

Photo 226102154 © Anton Skavronskiy | Dreamstime.com

 

Artificial intelligence gained its vision first. Now, the senses of speech and hearing are kicking in.


Riding on the high of tools like DALL-E, researchers at Microsoft have announced ‘VALL-E’, a text-to-speech generator that can imitate a person’s voice with a short, three-second audio sample. VALL-E can even mimic a speaker’s emotions and tone with its synthetic voices.

 

The machine is trained on 60,000 hours of English speech from LibriLight, an audio library collated by Meta. Besides replicating a person’s speech patterns, it can apply the synthesized audio on a text transcript with words that haven’t been uttered by the original speaker, as well as borrow elements from the “acoustic environment” in which the sample audio was set—recreating, say, background variations in a phone call.


Microsoft has shared some of VALL-E’s results on a dedicated website. By the looks or sounds of it, the AI does deliver some pretty convincing and animated audio, though others still come across as robotic. Take note that all these recordings derive from three-second clips; with a richer dataset, VALL-E will likely perform much better.

Advertisement
Advertisement

 

Surprised there isn't more chatter around VALL-E

This new model by @Microsoft can generate speech in any voice after only hearing a 3s sample of that voice 🤯

Demo → https://t.co/GgFO6kWKha pic.twitter.com/JY88vf4lYc

— Steven Tey (@steventey) January 9, 2023

 

Understandably, Microsoft isn’t yet making VALL-E’s abilities available to the public, considering how easy the model can be weaponized. Deepfakes are a growing threat that even has the FBI on high alert; the bureau has warned of a rising trend of criminals who are hiding behind different faces and voice-spoofing technology to climb their way up recruitment processes in hopes of stealing confidential information.

 

“Since VALL-E could synthesize speech that maintains speaker identity, it may carry potential risks in misuse of the model, such as spoofing voice identification or impersonating a specific speaker,” Microsoft acknowledges in its research paper.


However, the team adds that it’s possible to develop a system to detect if an audio clip was generated by VALL-E.

 

 

 

[via Ars Technica and Windows Central, cover photo 226102154 © Anton Skavronskiy | Dreamstime.com]

Receive interesting stories like this one in your inbox
Advertise here

More related news

Advertise here
Also check out these recent news
Web Design
Link to news page

When Your Website Goes Down, This Is the Page Customers Meet Instead

2027
Link to news page

2027 Already Has A Color Of The Year And It’s Beginning On ‘Grounded’ Territory

IKEA
Link to news page

IKEA & Xbox Press Play On Furniture & Storage Inspired By The Iconic Controller

Fashion
Link to news page

Vogue Presents ‘United Flags of Fashion’ With Top Designers For All 50 States

Coca-Cola
Link to news page

Coca-Cola Pours Fresh Life Into Its Iconic Branding With Worldwide Redesign