OpenAI Introduces ‘Whisper’ Speech-To-Text Bot That’s Also Multilingual
After the introduction of OpenAI’s ChatGPT, it’s now introducing a new system that provides a speech-to-text function.
Whisper API is a new version of the open-source Whisper system introduced last September. The model is an automatic speech recognition system that can understand multiple languages. The audio is then formatted into M4A, MP3, MP4, MPEG, MPGA, WAV, and WEBM.
In its development, Whisper was fed 680,000 hours of multilingual data from the internet so that it would be able to understand technical jargon, identify background noise from the voice that’s speaking, and even comprehend accents.
According to TechCrunch, Greg Brockman, president and chairman of OpenAI, posits that many companies still have yet to adopt this type of technology as it’s hard to train models to distinguish and pick up on accents and other dialect-related issues.
Still, Whisper has some limitations, one of which is that it is not yet proficient at “next-word” prediction due to OpenAI training it with “noisy data.” As a result, users are advised that the transcript may contain words they may not have spoken. This might result from the transcription software’s attempt to predict the following spoken word while transcribing audio at the same time.
Another thing to note is that certain languages might be more challenging to understand speakers from groups that have yet to be taught properly to the system.
Whisper API is priced at $0.006 per minute of audio if you’re interested.
[via TechCrunch and OpenAI, Photo 218932766 © Flynt | Dreamstime.com]