Voice-controlled assistants like Siri or Alexa rely on spoken commands given by the user. But not all environments are conducive for speaking out loud, so two researchers from Cornell University have created a device that counts on sight, and not sound.
A wearable camera dangles from the user’s neck like a large pendant, and this is able to detect voice commands even if there’s no spoken sound. The device is able to do so by measuring “skin deformation” in the neck and face from under the chin.
The little infrared camera is mounted on a 3D-printed necklace casing before being strung onto a silver chain, facing upwards, a press release details. To ensure it remains stable, there’s a wing on each side, and a coin placed at the bottom.
Named ‘SpeeChin’, the device is purportedly able to recognize some 54 English and 44 Chinese silent commands. Because it’s angled upwards, there are less privacy concerns regarding people in the surrounding environment as opposed to a camera facing the user’s face, which would also capture the people behind them.
“We feel a necklace is a form factor that people are used to, as opposed to ear-mounted devices, which may not be as comfortable,” explains doctoral student Ruidong Zhang, one of the duo.
“As far as silent speech, people may think, ‘I already have a speech recognition device on my phone.’ But you need to vocalize sound for those, and … the person may not be able to vocalize speech.”
It was reported that, after some image-based training, SpeeChin was able to recognize English and Mandarin commands with an average accuracy of 90.5% and 91.6% respectively during an initial test.
However, thissuccess wasn’t replicated when participants were moving or walking due to the differences in walking gait and head movements, though this might change with further development and more training.
SpeeChin and the researchers’ process is detailed in a study published in the Proceedings of the Association of Computing Machinery on Interactive, Mobile, Wearable and Ubiquitous Technologies journal. Zhang will also be presenting the paper at the UbiComp 2022 conference later this year.