Site navigation

Glasgow Uni Research Aims to Boost New Voice Recognition Tech

Thom Carter

,

glasgow uni research aims to boost new voice recognition tech
Recently undertaken analysis of physical speech processes could create new applications for voice recognition technologies, researchers say.

Led by engineers and physicists from the University of Glasgow, the research scrutinised the internal and external muscle movements of volunteers as they talked using a range of wireless sensing devices.

The data from 400 minutes of analysis is being made openly available to other researchers in a bid to aid the development of new technologies based on speech recognition.

The future technologies, it’s been suggested, could help people with speech impairments or voice loss by using sensors to read their lips and facial movements and provide them a synthesised voice.

The dataset could also enable voice-controlled devices like smartphones to read users’ lips as they speak silently, enabling silent speech recognition.

Further, it could help improve security for banking by analysing users’ distinctive facial movements before unlocking sensitive information.

To gather the data, the researchers asked 20 volunteers to speak a series of vowel sounds, words, and sentences while scans of their facial movements and recordings of their voices were collected.

The team used two different radar technologies to image the movement of the volunteers’ facial skin as they spoke, along with the movements of their tongue and larynx.

Vibrations on the surface of their skin was scanned via camera and a laser speckle detection system. A separate camera capable of measuring depth read their mouths as they shaped different sounds.

The University of Glasgow researchers collaborated with colleagues at the University of Dundee and University College London to synchronise and compile the dataset. It’s been named “RVTALL” for the radio frequency, visual, text, audio, laser and lip landmark information it contains.

Professor Abbasi, of the University of Glasgow’s James Watt School of Engineering, said: “This type of multi-modal sensing for speech recognition is still a relatively new field of research, and our review of existing public data found that there wasn’t much available to help support future developments.

“What we set out to do in collecting the RVTALL dataset was create a much more complete set of analyses of the visible and invisible processes which create speech to enable new research breakthroughs, and we’re pleased that we’re now able to share it.”


Recommended reading


Professor Muhammad Imran, leader of the University of Glasgow’s Communications, Sensing and Imaging group, is a co-author of the paper. He commented: “Contactless sensing has huge potential for improving speech recognition and creating new applications in communications, healthcare and digital security.

“We’re keen to explore in our own research group here at the University of Glasgow how we can build on previous breakthroughs in lip-reading using multi-modal sensors and find new uses everywhere from homes to hospitals.”

The researchers discussed how they conducted their multi-modal analysis of speech formation in the Scientific Data journal.

Funding from the Engineering and Physical Sciences Research Council and the Royal Society of Edinburgh supported the research.

Thom Carter

Staff Writer, DIGIT

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data