Site navigation

OpenAI Shares a Preview of its Sythetic ‘Voice Engine’

Graham Turner

,

OpenAI voice engine
The company claims that is aiming to foster discussions on the responsible utilisation of this technology, informed by ongoing dialogues and the outcomes of preliminary tests.

OpenAI has shared initial findings from a limited-scale preview of a model named Voice Engine, designed to craft custom, natural sounding voices using text input and a 15-second audio sample.

Voice Engine was initially developed in late 2022 and has since powered pre-set voices within OpenAI’s text-to-speech API, as well as offerings like ChatGPT Voice and Read Aloud. However, OpenAI has stated in a blog on its site that its approaching wider deployment of the technology with caution, recognising potential risks associated with the misuse of synthetic voices.

Consequently, the company claims that it is aiming to foster discussions on the responsible utilisation of this technology, informed by ongoing dialogues and the outcomes of preliminary tests.

The blog states: ‘We recognise that generating speech that resembles people’s voices has serious risks, which are especially top of mind in an election year. We are engaging with U.S. and international partners from across government, media, entertainment, education, civil society and beyond to ensure we are incorporating their feedback as we build.’

To explore the potential applications of Voice Engine, OpenAI conducted private tests with select partners, with the trials informing the company’s strategies, safeguards, and perspectives on how Voice Engine could be harnessed positively across various sectors. Some potential use-cases OpenAI include:

– Facilitating reading assistance for non-readers and children through emotive voices representing diverse speakers.
– Enabling seamless translation of content, such as videos and podcasts, to reach global audiences.
– Enhancing essential service delivery in remote areas.
– Supporting individuals with speech impairments.
– Assisting patients in recovering their voice, particularly those affected by speech-related conditions.


Recommended reading


Partners involved in testing Voice Engine adhere to usage policies that prohibit unauthorised impersonation and require explicit consent from the original speaker. OpenAI has also implemented safety measures such as ‘watermarking and proactive monitoring’, in a bid to trace and mitigate potential misuse.

OpenAI Continues its Flurry of Technological Reveals

Just last week, OpenAI revealed a series of short films which took advantage of its generative AI video production model, Sora.

The tech’s realistic depictions of mammoths and cityscapes took the internet by storm when it was first released, and studios have already adopted the technology to rapidly take ideas from conceptualisation into creation.

The videos range in subject matter, from advertisements to fake animal documentaries, to music videos.

Creators cite different reasons for adopting the technology, saying its advantages in development speed, as well as budget constraints are already creating fast fans.

“As great as Sora is at generating things that appear real, what excites us is its ability to make things that are totally surreal,” remarked Walter Woodman, a co-creator a “Air Head” which is shared below. “A new era of abstract expressionism.”

Graham Turner

Sub Editor

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data