OpenAI has shared initial findings from a limited-scale preview of a model named Voice Engine, designed to craft custom, natural sounding voices using text input and a 15-second audio sample.
Voice Engine was initially developed in late 2022 and has since powered pre-set voices within OpenAI’s text-to-speech API, as well as offerings like ChatGPT Voice and Read Aloud. However, OpenAI has stated in a blog on its site that its approaching wider deployment of the technology with caution, recognising potential risks associated with the misuse of synthetic voices.
Consequently, the company claims that it is aiming to foster discussions on the responsible utilisation of this technology, informed by ongoing dialogues and the outcomes of preliminary tests.
The blog states: ‘We recognise that generating speech that resembles people’s voices has serious risks, which are especially top of mind in an election year. We are engaging with U.S. and international partners from across government, media, entertainment, education, civil society and beyond to ensure we are incorporating their feedback as we build.’
To explore the potential applications of Voice Engine, OpenAI conducted private tests with select partners, with the trials informing the company’s strategies, safeguards, and perspectives on how Voice Engine could be harnessed positively across various sectors. Some potential use-cases OpenAI include:
– Facilitating reading assistance for non-readers and children through emotive voices representing diverse speakers.
– Enabling seamless translation of content, such as videos and podcasts, to reach global audiences.
– Enhancing essential service delivery in remote areas.
– Supporting individuals with speech impairments.
– Assisting patients in recovering their voice, particularly those affected by speech-related conditions.
Recommended reading
- OpenAI Introduces Sora, Its Realistic Text-to-video Model
- OpenAI Claims AI Training “Impossible” Without Copyrighted Material
- OpenAI Faces Lawsuit Over Potential ChatGPT GDPR Violations
Partners involved in testing Voice Engine adhere to usage policies that prohibit unauthorised impersonation and require explicit consent from the original speaker. OpenAI has also implemented safety measures such as ‘watermarking and proactive monitoring’, in a bid to trace and mitigate potential misuse.
OpenAI Continues its Flurry of Technological Reveals
Just last week, OpenAI revealed a series of short films which took advantage of its generative AI video production model, Sora.
The tech’s realistic depictions of mammoths and cityscapes took the internet by storm when it was first released, and studios have already adopted the technology to rapidly take ideas from conceptualisation into creation.
The videos range in subject matter, from advertisements to fake animal documentaries, to music videos.
Creators cite different reasons for adopting the technology, saying its advantages in development speed, as well as budget constraints are already creating fast fans.
“As great as Sora is at generating things that appear real, what excites us is its ability to make things that are totally surreal,” remarked Walter Woodman, a co-creator a “Air Head” which is shared below. “A new era of abstract expressionism.”





