Exclusive invitation-only access. Global public launch coming in 2026.
OpenAI just announced that it recently conducted a small-scale preview of a new tool called Voice Engine.
This is a voice cloning technology that can imitate any speaker by analyzing a 15-second audio sample. The company claims it produces "natural speech" with "emotive and realistic" voices.
Infinite possibilities... and serious risks
This technology, based on the company's existing speech synthesis API, has been in development since 2022. OpenAI already uses a version of the entire toolkit to power the predefined voices available in the current speech synthesis API and the text-to-speech feature. There are several examples on the company's official blog, and they sound eerily close to reality. I encourage you to listen to them and imagine the possibilities, both good and bad.
Potential uses and privacy concerns
OpenAI states that this technology could be useful for reading assistance, language translation, and helping people with sudden or degenerative speech disorders. However, malicious actors would certainly misuse this technology for serious deepfake deceptions, which is already an issue. In this regard, Voice Engine is not quite ready for the general public, as there are serious privacy concerns that need to be addressed before full deployment.
Safety measures and responsible approach
OpenAI acknowledges that this technology presents "serious risks, particularly concerning during an election year." The company states that it incorporates feedback from "American and international partners from diverse backgrounds, including government, media, entertainment, education, civil society, and beyond" to ensure that the product is launched with minimal risks. All preview testers agreed to OpenAI’s usage policies, which prohibit impersonation of another person without their consent or legal right.
Transparency and control over usage
Furthermore, anyone using the technology will be required to disclose to their audience that the voices are generated by AI. OpenAI has implemented safety measures such as digital watermarking to trace the origin of any audio and "proactive monitoring" of system use. When the product is officially launched, there will be a "list of banned voices" that detects and prevents AI-generated voices that are too similar to public figures.
Competitive pricing and future plans
Regarding the timing of deployment, OpenAI remains discreet. TechCrunch has uncovered some potential pricing data, and it appears they are underestimating competitors in the space, such as ElevenLabs. Voice Engine could cost $15 for one million characters, roughly 162,500 words. That's about the length of Stephen King’s "Shining." This seems like an economical way to create an audiobook. Marketing materials also mention a "HD" version that costs double, but the company hasn't detailed how that works.
Huge potential, but challenges to overcome
OpenAI continues to make major announcements this week. It just announced another partnership with its close friend Microsoft to build an AI-based supercomputer called "Stargate." The project apparently costs a staggering $100 billion, according to The Information.