HomeArtificial IntelligenceOpenAI Halts ChatGPT Voice After Scarlett Johansson Complaints

OpenAI Halts ChatGPT Voice After Scarlett Johansson Complaints

A familiar-sounding voice can become a real product problem

OpenAI has paused “Sky,” one of the voice options in ChatGPT, after users said it sounded similar to Scarlett Johansson. The comparison was especially hard to miss because Johansson played an AI assistant in the film Her, a role that has become a cultural shorthand for the warm, conversational voice many people imagine when they think about artificial intelligence.

OpenAI said Sky was not designed to mimic Johansson, or any specific individual. Still, it chose to pause the option in response to the reaction. That decision matters because the dispute is not merely about whether a voice was intentionally copied. For users, the practical question is often simpler: does this sound enough like a recognizable person that it changes how they experience the product?

Voice is more personal than a text interface. A chatbot’s written answers may be judged for accuracy, usefulness, or tone, but a speaking assistant also creates an immediate impression of identity. People are highly sensitive to cadence, pitch, inflection, and mannerisms. A voice does not need to be an exact replica to prompt association with a celebrity, a fictional character, or someone in a user’s own life.

User comparisons prompted action

Many users described Sky’s voice as bearing a striking resemblance to Johansson, particularly to her AI-assistant portrayal in Her. The reaction put OpenAI in an uncomfortable position. It had presented the voice as one option within ChatGPT, while the public discussion quickly attached a much more specific identity to it.

That gap between product intent and public perception is central to the issue. OpenAI’s explanation addresses how Sky was made: the company says it was not crafted as an imitation. But public trust is shaped by what people hear, not just by the internal process behind a voice selection. When a large number of users make the same association, a company has to consider the association itself as part of the product’s real-world effect.

Pausing Sky was therefore a cautious move, even without a detailed public explanation of the reasons behind it. It avoids treating user discomfort as a minor branding issue and acknowledges that AI voice technology carries questions that ordinary software design does not. A menu label, an icon, or a color scheme can be changed with little consequence. A voice can feel like a person.

How OpenAI says the voices were selected

Sky was one of five distinct voices available through ChatGPT. OpenAI said the voices were chosen through a five-month selection process involving more than 400 submissions from professional voice actors. The company worked with industry experts, including casting directors and talent agencies, and said it aimed to select voices that were diverse, natural, and appealing to a global audience.

That detail is important because it places the controversy outside the usual image of AI systems simply scraping or reproducing material without consent. OpenAI’s description points instead to a conventional talent-selection process involving professional actors and industry intermediaries. Yet the Sky episode shows that a formal process does not settle every ethical question. A voice can be professionally performed and still evoke a person who was never meant to be invoked.

OpenAI also chose not to disclose the identities of its voice actors, citing their privacy. That choice is understandable, particularly in an environment where a voice actor could face unwanted attention once their work becomes attached to a widely used AI product. At the same time, anonymity creates a difficult balance. It protects performers, but it also means the public has limited visibility into who created the voice and how the company assessed potential resemblance to well-known people.

There is no easy formula for resolving that tension. Naming performers could expose them to scrutiny they did not seek. Keeping them unnamed can make a company’s assurances harder for outsiders to evaluate. The useful standard is not total disclosure; it is a process that takes resemblance, consent, and foreseeable public reaction seriously before a voice reaches a major product.

GPT-4o raises the stakes for voice design

The dispute arrived as OpenAI was presenting GPT-4o, the latest iteration of its chatbot, as a major step in how people can interact with AI. GPT-4o can process audio, video, and text inputs in real time. It can respond promptly to audio prompts and detect emotions in voices.

Those capabilities make voice less of a decorative feature and more of an interface in its own right. A system that can listen, answer, and react to vocal emotion invites a more natural style of conversation than a traditional text chatbot. That may be useful, but it also makes design choices around voice far more consequential. The more lifelike the interaction feels, the more users will read personality, intent, and identity into it.

OpenAI had not yet made the voice conversation feature publicly available. Sam Altman, OpenAI’s CEO, said it would be introduced gradually over the next few weeks, initially exclusively to ChatGPT Plus subscribers. The gradual release is sensible in this context. Voice interaction changes the emotional texture of a chatbot, and early feedback can expose concerns that may not appear in a product demonstration or an internal review.

The Sky pause is a reminder that technical capability and public readiness are not the same thing. Fast audio responses and emotional detection may make an assistant seem more attentive. They can also make the relationship feel more intimate than users expect from software. Companies building these systems need to think not only about whether the technology works, but about the social signals it sends when it does.

The ethical issue is broader than one voice

OpenAI did not provide specific reasons for pausing Sky, beyond acknowledging the concerns raised by users and acting on them. That restraint leaves unanswered questions, but it also avoids pretending that the matter can be reduced to a single technical test. Similarity in a voice is not always a clean yes-or-no determination. It can be shaped by context, cultural memory, and the way a product is marketed or discussed.

The episode highlights a wider challenge for AI voice technology: consent is necessary, but it may not be sufficient. Companies also need to consider recognizability, the risk of unintended association, and how a synthetic or recorded voice will be received once it is placed inside a conversational machine. The consequences are not limited to celebrities. As voice tools become more common, ordinary people may also worry about voices that resemble their own or those of people they know.

OpenAI’s decision does not prove that every familiar-sounding AI voice is improper, nor does its statement establish that Sky was intended to imitate Johansson. What it does show is that companies cannot dismiss resemblance concerns simply by pointing to their intentions. Responsible development requires a willingness to respond when the product’s public meaning differs from the one its creators expected.

That is the standard OpenAI now has to meet as ChatGPT’s voice features expand. Protecting voice actors’ privacy, using professional selection processes, and listening to users are all meaningful steps. But trust will depend on what happens before and after a problem becomes public: how carefully voices are assessed, how clearly concerns are explained, and how quickly a company acts when an AI assistant begins to sound like someone it was never supposed to be.

More: Technology News

Yasir Khursheed
Yasir Khursheedhttps://www.squaredtech.co/
Meet Yasir Khursheed, a VP Solutions expert in Digital Transformation, boosting revenue with tech innovations. A tech enthusiast driving digital success globally.
RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular