Localization Is the Missing Test for Voice AI, Konecta CTO Says as Market Eyes $47.5 Billion
Voice AI agents are projected to grow from $2.4 billion in 2024 to $47.5 billion by 2034, but Konecta's group CTO argues in TechRadar that localization, not translation alone, is the test the industry has still to pass.
According to the TechRadar piece, the risk is that voice AI repeats a problem already familiar from outsourced customer service centers, where differences in language and culture create a disconnect between brand and customer. As voice AI advances, language itself may become less of a barrier, but speaking the same language is not the same as understanding someone, the CTO writes. Customers bring accents, habits and cultural expectations to every interaction, and they expect support to reflect their reality. Localization therefore needs to become a core test of voice AI performance rather than a translation exercise tacked on at the end of development.
The article points to Scotland as an example. A customer might say "aye" rather than "yes", call something small "wee", or talk about "getting the messages" when they mean going shopping. The language is English, but understanding the interaction requires familiarity with how that language is used locally. Cultural expectations around directness and politeness also differ: in some cultures, using titles and surnames is an important sign of respect, while in others, first names are the norm. A voice agent that is too informal can come across as overly familiar, while one that is too formal may feel distant or unnatural. These are not cosmetic details, the CTO argues; they shape an agent's ability to understand intent and respond appropriately, and ultimately determine whether customers trust the experience.
The article acknowledges the advances already made. ElevenLabs has built its business on AI-generated voices that reproduce not just speech but pacing, intonation and emotion, with technology spanning voice cloning, multilingual speech and conversational agents. Actor Matthew McConaughey has backed AI voice technology that can match his performance without requiring him to change the way he sounds. The days when synthetic speech was instantly recognizable as robotic, as with early digital voices such as "Microsoft Sam", are fading, and the line is becoming harder to distinguish. But the most advanced experiences are not universally available, meaning many people may still encounter underdeveloped AI voices. Humans are also good at noticing when something is slightly off, which produces the uncanny valley effect: an unnatural pause, misplaced emphasis or strangely enthusiastic response can break the illusion. A voice can sound perfectly human and still sound culturally out of place.
The conclusion drawn in the piece is that global brands should resist treating voice AI as a single product rolled out everywhere. A frontier model might provide the foundation, but its performance needs to be tested against local data and real customer interactions so it captures what "good" looks like in a given market: the empathy that de-escalates a frustrated customer, the specific terminology a market actually uses, and the intonation that reads as natural. Those qualities, the CTO writes, cannot be inferred by a frontier model alone.