Building VOXES
Giving AI a more representative voice.

Voice is deeply personal. Before we finish a sentence, the way we speak can communicate geography, culture, age, personality and identity. Our voices carry information beyond the words themselves.
As generative voice technology improved, we became increasingly interested in a contradiction. Artificial intelligence was becoming remarkably good at reproducing human speech, but many communities still did not hear themselves represented naturally in the technology.
That observation became part of the idea behind VOXES.
- 01
Voice AI reached a turning point.
Synthetic speech has existed for decades, but for most of that history it was relatively easy to identify. Computer-generated voices sounded mechanical, repetitive and disconnected from the subtleties of normal conversation.
Generative artificial intelligence changed that quickly. Modern text-to-speech systems can produce increasingly natural rhythm, emotion and pronunciation. Speech-to-text can turn conversations into usable information. Voice-to-voice technology can transform one performance while maintaining characteristics such as timing and expression. Conversational systems can combine language models with voice to allow people to interact with software by speaking naturally.
Together, these technologies are changing audio from a production format into an interface.
For businesses, creators and developers, that creates opportunities in advertising, education, accessibility, customer service, media, entertainment and software. A process that once required specialized studios, equipment and significant production time can increasingly become part of a digital workflow.
The technology itself, however, was only part of what interested us.
The more important question was whose voices would become part of it.
- 02
Representation sounds different.
Puerto Rico is a small geographic market with an unusually distinctive linguistic identity. Spanish on the island is influenced by history, migration, the Caribbean, the United States and generations of cultural exchange. Even within Puerto Rico, voices vary by geography, age, social context and individual experience.
That diversity is difficult to reduce to a generic label such as "Spanish."
The same is true throughout Latin America and the Caribbean. A voice from Puerto Rico does not sound like one from Mexico, Argentina, Colombia or Spain, even though the language may be mutually understood.
This matters because voice is not merely a method for transmitting information. It can create familiarity and trust. People recognize when someone sounds like their community, and they also recognize when an accent feels artificial or disconnected from the context.
VOXES begins with Puerto Rico because it is a market and culture we understand. The larger opportunity, however, is not limited to one island. It is to participate in a global voice ecosystem in which more communities can hear themselves represented naturally.
- 03
The technology is only one part of the voice.
As generative voice technology becomes more powerful, another question becomes equally important: who has the right to use a voice?
A human voice can be part of someone's professional identity. Actors, announcers, creators and performers may spend years developing a recognizable sound. Voice cloning changes the relationship between a recording and the person who created it because a voice can potentially continue generating new content long after the original recording session.
For VOXES, this makes consent, licensing and transparency fundamental parts of the product model.
Our approach is based on working with real voice talent through clear agreements that establish how a voice may be used. The objective is to create a licensed voice library where participation is intentional rather than assumed.
Technology should make new things possible without eliminating the rights of the people whose work makes those things possible.
That principle becomes especially important as synthetic media becomes easier to create.
- 04
Build a platform, not just a generator.
Generating a piece of audio is useful, but real work rarely ends with a single generation.
A creator may need to organize multiple projects. A company may need different voices for different brands. An agency may work across clients and languages. A development team may want to integrate voice capabilities directly into an application. A conversational system may combine transcription, language processing and generated speech in the same interaction.
VOXES is therefore being conceived as a voice platform rather than a single-purpose tool.
Text-to-speech, speech-to-text and voice-to-voice form part of that ecosystem. Voice libraries, projects, audio management and future conversational capabilities can extend it further.
The larger goal is to make sophisticated voice technology easier to use without hiding the complexity that matters, particularly around identity, licensing and control.
- 05
Puerto Rico is the starting point, not the limit.
Many technology products begin by trying to address the largest possible market. We are taking a different approach.
Starting from Puerto Rico gives VOXES a specific perspective. We understand the language, the cultural context and the importance of representing voices that global platforms may treat as a relatively small segment.
That specificity can become an advantage.
Products built around real communities can often solve problems that disappear when everything is designed for an abstract global user.
At the same time, the underlying need is much larger. Brands throughout the United States need Spanish-language content. Latin American creators need more voice options. Developers need multilingual capabilities. Global companies increasingly communicate with audiences that do not sound the same.
The opportunity is to begin locally while building technology capable of traveling far beyond the market where the idea began.
- 06
Where VOXES can go next.
Voice AI is evolving too quickly to define the future of the category through one feature.
The opportunity may include media production, advertising, education, accessibility, localization, conversational agents, customer experiences, developer tools and applications we have not yet considered.
Our objective is not to predict every use case.
It is to build a foundation flexible enough to participate as the technology evolves.
VOXES is ultimately an exploration of the relationship between human identity and machine-generated communication.
Artificial intelligence can generate speech.
The more interesting challenge is making sure the people listening can still recognize something human inside it.
Technology can generate a voice.
The challenge is making sure people can recognize themselves in it.
Related technologies
Related verticals