Have you ever noticed that a text message can’t ever really express a feeling? The same goes for an AI. The digital assistants in our lives are like that acquaintance who never quite gets the joke. They answer your questions, sure, but there’s always something… missing. It’s like they’re reading from a script, even when you’re just asking for the weather. The folks at Sesame are working to change that. They’ve made big steps toward bridging that gap between robotic recitations and truly human-like conversation with their new Conversational Speech Model (CSM).

Table of contents
What’s the Big Deal with “Voice Presence?”
Sesame, spearheaded by Brendan Iribe and Ankit Kumar, is on a mission to achieve what they call “voice presence.” What does that mean? Think about your best friend’s voice. You can tell when they’re excited, worried, or just being sarcastic, all from subtle shifts in their tone. That’s what Sesame is trying to capture in AI – that human quality that makes a conversation feel real.
They’re not just trying to make a computer that sounds human, although that’s part of it. They’re aiming for an AI that actually responds to the feeling behind your words. Imagine a digital assistant that can pick up on your frustration and offer a calming tone, or share in your excitement when you tell it good news. That is where the magic is.
Breaking Down the Tech (Without the Jargon)
The current problem is that making an AI actually understand context is tough. Traditional text-to-speech (TTS) systems are pretty good at turning text into spoken words, but they’re not great at understanding the why behind those words.
The Sesame team identified a core issue that older systems lacked:
- There are many ways to say the same sentence.
- Tone changes the meaning.
- The history of the conversation affects things.
The Sesame team tackled this with their Conversational Speech Model (CSM). The CSM is not just making sounds that is speech, it is looking at the whole picture of the communication flow.
Here’s the clever part: CSM uses a thing called “multimodal learning.” It does this to figure out the best way to say something. It’s like the AI is watching a video of your conversation, not just reading the transcript.
This new model breaks the mold by:
- Being a one-stop shop: CSM does everything in a single stage, making it faster and more expressive.
- Having a built-in report card: They developed a new way to evaluate how well the AI understands context, because the old tests just weren’t cutting it.
The Nitty-Gritty of the Conversational Speech Model
CSM operates by using transformers (the same type of AI behind tools like ChatGPT). It doesn’t just read the text; it also “listens” to the audio. By analyzing the audio, it can identify the rhythm and subtle emotional cues. It looks at RVQ tokens. The goal is to understand how a human would say it.
The process starts with interleaving text and audio. The first backbone does the predicting of the codebook. The audio decoder handles the reconstruction. The model’s design allows for low-latency responses.
Sounds Cool… But Does It Actually Work?
Sesame has been testing CSM with a huge dataset of audio (around a million hours!). They trained a few different versions, from “Tiny” to “Medium,” to see how well it could handle things like:
- Paralinguistics: Those little “ums” and “ahs” that make speech sound natural.
- Foreign words: Can it handle words from different languages without sounding like a robot trying to pronounce them?
- Emotional context: Can it adapt its tone to match the situation?
- Pronunciation: Can it get tricky words right, even when they’re spelled the same but pronounced differently (like “read” as in “I read a book” vs. “I have read a book”)?
The results are promising. In many areas, the CSM is getting close to human-level performance, especially when it comes to just sounding natural.

Where Conversational Speech Model Still Falls Short (For Now)
Even with all this progress, there’s still work to be done. One test showed that while people couldn’t always tell the difference between Conversational Speech Model (CSM) and a real human voice without context, they could definitely pick out the human when they had the full conversation to listen to. This means the AI is still missing some of those subtle cues that make human conversation flow naturally.
A Look at the Future of Voice AI
Sesame isn’t keeping all this tech to themselves. They’re planning to open-source parts of their work, which is great news for other developers who want to build on their progress.
Some of the hurdles Sesame plans to focus on in the future include:
- Getting the AI to work well with more languages.
- Making it more conversational.
- Helping the AI learn the flow of conversation.
The “Sesame’s new text to voice model is insane” post was right. The model struggles a bit at extreme speeds, but overall, it is groundbreaking. This isn’t just about making our voice assistants sound prettier. It’s about creating AI that can truly connect with us. It is about the future of communication, and this tech is a big step in that direction.
| Latest From Us
- Forget Towers: Verizon and AST SpaceMobile Are Launching Cellular Service From Space

- This $1,600 Graphics Card Can Now Run $30,000 AI Models, Thanks to Huawei

- The Global AI Safety Train Leaves the Station: Is the U.S. Already Too Late?

- The AI Breakthrough That Solves Sparse Data: Meet the Interpolating Neural Network

- The AI Advantage: Why Defenders Must Adopt Claude to Secure Digital Infrastructure


