Modulate: Interview With CEO And Co-Founder Carter Huffman About Audio AI And Voice Intelligence

Modulate is an audio AI company developing technology designed to understand what is happening in voice conversations beyond the transcript. Its flagship platform, Velma, analyzes signals including tone, emotion, cadence, intent, emphasis, synthetic speech, manipulation, and conversational behavior, while its Ensemble Listening Model architecture coordinates more than 100 specialized audio models. Pulse 2.0 interviewed Modulate CEO and Co-Founder Carter Huffman to learn more about the company’s origins, its audio-native AI technology, deepfake detection, and the future of voice intelligence.

Carter Huffman’s Background

Carter Huffman

When asked about his background, Huffman shared:

I’m the CEO and co-founder of Modulate. My co-founder Mike Pappas and I met as physics students at MIT, and after graduating I spent several years at NASA’s Jet Propulsion Laboratory working on onboard intelligence for autonomous spacecraft and low-power machine learning.

That experience shaped a lot of how I think about AI today: how do you build systems that are highly capable and accurate without assuming unlimited compute, memory or power?

That philosophy eventually became very relevant to the way we architected Modulate’s audio-native AI technology.

How Modulate Started

When asked how the idea for the company came together, Huffman explained:

I became fascinated by advances in generative AI around 2014 and 2015. At the time, most of the excitement was around images, and I started asking why similar techniques were not being applied more aggressively to audio.

That led to early work in real-time voice transformation and eventually to the broader question that still drives Modulate today: how can AI truly understand everything being communicated through voice, not just the words?

Favorite Memory

When asked about his favorite memory working for the company so far, Huffman said:

One of the most meaningful parts of building Modulate has been seeing technology that started in gaming expand into areas where it can have a very tangible real-world impact.

We originally built systems capable of understanding nuanced conversations at enormous scale in games such as Call of Duty, Grand Theft Auto Online and Rainbow Six Siege.

Today, that same underlying technology is being used in areas like fraud prevention, healthcare, AI agent supervision and online trust and safety.

Seeing that evolution from a very difficult technical problem into something that can genuinely protect people has been incredibly rewarding.

Core Products And Features

When asked about Modulate’s core products and features, Huffman detailed:

Modulate’s flagship platform is Velma, which is designed to understand what is actually happening in a voice conversation beyond the transcript.

It analyzes signals including tone, emotion, cadence, intent, emphasis, synthetic speech, manipulation and conversational behavior, and it can turn those signals into real-time actions or post-call intelligence.

Underneath Velma is our Ensemble Listening Model architecture, which orchestrates more than 100 specialized audio models.

We also offer individual capabilities such as transcription, deepfake detection and other audio-native APIs that developers can use independently.

Modulate

Balancing Privacy And Safety

When asked about recent challenges in the sector and how Modulate has addressed them, Huffman explained:

A big challenge we face, where we currently see a lot of positive opportunity, is how to balance privacy and safety in Audio AI and voice analysis.

As an example, twelve months ago my grandmother was the victim of a deepfake fraud attack that mimicked my voice. She’s 92 years old and suffers from dementia. She has no way to defend herself against these attacks, but Modulate’s technology could have detected the deepfake and stopped it.

Scams targeting the elderly in the U.S. have resulted in billions of dollars per year stolen from America’s most vulnerable population. Of course we want to deploy our models to stop that.

But at the same time, we can’t ethically or legally send every phone conversation to our cloud service for analysis, even if the result is beneficial and prevents harm.

We face similar tradeoffs when defending hospitals from social manipulation attacks, banks from fraud, and other sensitive but vulnerable groups.

We achieve that balance through strict data residency and controls, flexible deployment capabilities including detecting synthetic voices privately on-device, and similar strategies.

But finding that right balance is both difficult to navigate and critically important to get right.

How The Technology Has Evolved

When asked how Modulate’s technology has evolved since launching, Huffman said:

We initially built Modulate for online gaming, where understanding context was essential to distinguish normal competitive banter from harassment, threats or other harmful behavior.

Solving that problem forced us to understand not only the transcript, but emotion, tone, intent and how a conversation evolves over time.

Over the years, we realized that the same underlying audio intelligence could apply much more broadly, so we expanded into transcription, deepfake detection, fraud prevention, AI agent supervision, healthcare, customer experience and other voice applications.

Key Company Milestones

When asked about some of Modulate’s most significant milestones, Huffman highlighted:

A few milestones stand out.

Our technology has now analyzed more than 600 million hours of audio, and we currently process more than 10 million hours each month.

We have also ranked #1 on public Hugging Face benchmarks for both transcription and deepfake speech detection, while expanding from our original gaming deployments into healthcare, fraud and security, customer service, AI agents and other professional applications.

Most recently, we announced $25 million in new funding, bringing total capital raised to $60 million.

Beyond these numbers, Modulate has been deployed in hospitals, where our technology is used to protect their businesses from social engineering and deepfake attacks.

Additionally, our technology has been used to detect and stop harassment, extremism, child exploitation, and other harms for tens of millions of unique users across the globe.

Customer Success Stories

When asked to share a specific customer success story, Huffman said:

One of the most impactful successes with customers has been preventing social engineering and deepfake attacks on hospitals, which ultimately lets us protect patient data and in some cases save lives.

Social engineering is responsible for a vast number of successful hacks on hospitals and healthcare companies globally, with devastating consequences, stealing sensitive health information from patients and selling it on the Dark Web, deploying ransomware that halts hospital operations, and more.

One of our hospital customers that has deployed Modulate’s Velma model to detect and escalate social engineering attacks on their IT helpdesk found over a dozen incidents in a single month, any one of which can escalate into a full hack within minutes.

Velma finds attacks and escalates them in only a few seconds, stopping the attack and closing off the window of opportunity for the hacker to exfiltrate data or deploy ransomware.

Because of the critical nature of hospital operations, our technology is literally saving lives, alongside protecting sensitive health data and avoiding tens of millions of dollars in damage to hospitals.

Funding And Revenue Metrics

When asked whether Modulate could discuss funding or revenue metrics, Huffman shared:

We recently raised $25 million in new funding led by Future Ventures, with participation from Hyperplane and Lakestar, bringing Modulate’s total funding to $60 million.

The capital is being used to accelerate AI and machine learning research, expand product and engineering, grow our developer ecosystem, deepen partnerships and broaden the models, APIs and deployment options available to customers.

We are not publicly disclosing revenue figures at this time, but the investment comes as we are seeing growing adoption across a wider range of voice AI applications.

Total Addressable Market

When asked about the total addressable market Modulate is pursuing, Huffman explained:

Modulate is a frontier Audio AI company, whose platform was built to support the entire range of Audio and Voice AI models and capabilities needed both today and in the future.

Billions of hours of voice conversations occur over digital channels every day, and Voice AI can be beneficial to every single one.

Today, even individual capabilities within the broader Voice AI umbrella have market sizes within the tens of billions. The overall Voice and Audio AI space is a $1 trillion-plus market and growing.

Competitive Differentiation

When asked what differentiates Modulate from its competition, Huffman explained:

The biggest difference is that most voice AI companies are focused on one slice of the conversation, often transcription or speech generation.

Modulate is focused on understanding the entire conversation: the words, emotion, tone, intent, behavior, synthetic speech and how all of those signals interact.

Our Ensemble Listening Model architecture also lets us combine more than 100 specialized models rather than relying on one enormous foundation model, which gives us strong accuracy while dramatically reducing compute requirements.

It also allows us to add new capabilities quickly without having to retrain an entire frontier model every time.

Modulate

Future Goals

When asked about Modulate’s future goals, Huffman said:

Our biggest goal is to make audio-native intelligence a foundational part of how voice AI is built.

In the near term, that means expanding the capabilities within Velma, adding more specialized models, growing our engineering and AI/ML teams, investing in developer tools and APIs, and supporting more deployment environments including on-premise and on-device use cases.

Longer term, we believe voice will become a much more natural way for people to interact with technology, and for that future to work well, AI needs to understand how people speak, not simply transcribe what they say.

The Future Of Voice AI

When invited to discuss another topic, Huffman concluded:

Voice AI is at an inflection point right now. It still feels a little like text AI did before ChatGPT: the technology is clearly powerful, but interacting with it often still has rough edges.

Over the next few years, I think we will move from typing, clicking and navigating interfaces to simply telling technology what we want it to do using natural speech.

For that to happen, machines need to understand the full meaning of human speech, including tone, emotion, context and intent.

That is the layer we are building at Modulate.