Modulate: Audio-Native AI Company Raises $25 Million Led By Future Ventures

Modulate, an audio-native AI company developing models that help machines understand emotion, tone, intent, synthetic speech and other signals in human conversations, has raised $25 million in new funding led by Future Ventures, with participation from Hyperplane and Lakestar. The round brings the Boston-based company’s total funding to $60 million.

The new capital will support AI and machine learning research, product development, engineering, developer relations and partnerships as Modulate expands its models, APIs and deployment options.

Modulate is focused on a growing challenge created by voice AI: understanding a conversation requires more than simply converting speech into text. Transcripts can capture spoken words but may lose information such as tone, emotional state, emphasis and whether a voice is synthetic.

The company’s technology analyzes audio directly to provide those additional signals. Its models are being used across fraud and deepfake detection, AI agent supervision, customer experience, trust and safety and other voice applications.

Modulate’s technology now analyzes more than 10 million hours of audio each month and has processed more than 600 million hours in total. The company also said its transcription and deepfake detection technologies have reached the top position on public Hugging Face benchmarks.

Its flagship Velma platform is designed to recognize emotion, tone, intent, emphasis, synthetic speech and conversational behavior. Those signals can be combined to identify higher-level events such as suspected fraud, harassment, customer dissatisfaction, policy violations or problems involving voice AI agents.

According to Modulate, Velma delivers twice the accuracy of traditional large language models when detecting true positives while generating seven times fewer false positives. The platform can operate in real time, enabling applications to respond to events while conversations are still occurring.

Velma is powered by Modulate’s Ensemble Listening Model, or ELM, architecture. Rather than relying on a single massive foundation model, ELM coordinates more than 100 specialized audio models and combines them depending on the task.

Modulate said that approach has demonstrated as much as 1,000 times greater efficiency than using a single large model, reducing the computing power, energy and memory required to analyze audio at scale.

The company’s transcription API is priced at $0.03 per hour for batch processing, while its deepfake detection technology has achieved 98.9% accuracy on public benchmark data.

Modulate’s technology is already being applied to several emerging voice-related challenges. Its models can help healthcare organizations identify deepfake attacks, enable AI agents to recognize emotion more effectively, detect harassment and child grooming in online conversations, monitor voice agent performance and provide voice masking in higher-risk environments.

The company sees growing opportunities as organizations deploy more conversational AI. Businesses using voice agents increasingly need tools capable of determining whether agents understand customers, respond appropriately and perform as intended.

Deepfake detection is another expanding use case as synthetic voice technology becomes increasingly realistic and accessible. Financial services companies, healthcare organizations and contact centers are among the businesses that could use audio intelligence to identify suspicious calls or synthetic speech.

Modulate plans to use the funding to make those capabilities more accessible to outside developers. The company is developing additional industry models, SDKs and APIs, expanding developer relations and partnerships, and supporting additional deployment environments.

The broader goal is to create an audio intelligence layer developers can integrate into voice agents, security products, communications platforms and other applications without having to build specialized audio models themselves.

With $60 million in total funding and more than 600 million hours of audio already processed, Modulate plans to expand its team and infrastructure as it works to establish audio-native intelligence as a foundational component of the growing voice AI ecosystem.

KEY QUOTES:

“Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript. We’re already using audio-native AI to protect organizations from deepfake attacks, help voice agents understand emotion and respond with more empathy, identify dangerous behavior in online conversations, and monitor whether voice agents are actually performing the way they’re supposed to. Underneath all of that are more than a hundred specialized models working together to understand what’s really happening across audio, with dramatically less cost and compute than traditional large models. Developers shouldn’t have to rebuild the audio intelligence layer every time they create a new voice experience. Our mission is to build the models and infrastructure that let them focus on the application they want to create. The opportunity facing audio-native AI is expanding incredibly quickly. We’ve built the technology and proven it at scale, and this investment lets us grow the team and move faster to meet that demand.”

Carter Huffman, CEO and Co-Founder of Modulate

“Modulate has gained a significant technical lead in audio-native AI, and the market opportunity is expanding quickly. The team has proven these models in some of the most demanding voice environments in the world, and we’re now seeing the need for that technology to expand well beyond where it started into AI agents, security, customer experience, and more. This investment will help Modulate move faster, grow the team, put its models into the hands of more developers and partners, and establish audio intelligence as a foundational layer of the AI stack.”

Steve Jurvetson, Co-Founder of Future Ventures and Board Member of SpaceXAI