The round saw participation from existing investors Sierra Ventures and 3one4 Capital. The company plans to use the fresh capital to accelerate the deployment of its speech-to-speech technology for business communications.
Founded by Sudarshan Kamath and Akshat Mandloi, Smallest.ai is building a vertically integrated AI stack for real-time voice applications. The company has developed its own speech, language and voice-processing models rather than relying on separate third-party APIs for each stage of a conversation.
Addressing Latency in Voice AI
Traditional voice automation systems typically combine separate speech-to-text, large language model and text-to-speech systems. Data moving between these systems can add latency to a conversation. Delays of more than 200–300 milliseconds can make voice interactions feel less natural. Smallest.ai's architecture combines the conversational processing pipeline within a unified execution layer. The company controls the model stack and infrastructure, allowing it to optimise different stages of voice processing together.
Its proprietary stack consists of four models: Lightning, Pulse, Electron and Hydra. Lightning is the company's text-to-speech model, which can generate 10 seconds of audio in around 100 milliseconds while operating on less than 1 GB of video random-access memory (VRAM).
Pulse is a specialised transcription model that delivers the first transcript in under 300 milliseconds, allowing the voice agent to begin processing an interaction while the conversation is still underway. Electron is a small language model designed for real-time dialogue. Its architecture focuses on delivering reasoning capabilities without relying solely on a higher parameter count, helping reduce response latency.
Hydra is a speech-to-speech model that processes incoming audio waveforms directly instead of first converting them into text. Its asynchronous listening and reasoning architecture allows the system to execute tool calls during an ongoing conversation. The models operate through a unified orchestration layer designed to support conversational features such as pauses, inflection and other elements of spoken interaction.
Infrastructure Built for High-Volume Voice Traffic
Smallest.ai has also developed its own inference infrastructure to handle high-concurrency workloads. Enterprise contact centres can experience sudden increases in call volumes during billing cycles, promotional campaigns and service disruptions, which can place pressure on API-dependent voice systems and lead to rate limits or performance degradation.
The company's infrastructure allows it to manage load queues and provision computing resources based on demand. Smallest.ai also supports on-premise deployment for enterprises that require greater control over their data and infrastructure. The platform has compliance certifications and standards covering HIPAA, GDPR, SOC 2 Type II and ISO 27001.
Enterprise Deployments and Cost Reduction
Smallest.ai has moved from its early research stage to production deployments, with companies including RingCentral and Truecaller using its technology for voice interactions. The company said its platform currently powers millions of voice interactions each month. Its model and infrastructure optimisation has also reduced text-to-speech operating costs from around USD 0.20 per minute to USD 0.01 per minute at scale, according to the company.
3one4 Capital first backed Smallest.ai at its pre-seed stage, when the company was still developing its initial research direction. The Series A investment marks the continued participation of the venture capital firm alongside Seligman Ventures and Sierra Ventures. With the latest funding, Smallest.ai will focus on expanding the deployment of its voice AI infrastructure and developing applications for enterprise communications.
This article was originally
published by the