
Smallest.ai, a startup founded in late 2024, has raised a $13 million Series A to build voice AI systems aimed at making automated phone conversations feel less like bot interactions and more like human exchanges. TechCrunch reported that the round was led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital, bringing the company’s total funding to more than $21 million.
The pitch is notable because it cuts against one of the dominant assumptions in generative AI: that better conversational systems mainly come from scaling large language models. According to TechCrunch, Smallest.ai is instead building smaller, specialized voice models for real-time conversation, arguing that phone-based AI needs to listen, think, and speak in parallel rather than wait for a full prompt before generating an answer. If that approach works in production, it could matter for a fast-growing part of enterprise AI: customer-facing voice systems where even brief delays make interactions feel robotic.
According to TechCrunch, Smallest.ai’s core argument is that conventional LLM behavior creates too much latency for live conversation. In text chat, users tolerate a pause while a model processes a prompt and writes back. In a phone call, that same pause can break the flow immediately.
CEO Sudarshan Kamath told TechCrunch that the company is designing its model to behave more like people do in conversation, processing speech continuously and preparing a response while the other person is still talking. The company describes this as a real-time intelligence layer for narrow conversational tasks, especially customer support-style exchanges.
SiliconANGLE’s coverage characterizes the design as an “asynchronous voice AI architecture,” which broadly aligns with the TechCrunch description, although the fuller technical details were not available in the provided source text. Based on TechCrunch’s reporting, the practical idea is that a smaller model handles the rapid back-and-forth, while a larger system is called in only when the conversation moves beyond the smaller model’s domain.
That handoff model is central to the startup’s strategy. TechCrunch reported that if Smallest.ai’s voice system encounters a question outside its limited knowledge base, it passes the query to a larger foundational model and briefly places the caller on hold to gather the answer. Kamath framed that as analogous to how a human support agent might pause to check information rather than improvise.
The funding arrives as investors continue to look for application-layer AI companies with a clearer route to revenue than broad foundation model builders. Voice AI has become a particularly active category because it touches call centers, sales, appointment scheduling, collections, and other business functions with obvious automation budgets.
Smallest.ai is entering that market with a narrower focus than some of the best-known voice startups. TechCrunch reported that the company is concentrating on real-time conversational agents for enterprise use, rather than adjacent voice media use cases such as dubbing or podcast production.
That distinction matters. The performance requirements for a synthetic narration tool are very different from those for an inbound support call. In enterprise telephony, low delay, interruption handling, accent robustness, and noisy-line performance can be more important than purely natural voice quality. TechCrunch said Smallest.ai is focusing on exactly those voice-specific issues, including diverse accents, support for dozens of languages, and operation in noisy environments.
For investors, that focus may look attractive because it targets a use case with measurable business outcomes. Enterprises buying voice automation are usually less interested in novelty than in whether the system can reduce handling time, avoid transfers, improve containment rates, and preserve customer satisfaction. Smallest.ai has not publicly shared such production metrics in the supplied reporting, but the market logic is clear: if its latency advantage is real, that could become a meaningful wedge against larger general-purpose systems.
TechCrunch reported that Smallest.ai’s customers already include RingCentral and Truecaller, both established names in communications and voice-related products. The report does not specify the scope of those relationships, deployment size, or whether the companies are using Smallest.ai in production, pilots, or as infrastructure components. That limits how much can be inferred from the customer list alone, but the names do suggest the startup has at least early validation from companies that understand voice workflows.
Kamath also told TechCrunch that potential customers include customer support software companies such as Sierra and Decagon. His argument, as paraphrased in the report, is that customer service platforms may not want to build their own specialized voice stack if doing so distracts from their main product.
That positioning places Smallest.ai in competition with several layers of the stack. TechCrunch identified ElevenLabs as a voice AI leader and named Cartesia and Sarvam as competitors or adjacent players. Each of those companies brings a different strength: ElevenLabs is strongly associated with synthetic voice quality and broad developer mindshare; Cartesia has pushed low-latency speech infrastructure; Sarvam is linked to regional language capabilities. Smallest.ai appears to be trying to win not by being the broadest voice platform, but by being the best fit for real-time enterprise calls.
The company is also effectively competing with in-house efforts by larger platforms. Enterprises using RingCentral-style communications software or building AI call flows on their own stacks may eventually decide whether to buy a dedicated speech layer, assemble one from multiple vendors, or rely on increasingly multimodal general models.
The strongest confirmed facts in the source material are the financing itself, the participating investors, the company’s founding timeframe, and the names of cited customers. Those details come from TechCrunch’s reporting on the round, with SiliconANGLE separately reporting the raise and describing the company’s architecture in similar terms.
Several of the more ambitious claims remain the startup’s own framing rather than independently verified results. Most notably, Kamath told TechCrunch that the company wants its models to “break the Turing test” in phone calls so users cannot tell whether they are speaking with AI or a human. That is an aspiration, not a demonstrated benchmark in the available evidence.
Likewise, the claim of “virtually zero response lag,” as described by TechCrunch, should be read as a product assertion rather than a published measurement. The supplied sources do not include latency data, word error rates, interruption recovery numbers, multilingual benchmark results, or side-by-side comparisons against ElevenLabs, Cartesia, or other voice AI systems.
The customer references also need caution. While RingCentral and Truecaller are named as customers, the sources do not include contract values, usage volume, renewal data, or direct customer testimonials. For enterprise buyers and builders, that means the real test is still ahead: whether Smallest.ai can move from promising demos and early integrations to durable, high-volume deployments.
For builders, Smallest.ai’s strategy reinforces a broader design shift in AI agents: not every interaction should run through a large general model from start to finish. In latency-sensitive applications, a layered system may be more practical, with a small fast model handling interaction management and a larger model reserved for complex reasoning.
That idea could influence how teams architect AI agents across phone support, interactive voice response replacements, outbound sales calls, and industry-specific assistants. The key tradeoff is between speed and breadth. A narrow model can feel more natural in a tight domain, but it may fail more often outside that domain and require graceful escalation. Smallest.ai’s “put the caller on hold and query a larger model” pattern is one possible solution, though it introduces complexity around orchestration, reliability, and customer experience.
For enterprise AI buyers, the story is less about novelty than procurement criteria. A vendor promising human-like calls must still answer standard operational questions: how often does it interrupt incorrectly, how well does it perform with accents, how does it behave in noisy environments, what are the fallback rules, and what governance exists around disclosure and consent in automated calling. The supplied reports show Smallest.ai emphasizing several of those technical pain points, but not yet providing public proof points.
There is also a strategic angle for customer support platforms such as Sierra and Decagon. If specialized speech layers improve quickly, some software vendors may prefer to integrate external voice infrastructure rather than build proprietary speech systems. But if multimodal foundation models close the latency gap, the window for standalone voice specialists could narrow.
The next signals around Smallest.ai will likely be operational rather than fundraising-related. One will be whether RingCentral or Truecaller, or other named partners, publicly describe specific use cases or measurable outcomes.
Another will be technical disclosure. If Smallest.ai wants to separate itself from ElevenLabs, Cartesia, and Sarvam, it may eventually need to publish clearer evidence on latency, barge-in handling, multilingual quality, and performance in real support environments.
A third issue is product scope. The company’s current case is strongest if enterprises continue to value dedicated voice systems over end-to-end LLM stacks. Any moves by major model providers into lower-latency, call-ready speech interaction could intensify that pressure.
Finally, watch the market’s tolerance for human-like AI calls. As systems get harder to distinguish from people, legal, compliance, and trust questions will become more important for enterprise AI deployments, especially in regulated support and outbound calling contexts.
Smallest.ai is making a credible bet on a real weakness in today’s AI stack: voice conversations punish latency more harshly than chat does. The startup’s emphasis on small, specialized models and selective handoff to larger systems fits a growing reality in AI product design, where orchestration often matters as much as raw model capability.
But this category is unforgiving. Enterprise voice buyers do not just need a convincing demo; they need reliability across accents, interruptions, domain boundaries, and compliance edge cases. Smallest.ai’s raise suggests investors believe that problem is worth solving with dedicated infrastructure. The harder question is whether the company can turn its architectural thesis into reproducible production performance before larger platforms and voice incumbents absorb the same lesson.
Smallest.ai has raised $13M to build low-latency voice AI for enterprise calls, betting small specialized models can make agents sound human.