Google DeepMind launches Gemini 3.8 Live audio models with low API pricing, challenging OpenAI’s GPT-Live-1 on cost, quality and developer adoption.

Google DeepMind has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two audio models designed for developers building conversational voice applications. Available through the Gemini API and Google AI Studio, the models give Google a new entry in the competition with OpenAI’s GPT-Live-1 while setting a substantially lower published price for voice conversations.
The launch matters because real-time speech systems are moving beyond transcription and basic chat. According to The Decoder, Gemini 3.8 Live can support voice agents that make API calls in the background, process visual input and continue speaking during those operations. Google also says the models support more than 97 languages, although the source evidence does not provide further detail on language-by-language performance or availability.
Gemini 3.8 Live is positioned as the standard real-time audio model, while Gemini 3.8 Live Extended Thinking adds a reasoning-oriented variant. Both are available to developers rather than being presented only as consumer-facing features. The release includes sample applications on GitHub, according to The Decoder, giving product teams a starting point for testing voice workflows.
The feature set points toward applications that combine speech with tool use and multimodal context. A voice agent could, for example, interpret visual input while handling an external API request and maintaining a spoken interaction. The source does not identify specific enterprise integrations, production customers or general availability conditions beyond access through the Gemini API and Google AI Studio.
That distinction is important for builders. A model that can speak naturally is only one part of a voice product. Developers also need reliable tool execution, interruption handling, latency controls, logging and safeguards around actions triggered from spoken requests. Google’s described capabilities address some of that architecture, but the available evidence does not establish how consistently the models perform in production environments.
The clearest competitive difference is cost. The Decoder reports that Google charges $0.005 per minute for audio input and $0.018 per minute for audio output. Based on the publication’s calculation, an hour of voice conversation costs about $1.38 under Google’s pricing assumptions.
The same report lists OpenAI’s GPT-Live-1 at $0.05 per minute, putting an hour of conversation at roughly $3 or more. The exact total for any application will depend on the ratio of input to output audio, pauses, retries, tool calls and other billable activity, so the hourly comparison should be treated as an illustrative estimate rather than a universal operating cost.
Even with that qualification, the pricing gap could affect how teams design voice products. Lower inference costs make it easier to test longer conversations, offer voice features to more users and run experiments before a business model is proven. For startups, cost can determine whether a voice workflow is viable at all. For larger companies, it can influence whether speech is offered as a primary interface or as an expensive add-on.
Price alone does not settle the choice. The Decoder says OpenAI’s model should sound more natural because GPT-Live-1 uses full duplex, allowing it to listen and speak at the same time. The publication also judged OpenAI’s demonstrations to sound better, while characterizing Google’s apparent trade-off as stronger price efficiency rather than clearly superior conversational quality.
The strongest performance claim in the report comes from the Artificial Analysis Speech-to-Speech Leaderboard. The Decoder says Gemini 3.8 Live Extended Thinking scored 82.6 percent and ranked first ahead of OpenAI’s latest GPT-Live-1 models.
That is a benchmark result, not proof that Google’s model will produce the best experience in every application. Leaderboard scores can depend on test design, prompts, evaluation criteria and model versions. The source evidence does not provide the leaderboard’s methodology, confidence intervals or a breakdown of where the models gained or lost points.
The report’s assessment of conversational naturalness is also based on listening to demonstrations. It is useful market context, but it is not a controlled evaluation of latency, interruption recovery, factual accuracy, tool-use reliability or user satisfaction. Neither the source nor the available evidence provides independent adoption figures, customer references or production reliability data for Gemini 3.8 Live.
For teams evaluating the release, the practical test is likely to be task-level performance rather than a single speech-to-speech score. Builders should compare response latency, turn-taking, barge-in behavior, pronunciation, multilingual consistency and the cost of conversations with different speaking patterns. They should also test whether the model correctly handles visual context and API actions without creating unsafe or confusing interruptions.
Google’s strategy is straightforward: make real-time voice development cheaper while offering enough multimodal and tool-use capability to attract developers already experimenting with AI agents. Access through the Gemini API and Google AI Studio lowers the barrier to prototyping, and the GitHub samples may help teams move from a demonstration to an initial application more quickly.
The main opportunity is for products where voice interaction is frequent but does not require the most humanlike conversation possible. Customer support triage, internal assistants, spoken search, field-service tools and hands-free workflows could all benefit from lower audio costs. However, applications involving sales, health, finance or other sensitive conversations may place greater weight on natural turn-taking, predictable behavior, data controls and auditability than on price.
The launch also raises a deployment question for companies choosing between model providers. A lower per-minute rate can be offset by engineering work if the model needs more retries, produces more awkward turns or requires additional orchestration to match a competitor’s experience. Teams should therefore calculate total cost per completed task, not just audio token or minute pricing.
The first signal will be independent testing of Gemini 3.8 Live against GPT-Live-1 across latency, interruptions, visual reasoning, multilingual speech and tool execution. The Artificial Analysis ranking provides an initial reference point, but broader evaluations will show whether the result holds outside its benchmark setting.
Developers should also watch for evidence of real production use: named customers, public case studies, usage data and clearer service limits. Google’s pricing may attract experimentation, but sustained adoption will depend on reliability, availability and the quality of the surrounding API tooling.
Finally, model updates from both companies could quickly change the comparison. OpenAI may respond with lower pricing or improved access, while Google may prioritize more natural turn-taking. The competitive question is not simply which model is cheapest today, but which provider can deliver acceptable voice quality and dependable agent behavior at scale.
Gemini 3.8 Live gives developers a credible reason to reconsider the economics of voice AI. Google’s reported pricing advantage is large enough to change product experiments, particularly for applications built around frequent or extended conversations.
But the release does not eliminate the central trade-off between cost and interaction quality. The most useful evaluation will combine the reported benchmark with real workflows, where latency, interruptions, tool calls and safety matter as much as a leaderboard score. For now, Google has made the cost case; developers still need to establish whether the experience is strong enough for their users.