Google's Gemini 3.8 Live models stream voice and visual context, handle interruptions and can keep a conversation moving while tools run. The published audio rate starts at $0.005 per input minute and $0.018 per output minute, before the rest of the call stack.
RELATED BUYER PROFILEReview Gemini 3.8 Live pricing, access and test-pending statusPricing · access · strengths · limitations →RELATED DECISION GUIDERun a same-call voice model pilotWorkflow · evidence · risk · governance →CONTINUE THE DECISIONReview the Extended Thinking product profileEvidence · workflow · next action →CONTINUE THE DECISIONCompare the live model with an existing shortlistEvidence · workflow · next action →CONTINUE THE DECISIONEstimate complete resolved-call costEvidence · workflow · next action →CONTINUE THE DECISIONBuild a controlled customer-support stackEvidence · workflow · next action →What you need to know
- Published model codes: gemini-3.8-live and gemini-3.8-live-extended-thinking
- Standard API audio rates are listed together at $0.005/min input and $0.018/min output
- The practical decision is completed-call quality and total resolved-task cost—not the model demo or audio rate alone
The verified launch boundary
Short answer: two live models, two jobs
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September and updated its announcement on 17 September. Live is positioned as the default for low-latency voice agents and fluid dialogue. Extended Thinking is for higher-complexity tasks that need more background reasoning and multi-step coordination. Both accept text, images, audio and video and can return text and audio. The model card lists a context window up to 128K and output up to 64K. HubAI has not independently reproduced Google's launch demonstrations or benchmark results.
What changed inside the conversation
The Live API uses a stateful WebSocket connection and supports interruption, transcripts and tool use. Google says the new models can make asynchronous function calls while continuing to talk, so a user does not always wait in silence for a lookup or action. Extended Thinking can narrate progress while it reasons through a more complex task. That is useful for support, troubleshooting, training and booking flows—but a natural voice does not prove that the action, source or recommendation is correct.
The API price is a component, not a completed-call price
Google's pricing page lists both Gemini 3.8 Live variants together. The standard paid rate is $3 per million audio-input tokens, shown as $0.005 per minute, and $12 per million audio-output tokens, shown as $0.018 per minute. Text input is $0.75 per million tokens, image or video input is $1 per million, and text output including thinking tokens is $4.50 per million. A free tier is listed. Google Search grounding can add charges after the shared free allowance. Telephony, media transport, storage, monitoring, tool APIs, failed calls and human escalation are outside the headline audio rate.
Language coverage needs a feature-level check
Google's launch announcement says Live automatically detects and transitions between 97 languages. The current Live API overview describes 70 supported conversational languages and 70-plus for live translation. These statements concern different features and do not establish equal speech quality, accent handling or tool completion in every language. A UK buyer should test the exact English accent, language switches and regulated phrases expected in the real call flow rather than treating the highest published number as universal coverage.
Access differs by route
Both model codes are available through the Gemini API and Google AI Studio. Google says Gemini 3.8 Live is also rolling into Search Live, while Extended Thinking is rolling into Gemini Live and selected Google AI subscription and Workspace surfaces. Enterprise access begins in private preview, with broader customer-experience availability described as coming soon. Rollout statements are not a guarantee that every UK account, plan or administrator can enable every feature today. Confirm the signed-in product, region, contract and data terms before designing a launch.
Known limits still matter
Google's model card says the models can hallucinate and may experience occasional slowness or timeouts; it lists a January 2025 knowledge cutoff. Audio output carries SynthID watermarking, but provenance does not replace consent, disclosure or recording controls. A voice agent can sound confident while using stale knowledge or a failed tool result. Keep critical actions behind explicit confirmation, preserve transcripts and tool logs where lawful, and provide a human route when the model cannot complete the task.
HubAI buyer verdict: prove the simpler route first
Pilot Gemini 3.8 Live before Extended Thinking unless the task clearly requires deeper live reasoning. Freeze one call script and score interruption recovery, factual accuracy, task completion, tool-call success, time to human escalation and cost per accepted resolution. Then run the same calls through Extended Thinking and a current alternative such as GPT-Live-1. Upgrade only when the improvement survives real accents, background noise, long sessions and failure recovery. This is a same-task buying recommendation, not a claim that one model is universally best.
Evidence limits and editorial disclosure
Independent editorial coverage; not sponsored. Product, price, capability and benchmark statements are attributed to Google and were checked against the launch post, Developer API pricing, Live API documentation and model card on 19 September 2026. The two catalogue profiles are marked Test pending and have no HubAI Score. The visible UK Google Trends top 25 contained no direct AI query during this review, so no trend volume is attached. The cover is an original conceptual illustration, not a product screenshot or Google asset.
HUBAI VIEWStart with the latency-led Live model. Pay for deeper live reasoning only when Extended Thinking improves accepted task completion after latency, tool failures, escalation and complete cost are counted.
Buyer decision signal: New voice models · API and product rollout
What to verify next
1Record the exact model code, account route and region used in the pilot
2Test interruption, accent handling, silence and background noise
3Measure successful tool calls and the behaviour after a timeout
4Price telephony, infrastructure, grounding, tools and human escalation
5Require explicit confirmation before consequential actions
6Compare cost per accepted resolution against Live, Extended Thinking and one alternative
Read the evidence
Capabilities, availability and prices can change. HubAI keeps analysis separate from the underlying official material.
01Google launch announcement — 15 September, updated 17 September 2026Open source ↗02Gemini Developer API pricing — checked 19 September 2026Open source ↗03Gemini Live API architecture and implementation guideOpen source ↗04Google DeepMind Gemini 3.8 Audio model cardOpen source ↗05Gemini 3.8 Live model page and model codeOpen source ↗
Products
Products
Products
Products