Google's Gemini 3.8 Live models stream voice and visual context, handle interruptions and can keep a conversation moving while tools run. The published audio rate starts at $0.005 per input minute and $0.018 per output minute, before the rest of the call stack.

RELATED BUYER PROFILEReview Gemini 3.8 Live pricing, access and test-pending statusPricing · access · strengths · limitations →RELATED DECISION GUIDERun a same-call voice model pilotWorkflow · evidence · risk · governance →CONTINUE THE DECISIONReview the Extended Thinking product profileEvidence · workflow · next action →CONTINUE THE DECISIONCompare the live model with an existing shortlistEvidence · workflow · next action →CONTINUE THE DECISIONEstimate complete resolved-call costEvidence · workflow · next action →CONTINUE THE DECISIONBuild a controlled customer-support stackEvidence · workflow · next action →
THE BRIEF IN 30 SECONDS

What you need to know

  • Published model codes: gemini-3.8-live and gemini-3.8-live-extended-thinking
  • Standard API audio rates are listed together at $0.005/min input and $0.018/min output
  • The practical decision is completed-call quality and total resolved-task cost—not the model demo or audio rate alone
THE NUMBERS

The verified launch boundary

$0.005published audio input price per minute
$0.018published audio output price per minute
128Kmaximum input context in the model card
64Kmaximum output tokens in the model card
2separate model codes for latency and deeper reasoning
Pendingindependent HubAI same-call test
01
THE CONTEXT

Short answer: two live models, two jobs

Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September and updated its announcement on 17 September. Live is positioned as the default for low-latency voice agents and fluid dialogue. Extended Thinking is for higher-complexity tasks that need more background reasoning and multi-step coordination. Both accept text, images, audio and video and can return text and audio. The model card lists a context window up to 128K and output up to 64K. HubAI has not independently reproduced Google's launch demonstrations or benchmark results.

02
WHY IT MATTERS

What changed inside the conversation

The Live API uses a stateful WebSocket connection and supports interruption, transcripts and tool use. Google says the new models can make asynchronous function calls while continuing to talk, so a user does not always wait in silence for a lookup or action. Extended Thinking can narrate progress while it reasons through a more complex task. That is useful for support, troubleshooting, training and booking flows—but a natural voice does not prove that the action, source or recommendation is correct.

03
WHAT HAPPENS NEXT

The API price is a component, not a completed-call price

Google's pricing page lists both Gemini 3.8 Live variants together. The standard paid rate is $3 per million audio-input tokens, shown as $0.005 per minute, and $12 per million audio-output tokens, shown as $0.018 per minute. Text input is $0.75 per million tokens, image or video input is $1 per million, and text output including thinking tokens is $4.50 per million. A free tier is listed. Google Search grounding can add charges after the shared free allowance. Telephony, media transport, storage, monitoring, tool APIs, failed calls and human escalation are outside the headline audio rate.

04
THE CONTEXT

Language coverage needs a feature-level check

Google's launch announcement says Live automatically detects and transitions between 97 languages. The current Live API overview describes 70 supported conversational languages and 70-plus for live translation. These statements concern different features and do not establish equal speech quality, accent handling or tool completion in every language. A UK buyer should test the exact English accent, language switches and regulated phrases expected in the real call flow rather than treating the highest published number as universal coverage.

05
THE CONTEXT

Access differs by route

Both model codes are available through the Gemini API and Google AI Studio. Google says Gemini 3.8 Live is also rolling into Search Live, while Extended Thinking is rolling into Gemini Live and selected Google AI subscription and Workspace surfaces. Enterprise access begins in private preview, with broader customer-experience availability described as coming soon. Rollout statements are not a guarantee that every UK account, plan or administrator can enable every feature today. Confirm the signed-in product, region, contract and data terms before designing a launch.

06
THE CONTEXT

Known limits still matter

Google's model card says the models can hallucinate and may experience occasional slowness or timeouts; it lists a January 2025 knowledge cutoff. Audio output carries SynthID watermarking, but provenance does not replace consent, disclosure or recording controls. A voice agent can sound confident while using stale knowledge or a failed tool result. Keep critical actions behind explicit confirmation, preserve transcripts and tool logs where lawful, and provide a human route when the model cannot complete the task.

07
THE CONTEXT

HubAI buyer verdict: prove the simpler route first

Pilot Gemini 3.8 Live before Extended Thinking unless the task clearly requires deeper live reasoning. Freeze one call script and score interruption recovery, factual accuracy, task completion, tool-call success, time to human escalation and cost per accepted resolution. Then run the same calls through Extended Thinking and a current alternative such as GPT-Live-1. Upgrade only when the improvement survives real accents, background noise, long sessions and failure recovery. This is a same-task buying recommendation, not a claim that one model is universally best.

08
THE CONTEXT

Evidence limits and editorial disclosure

Independent editorial coverage; not sponsored. Product, price, capability and benchmark statements are attributed to Google and were checked against the launch post, Developer API pricing, Live API documentation and model card on 19 September 2026. The two catalogue profiles are marked Test pending and have no HubAI Score. The visible UK Google Trends top 25 contained no direct AI query during this review, so no trend volume is attached. The cover is an original conceptual illustration, not a product screenshot or Google asset.

HUBAI VIEW

Start with the latency-led Live model. Pay for deeper live reasoning only when Extended Thinking improves accepted task completion after latency, tool failures, escalation and complete cost are counted.

Buyer decision signal: New voice models · API and product rollout
BUYER ACTION PLAN

What to verify next

1Record the exact model code, account route and region used in the pilot

2Test interruption, accent handling, silence and background noise

3Measure successful tool calls and the behaviour after a timeout

4Price telephony, infrastructure, grounding, tools and human escalation

5Require explicit confirmation before consequential actions

6Compare cost per accepted resolution against Live, Extended Thinking and one alternative

Build a side-by-side comparison →
PRIMARY SOURCES

Read the evidence

Capabilities, availability and prices can change. HubAI keeps analysis separate from the underlying official material.

01Google launch announcement — 15 September, updated 17 September 2026Open source ↗02Gemini Developer API pricing — checked 19 September 2026Open source ↗03Gemini Live API architecture and implementation guideOpen source ↗04Google DeepMind Gemini 3.8 Audio model cardOpen source ↗05Gemini 3.8 Live model page and model codeOpen source ↗
Corrections & updates

Published and reviewed by the HubAI Intelligence Desk. Material product or policy changes are recorded with an updated timestamp.

Our editorial standard →