Google says Gemini 3.8 Flash is generally available with a 1M-token context window, up to 64k output tokens, tunable thinking and introductory API pricing for agent and software-engineering work.
What you need to know
- Generally available for production
- 1M-token context and tunable thinking
- Introductory price is not the full workload cost
What Google launched
Gemini 3.8 Flash is available through the Gemini API with low, medium and high thinking settings, a one-million-token context window and up to 64,000 output tokens. Google positions it for multi-file engineering, tool execution and resilient multi-step agents.
Why Flash can still become expensive
Per-token pricing is only one variable. A model that thinks longer, calls more tools or emits more output can raise the cost of a completed workflow even when its headline rate is competitive. Teams should record total input, output, retries and human correction time.
The HubAI test
Use ten representative jobs: repository navigation, multi-file change, tool failure recovery, long-context synthesis and one deliberately ambiguous request. Compare completion rate and review burden against the model already in production before migrating.
HUBAI VIEWThe buyer question is total cost per completed task—not whether the model carries the Flash label.
Buyer decision signal: New model · GA
What to verify next
1Production model ID and migration path
2Total tokens per accepted task
3Tool-call reliability
4Regional and data controls
5Introductory pricing end date
Read the evidence
Capabilities, availability and prices can change. HubAI keeps analysis separate from the underlying official material.
01Google AI: Gemini 3.8 Flash guideOpen source ↗
Safety
Products
Products
Regulation