Skip to content

Gemini model series

Google's proprietary family of multimodal foundation models, from the original Gemini releases through the agent-oriented 3.5 line and the efficiency-focused 3.6 Flash.

Cameron
Jul 21, 20266 min read

he Gemini model series is Google DeepMind's family of proprietary, natively multimodal foundation models. Gemini models accept combinations of text, code, images, audio, and video, and Google deploys them through consumer products, developer APIs, enterprise services, and agent systems. The series has evolved from multimodal understanding and long context toward explicit reasoning, tool use, computer control, and long-running agent workflows.

Google does not publish Gemini model weights, parameter counts, or a complete account of its training data. Public descriptions therefore establish product capabilities, interfaces, evaluations, and safety claims rather than a reproducible model specification.

Model families and names

Gemini version numbers identify generations, while suffixes identify deployment roles. Pro models emphasize capability on difficult tasks. Flash models balance capability, latency, and price for production workloads. Flash-Lite models target high-volume and latency-sensitive work. Nano names models intended for on-device use. Google also publishes specialized variants and systems, including Deep Think reasoning modes, image and audio models, robotics models, and the security-focused Flash Cyber.

The numbering is not a simple size ranking. Gemini 3.6 Flash is based on Gemini 3.5 Flash, while Gemini 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite. A newer decimal version may improve one branch of the family without replacing every model in the preceding branch.

Gemini 1 and 1.5

Gemini 1.0 launched in December 2023 in Ultra, Pro, and Nano variants. Google described it as natively multimodal because it was trained from the beginning across multiple kinds of input rather than combining separately trained text and perception systems after training.

Gemini 1.5 introduced a mixture-of-experts architecture and made long context a defining feature of the series. Gemini 1.5 Pro entered preview with a standard 128,000-token window and an experimental one-million-token window. The longer window allowed one request to contain large document collections, substantial codebases, hours of audio, or extended video.

Gemini 2 and 2.5

Gemini 2.0 moved the series toward what Google called the "agentic era." Gemini 2.0 Flash added native tool use, multimodal input, and experimental image and audio output. Google paired the model with agent prototypes such as Project Astra, the browser-operating Project Mariner, and the coding agent Jules.

Gemini 2.5 made reasoning before answering a standard model capability. Its initial Pro release combined a stronger base model with post-training intended to improve mathematics, science, coding, and context-dependent reasoning. Google said this kind of reasoning would be built directly into later models rather than offered only as a separate experimental mode.

Gemini 3 and 3.1

Gemini 3 launched in November 2025 with Gemini 3 Pro and a Deep Think mode. It combined one-million-token context, multimodal reasoning, coding, and more consistent tool use. Google released it across the Gemini app, Search, AI Studio, Vertex AI, Gemini CLI, and Google Antigravity.

Antigravity is Google's agent-oriented development environment. Its agents operate across an editor, terminal, and browser, plan multi-step work, and return artifacts that expose what they changed. Later 3.1 releases extended the generation with updated Pro, Deep Think, and Flash-Lite models.

Gemini 3.5

Gemini 3.5 began in May 2026 with Gemini 3.5 Flash. Google positioned Flash as an agent and coding workhorse rather than a reduced version of a released Pro model. It supports multimodal input, a one-million-token context window, adjustable reasoning effort, tool use, and computer-use workflows.

The surrounding execution system is part of the release's practical capability. Antigravity can run parallel subagents, scheduled work, and persistent tasks. Managed Agents in the Gemini API provide resumable, isolated Linux environments with retained files and state. The distinction matters because benchmark and product behavior belong to a model-and-harness combination, not to model weights alone.

The 3.5 branch expanded in July 2026:

  • Gemini 3.5 Flash-Lite is a high-throughput model based on 3.1 Flash-Lite. It accepts text, images, audio, and video, has a one-million-token context window and a 64,000-token output limit, and supports configurable reasoning effort. At release, Google priced it at $0.30 per million input tokens and $2.50 per million output tokens. Artificial Analysis measured about 350 output tokens per second and an Intelligence Index score of 36, up from 25 for 3.1 Flash-Lite.
  • Gemini 3.5 Flash Cyber is a specialized cybersecurity model used inside Google's CodeMender agent. CodeMender invokes several model instances to search, validate, and combine vulnerability findings. Because vulnerability discovery is useful for both defense and offense, Google limits the model to governments and trusted partners through a pilot program.
  • Gemini 3.5 Pro was announced in May with a planned June release. As of July 21, Google described it as testing with partners and did not provide a new public release date.

Gemini 3.6

Gemini 3.6 Flash launched on July 21, 2026 as an efficiency-focused successor to 3.5 Flash. The model card describes it as based on 3.5 Flash, with multimodal input, a one-million-token context window, a 64,000-token text-output limit, and a March 2026 knowledge cutoff.

Google reported higher results than 3.5 Flash on long-horizon software engineering, machine-learning engineering, computer use, and knowledge-work evaluations. It also reported fewer output tokens, reasoning steps, and tool calls per completed workflow. The API launched at $1.50 per million input tokens and $7.50 per million output tokens, compared with $9.00 per million output tokens for 3.5 Flash.

Independent pre-release testing by Artificial Analysis measured an Intelligence Index score of 50, equal to 3.5 Flash, while output speed rose to about 304 tokens per second and average task time fell by more than half. The release therefore documents a change in production efficiency without a corresponding increase in that aggregate intelligence score. Specific agent evaluations did improve, so the result also illustrates why capability, latency, token use, and cost per completed task need separate measurements.

The same announcement said that Gemini 4 had entered Google's most ambitious pre-training run to date. That statement describes work in progress, not a released capability.

Deployment and evaluation

Gemini is distributed through several overlapping surfaces:

  • the Gemini app and AI Mode in Search for consumer use;
  • Google AI Studio and the Gemini API for application development;
  • Vertex AI, Gemini Enterprise, and the Gemini Enterprise Agent Platform for organizations;
  • Gemini CLI, Android Studio, and Antigravity for software development and agents.

The model selected, reasoning level, tool access, agent harness, sandbox, and product surface can all change observed behavior. Vendor benchmarks frequently use custom harnesses or tools, while independent evaluators may use different prompts, budgets, and execution environments. Comparisons should therefore identify the exact model version and evaluation date rather than treating "Gemini" as one stable system.

Sources

Did you enjoy this article?

Recommend it — Standard Reader surfaces well-loved writing to more readers across the network.

Across the AtmosphereDiscussions