Model launches

Google's Gemini 3.6 Flash improves multimodal features but not reasoning score

Google released Gemini 3.6 Flash, extending multimodal capabilities (text, images, video, audio, PDFs), a one‑million‑token context window, and faster, cheaper responses.

Google's Gemini 3.6 Flash improves multimodal features but not reasoning score

Google has launched Gemini 3.6 Flash, extending the model's multimodal capabilities and operational performance. The update adds support for text, images, video, audio and PDFs, introduces a one‑million‑token context window, and includes functionality for the model to use a computer. Google also reports faster responses, lower cost, and less verbose outputs.

What was introduced

  • Input types: text, images, video, audio, PDF.
  • Context window: one million tokens.
  • New capability: model-level computer usage features.
  • Operational changes: faster latency, reduced cost, and more concise outputs.

Independent evaluation found no intelligence gain

Artificial Analysis measured Gemini 3.6 Flash and reported an overall intelligence score of 50, the same value assigned to Gemini 3.5 Flash. According to that assessment, the upgrade did not raise the model's reasoning or judgment capabilities despite the broader multimodal support.

Competitive landscape and market context

Observers pointed out that competing vendors are emphasizing different priorities. OpenAI's GPT-5.6 is positioned to compete on reasoning, coding, and delivering completed work, and the company has framed native multimodality as a baseline capability. Kimi's K3 similarly labeled native multimodality as "table stakes" and focused its launch on agentic execution.

Anthropic provides a contrasting approach: Claude accepts visual inputs but produces only text outputs, while Fable 5 is ranked highly on the Intelligence Index for language, coding, judgment, and long‑horizon autonomy.

Why this matters

The Gemini 3.6 Flash release highlights a split in industry priorities. Multimodality delivers impressive demos and broadened input handling, but many stakeholders argue that improvements in language understanding, reasoning, and autonomous execution are more impactful for real work. According to Artificial Analysis's scoring, Google prioritized polishing multimodal inputs and operational metrics rather than increasing the underlying intelligence score.

Brief takeaways

Gemini 3.6 Flash expands multimodal input types and operational features (including a one‑million‑token context window and computer‑use capabilities), and it is reportedly faster and cheaper with shorter outputs. However, an independent evaluation finds its overall intelligence score unchanged at 50, matching Gemini 3.5 Flash. Competitors are placing heavier emphasis on reasoning, coding, and agentic capabilities.