<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Gemini 3.6 Flash on RockB</title><link>https://baeseokjae.github.io/tags/gemini-3.6-flash/</link><description>Recent content in Gemini 3.6 Flash on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 21 Jul 2026 19:02:04 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/gemini-3.6-flash/index.xml" rel="self" type="application/rss+xml"/><item><title>Gemini 3.6 Flash Cyber, 3.5 Flash-Lite, and 3.6 Flash: Google's New Model Family Compared</title><link>https://baeseokjae.github.io/posts/gemini-3-6-flash-cyber-review/</link><pubDate>Tue, 21 Jul 2026 19:02:04 +0000</pubDate><guid>https://baeseokjae.github.io/posts/gemini-3-6-flash-cyber-review/</guid><description>Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on July 21, 2026. Here is how they compare on benchmarks, pricing, and use cases.</description><content:encoded><![CDATA[<p>Google launched three new Flash models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Together, they form a three-tier strategy covering general-purpose workhorse AI, ultra-low-cost high-throughput inference, and specialized cybersecurity applications — each with a 1M token context window and the latest Frontier Safety safeguards.</p>
<h2 id="what-is-googles-new-flash-model-family">What Is Google&rsquo;s New Flash Model Family?</h2>
<p>On July 21, 2026, Google announced a major expansion of its Gemini Flash lineup with three distinct models designed for different segments of the AI market. The new family consists of Gemini 3.6 Flash (the upgraded general-purpose workhorse), Gemini 3.5 Flash-Lite (a cost-optimized high-speed model), and Gemini 3.5 Flash Cyber (a specialized model fine-tuned for cybersecurity applications). Each model shares the 1M token context window and supports text, image, speech, and video input, but they differ dramatically in pricing, speed, benchmark performance, and access restrictions.</p>
<h2 id="how-does-gemini-36-flash-improve-over-35-flash">How Does Gemini 3.6 Flash Improve Over 3.5 Flash?</h2>
<h3 id="token-efficiency-and-pricing">Token Efficiency and Pricing</h3>
<p>Gemini 3.6 Flash delivers a 17% reduction in output token consumption compared to Gemini 3.5 Flash, according to the Artificial Analysis Intelligence Index. This means developers get more done with fewer tokens — fewer unwanted code edits, reduced execution loops, and higher precision in agentic workflows. The pricing is set at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, making it cheaper than 3.5 Flash ($9 per 1M output tokens) while delivering superior performance.</p>
<h3 id="benchmark-performance-gains">Benchmark Performance Gains</h3>
<p>The benchmark improvements from 3.5 Flash to 3.6 Flash are substantial across coding, machine learning, and computer use tasks:</p>
<table>
  <thead>
      <tr>
          <th>Benchmark</th>
          <th>Gemini 3.5 Flash</th>
          <th>Gemini 3.6 Flash</th>
          <th>Improvement</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>DeepSWE</td>
          <td>37%</td>
          <td>49%</td>
          <td>+12.4%</td>
      </tr>
      <tr>
          <td>MLE Bench</td>
          <td>49.7%</td>
          <td>63.9%</td>
          <td>+14.2%</td>
      </tr>
      <tr>
          <td>OSWorld-Verified</td>
          <td>78.4%</td>
          <td>83.0%</td>
          <td>+4.6%</td>
      </tr>
      <tr>
          <td>GDPval-AA v2</td>
          <td>1349</td>
          <td>1421</td>
          <td>+72 points</td>
      </tr>
  </tbody>
</table>
<p>On the Artificial Analysis Intelligence Index, Gemini 3.6 Flash scores 50, ranking 21st out of 187 models. More impressively, it achieves 303.6 output tokens per second — the fastest speed of any model on the index, ranking 1st out of 187. This combination of improved accuracy and top-tier speed makes it an exceptional choice for production agentic workflows where both quality and latency matter.</p>
<h3 id="enhanced-safety-safeguards">Enhanced Safety Safeguards</h3>
<p>Google has implemented enhanced Frontier Safety safeguards for 3.6 Flash, specifically targeting CBRN (chemical, biological, radiological, and nuclear) and cyber offense capabilities. The knowledge cutoff has also been advanced from January 2025 to March 2026, giving the model access to more recent information.</p>
<h3 id="availability">Availability</h3>
<p>Gemini 3.6 Flash is available in the Gemini app, Google Antigravity, AI Studio, and Android Studio. It also introduces computer use as a built-in client-side tool via the Gemini API and Gemini Enterprise, signaling Google&rsquo;s commitment to agentic AI as the next major paradigm.</p>
<h2 id="what-makes-gemini-35-flash-lite-different">What Makes Gemini 3.5 Flash-Lite Different?</h2>
<h3 id="speed-and-pricing">Speed and Pricing</h3>
<p>Gemini 3.5 Flash-Lite is designed for cost-sensitive, high-throughput applications. It delivers 350 output tokens per second, ranking 3rd out of 152 models on the Artificial Analysis speed index. At just $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, it is dramatically cheaper than the full 3.6 Flash while still offering strong performance.</p>
<h3 id="outperforming-3-flash-despite-being-a-lite-model">Outperforming 3 Flash Despite Being a &ldquo;Lite&rdquo; Model</h3>
<p>One of the most surprising findings is that 3.5 Flash-Lite outperforms the previous-generation 3 Flash on several key benchmarks:</p>
<table>
  <thead>
      <tr>
          <th>Benchmark</th>
          <th>3 Flash</th>
          <th>3.5 Flash-Lite</th>
          <th>Improvement</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>SWE-Bench Pro</td>
          <td>49.6%</td>
          <td>54.2%</td>
          <td>+4.6%</td>
      </tr>
      <tr>
          <td>OSWorld-Verified</td>
          <td>65.1%</td>
          <td>74.0%</td>
          <td>+8.9%</td>
      </tr>
      <tr>
          <td>Terminal-Bench 2.1</td>
          <td>31% (3.1 Flash-Lite)</td>
          <td>54%</td>
          <td>+23%</td>
      </tr>
      <tr>
          <td>GDM-MRCR v2</td>
          <td>60.1% (3.1 Flash-Lite)</td>
          <td>72.2%</td>
          <td>+12.1%</td>
      </tr>
  </tbody>
</table>
<p>The comparison to 3.1 Flash-Lite is particularly striking: 3.5 Flash-Lite achieves 54% on Terminal-Bench 2.1 versus just 31% for 3.1 Flash-Lite, and 72.2% on GDM-MRCR v2 versus 60.1%. This means developers migrating from 3.1 Flash-Lite to 3.5 Flash-Lite get substantially better quality at a similar price point.</p>
<p>On the Artificial Analysis Intelligence Index, 3.5 Flash-Lite scores 36, ranking 12th out of 152 models. While this is lower than 3.6 Flash&rsquo;s score of 50, the cost-to-performance ratio is exceptional — at roughly one-fifth the input cost of 3.6 Flash, it makes agentic search, document processing, and large-scale batch inference economically viable.</p>
<h3 id="computer-use-and-multi-level-thinking">Computer Use and Multi-Level Thinking</h3>
<p>3.5 Flash-Lite also supports computer use as a client-side tool and offers multi-level thinking configurations, allowing developers to trade off between speed and reasoning depth depending on the task. It is available in the Gemini app and is rolling out to Google Search, making it the most accessible model in the new family.</p>
<h2 id="what-is-gemini-35-flash-cyber-and-why-does-it-matter">What Is Gemini 3.5 Flash Cyber and Why Does It Matter?</h2>
<h3 id="a-specialized-security-model">A Specialized Security Model</h3>
<p>Gemini 3.5 Flash Cyber is a fine-tuned variant of the Flash architecture specifically designed for cybersecurity applications. It reaches competitive frontier performance on the CyberGym benchmark, a specialized evaluation for cybersecurity AI capabilities. This is not a general-purpose model — it is purpose-built for security operations, vulnerability analysis, and defensive cyber tasks.</p>
<h3 id="codemender-and-the-multi-agent-security-approach">CodeMender and the Multi-Agent Security Approach</h3>
<p>3.5 Flash Cyber powers CodeMender, Google&rsquo;s multi-agent system for cybersecurity. CodeMender uses multiple specialized AI agents working together to identify vulnerabilities, generate patches, and verify fixes. The model&rsquo;s fine-tuning for security-specific tasks means it understands code patterns, attack vectors, and defensive strategies at a level that general-purpose models cannot match.</p>
<h3 id="restricted-access-and-dual-use-considerations">Restricted Access and Dual-Use Considerations</h3>
<p>Access to 3.5 Flash Cyber is initially limited to governments and trusted partners as a limited-access pilot. This restricted distribution model reflects the dual-use nature of advanced cybersecurity AI — the same capabilities that can defend systems can also be used to attack them. Google&rsquo;s approach raises important questions about AI safety versus open access in cybersecurity, a topic that generated over 333 comments on Hacker News within hours of the announcement.</p>
<p>The restricted access model for 3.5 Flash Cyber stands in contrast to the open availability of 3.6 Flash and 3.5 Flash-Lite, highlighting Google&rsquo;s tiered approach to AI deployment based on risk assessment.</p>
<h2 id="which-model-should-you-choose">Which Model Should You Choose?</h2>
<h3 id="pricing-comparison">Pricing Comparison</h3>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th>Input Price (per 1M tokens)</th>
          <th>Output Price (per 1M tokens)</th>
          <th>Speed (tokens/s)</th>
          <th>Intelligence Index Score</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Gemini 3.6 Flash</td>
          <td>$1.50</td>
          <td>$7.50</td>
          <td>303.6 (#1/187)</td>
          <td>50 (#21/187)</td>
      </tr>
      <tr>
          <td>Gemini 3.5 Flash-Lite</td>
          <td>$0.30</td>
          <td>$2.50</td>
          <td>350.2 (#3/152)</td>
          <td>36 (#12/152)</td>
      </tr>
      <tr>
          <td>Gemini 3.5 Flash Cyber</td>
          <td>Not publicly priced</td>
          <td>Not publicly priced</td>
          <td>Not disclosed</td>
          <td>Frontier on CyberGym</td>
      </tr>
  </tbody>
</table>
<h3 id="use-case-recommendations">Use Case Recommendations</h3>
<p><strong>Choose Gemini 3.6 Flash when:</strong></p>
<ul>
<li>You need the best general-purpose intelligence in the Flash family</li>
<li>You are building production agentic workflows that demand high accuracy</li>
<li>Computer use and tool-calling are core to your application</li>
<li>You want the fastest inference speed available (303.6 tokens/s)</li>
<li>Your budget allows $7.50 per 1M output tokens</li>
</ul>
<p><strong>Choose Gemini 3.5 Flash-Lite when:</strong></p>
<ul>
<li>You are processing large volumes of documents or search queries</li>
<li>Cost efficiency is your primary concern ($0.30 per 1M input tokens)</li>
<li>You need high throughput for batch inference or real-time applications</li>
<li>You are migrating from 3.1 Flash-Lite and want a significant quality upgrade</li>
<li>Your application can tolerate slightly lower intelligence for dramatically lower cost</li>
</ul>
<p><strong>Choose Gemini 3.5 Flash Cyber when:</strong></p>
<ul>
<li>You work in cybersecurity operations or vulnerability research</li>
<li>You are a government agency or trusted partner with access to the pilot</li>
<li>You need AI specifically fine-tuned for security tasks</li>
<li>Defensive cyber operations are your primary use case</li>
</ul>
<h2 id="what-is-the-bigger-picture-for-googles-ai-strategy">What Is the Bigger Picture for Google&rsquo;s AI Strategy?</h2>
<p>Google also announced that Gemini 3.5 Pro is currently testing with partners, and Gemini 4 pre-training has started. This signals that Google is already looking beyond the current generation, with Gemini 4 expected to represent a significant architectural leap.</p>
<p>The three-tier Flash strategy — workhorse (3.6 Flash), cost-efficient scale (3.5 Flash-Lite), and specialized security (3.5 Flash Cyber) — shows Google&rsquo;s maturation as an AI platform provider. Rather than offering a single model, Google is segmenting the market by use case, price sensitivity, and safety requirements.</p>
<p>The inclusion of computer use as a built-in client-side tool across the Flash family is a strong signal that Google believes agentic AI — where models take actions in digital environments rather than just generating text — is the next major paradigm shift. At $0.30 per 1M input tokens for 3.5 Flash-Lite, agentic search and document processing become economically viable at scale, potentially unlocking entirely new categories of AI applications.</p>
<h2 id="frequently-asked-questions">Frequently Asked Questions</h2>
<h3 id="what-is-the-difference-between-gemini-36-flash-and-35-flash">What is the difference between Gemini 3.6 Flash and 3.5 Flash?</h3>
<p>Gemini 3.6 Flash is the upgraded version of 3.5 Flash, offering 17% fewer output tokens, improved benchmark scores across the board (DeepSWE 49% vs 37%, MLE Bench 63.9% vs 49.7%), lower pricing at $7.50 per 1M output tokens versus $9, and enhanced Frontier Safety safeguards with a knowledge cutoff advanced to March 2026.</p>
<h3 id="how-much-does-gemini-35-flash-lite-cost">How much does Gemini 3.5 Flash-Lite cost?</h3>
<p>Gemini 3.5 Flash-Lite costs $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, making it the most cost-effective model in Google&rsquo;s Flash family. It delivers 350 output tokens per second, ranking 3rd out of 152 models on the Artificial Analysis speed index.</p>
<h3 id="is-gemini-35-flash-cyber-available-to-the-public">Is Gemini 3.5 Flash Cyber available to the public?</h3>
<p>No, Gemini 3.5 Flash Cyber is currently available only to governments and trusted partners through a limited-access pilot program. It powers CodeMender, Google&rsquo;s multi-agent cybersecurity system, and is not intended for general-purpose use.</p>
<h3 id="which-gemini-flash-model-is-best-for-agentic-ai-workflows">Which Gemini Flash model is best for agentic AI workflows?</h3>
<p>Gemini 3.6 Flash is the best choice for agentic AI workflows due to its top-ranked speed (303.6 tokens/s, #1/187), strong benchmark performance (83.0% on OSWorld-Verified), and built-in computer use as a client-side tool. For cost-sensitive agentic workflows, 3.5 Flash-Lite at $0.30/1M input tokens makes large-scale agentic operations economically viable.</p>
<h3 id="what-is-the-context-window-size-for-all-three-models">What is the context window size for all three models?</h3>
<p>All three models — Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — share a 1 million token context window. They also all support text, image, speech, and video input with text output, and reasoning variants are available for 3.6 Flash and 3.5 Flash-Lite.</p>
]]></content:encoded></item></channel></rss>