출력 토큰 17% 줄고 에이전트 점수는 오른다 — Google Gemini Flash 3종 동시 공개More for Less: Google Drops Three Gemini Flash Models at Once

Google이 2026년 7월 21일 Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber를 동시에 공개했어요. 비용 효율·최고 속도·사이버보안이라는 세 방향으로 Flash 라인업을 확장한 발표예요.Google unveiled Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber in a single announcement. The trio extends the Flash lineup across API efficiency, maximum throughput, and cybersecurity specialization.

Source

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

한 발표, 세 방향 — Google Flash 라인업, 목적별로 쪼개졌어요

“모델 하나로 다 커버”에서 “목적별 모델”로 넘어가는 흐름이 빨라지고 있어요. 2026년 7월 21일, Google이 Flash 계열 신모델 3종을 한꺼번에 공개했어요. Gemini 3.6 Flash(정밀도·효율 강화), Gemini 3.5 Flash-Lite(속도·저가 특화), Gemini 3.5 Flash Cyber(사이버 취약점 전용 fine-tuning)예요.

세 모델이 같은 날 나온 건 우연이 아니에요. OpenAI가 추론 특화 라인을 강화하고, Anthropic이 Haiku-Sonnet-Opus로 역할을 분리하는 동안, Google은 Flash 계열에서 같은 전략을 택한 거예요.

핵심 요약

  • Gemini 3.6 Flash: 출력 토큰 17% 감소, 에이전트 벤치마크 상승(DeepSWE 49%, MLE Bench 63.9%, OSWorld 83.0%). 가격 $1.50/$7.50/1M 토큰
  • Gemini 3.5 Flash-Lite: 초당 350 토큰으로 3.5 시리즈 최고 속도. 가격 $0.30/$2.50/1M 토큰
  • Gemini 3.5 Flash Cyber: 취약점 탐지·검증·패치 전용. 정부·신뢰 파트너 제한 파일럿(CodeMender 통해). 이미 Chrome, Android, Cloud 내부 배포 완료
  • Gemini 3.5 Pro는 곧 공개 예정이고, Gemini 4는 사전 학습 중이에요

One Announcement, Three Directions

The shift from “one model does everything” to “purpose-built models” is accelerating. On July 21, 2026, Google published three new Flash models simultaneously: Gemini 3.6 Flash (precision and efficiency), Gemini 3.5 Flash-Lite (speed and low cost), and Gemini 3.5 Flash Cyber (fine-tuned for cybersecurity).

This isn’t a coincidence. As OpenAI deepens its reasoning lineup and Anthropic separates Haiku, Sonnet, and Opus by role, Google is making the same structural bet inside the Flash family.

TL;DR

  • Gemini 3.6 Flash: 17% fewer output tokens, higher agent benchmark scores (DeepSWE 49%, MLE Bench 63.9%, OSWorld 83.0%). $1.50/$7.50/1M tokens
  • Gemini 3.5 Flash-Lite: 350 tokens/sec — fastest in the 3.5 series. $0.30/$2.50/1M tokens
  • Gemini 3.5 Flash Cyber: Vulnerability detection, validation, and patching. Limited pilot for government and trusted partners via CodeMender. Already deployed across Chrome, Android, Cloud, Ads, and YouTube
  • Gemini 3.5 Pro coming soon, Gemini 4 in pre-training

모델별로 뜯어보면

Gemini 3.6 Flash — 에이전트용 업그레이드

3.5 Flash 대비 가장 눈에 띄는 변화는 “같은 일을 더 적은 토큰으로” 처리하는 방향이에요. 외부 평가 기관 Artificial Analysis Index 기준으로 출력 토큰이 17% 줄었고, 일부 벤치마크에서 최대 65% 효율 향상도 확인됐어요(Datacurve 기준). 코드 에디터에서 불필요한 편집이 줄고 실행 루프가 짧아진다는 뜻이에요.

에이전트 성능 수치도 올랐어요.

벤치마크3.6 Flash3.5 Flash
DeepSWE49%37%
MLE Bench63.9%49.7%
OSWorld-Verified83.0%78.4%
GDPval-AA v214211349

가격은 $1.50/$7.50/1M 토큰이에요. Harvey(법률 AI)가 테스트한 결과 문서 검토 작업이 전작 대비 12% 빨라졌다고 밝혔어요. 이미 Figma, Harvey, Hebbia, JetBrains와 협력 배포 중이에요.

Gemini 3.5 Flash-Lite — 속도를 원하면 여기

3.5 시리즈에서 속도만 놓고 보면 Flash-Lite가 가장 빨라요. 초당 350 출력 토큰이에요. 가격은 $0.30/$2.50/1M 토큰으로 경쟁 모델 중 최저 수준이에요.

벤치마크3.5 Flash-Lite이전 세대
Terminal-Bench 2.154%3.1 Flash-Lite: 31%
SWE-Bench Pro54.2%3 Flash: 49.6%
OSWorld-Verified74.0%3 Flash: 65.1%
GDPval-AA v21140이전: 642

코드 에이전트나 데이터 처리처럼 속도가 병목이 되는 작업에 잘 맞아요.

Gemini 3.5 Flash Cyber — 취약점 탐지 전용

이번 발표에서 가장 독특한 포지셔닝이에요. 3.5 Flash를 기반으로 취약점 탐지·검증·패치에 맞춰 fine-tuning한 모델인데, 일반 API로는 판매하지 않아요. 정부와 신뢰된 파트너만 CodeMender를 통해 파일럿 접근이 가능해요.

이미 내부 배포는 됐어요. Chrome, Android, Cloud, Ads, YouTube에서 취약점 탐지에 실제로 쓰이고 있거든요. V8 JS 엔진 대상 테스트에서 고유 이슈 55개를 발견했어요. 같은 테스트에서 3.5 Flash는 47개, Claude Opus 4.6은 36개였어요(Google DeepMind 블로그 기준).

오용 리스크를 줄이기 위해 “신뢰된 방어자에게만” 접근을 제한한다는 게 Google의 설명이에요.

왜 중요한가

이 발표가 흥미로운 건 성능 수치보다 전략 방향 때문이에요. Google이 Flash 하나로 다 잡으려 하지 않고, 목적별로 쪼갰거든요.

사이버보안 모델을 공개 배포 없이 신뢰 파트너에게만 푸는 것도 중요해요. AI가 사이버 공격에 쓰일 수 있는 수준이 높아질수록 “누구에게 어떤 모델을 어떻게 주느냐”가 AI 거버넌스의 실질적 쟁점이 되고 있어요. Flash Cyber 발표는 Google이 그 문제를 회사 정책으로 명시한 사례예요.

동시에 Gemini 3.5 Pro 출시가 임박했고, Gemini 4가 사전 학습 중이에요. 내가 보기엔 이번 Flash 3종은 상위 모델이 나오기 전 API 생태계를 촘촘히 채우는 포석이에요.

핵심 통찰

Flash 라인업이 분화된다는 건, Google이 “다목적 모델 하나” 전략보다 “상황별 최적 모델” 전략으로 무게중심을 옮기고 있다는 얘기예요.

API 가격 경쟁이 벌어지는 시장에서 모델을 목적별로 쪼개는 건 고객이 “필요한 성능에 필요한 비용만 쓰게” 만드는 전략이에요. Flash-Lite가 $0.30/1M으로 내려가는 동안 Cyber는 가격조차 공개하지 않아요. 희소성과 가격 차별화를 동시에 챙기는 구조예요.

My Take

Flash Cyber가 가장 흥미로워요.

AI가 사이버 공격에 쓰일 수 있다는 우려는 오래됐는데, Google은 그 모델을 방어 목적으로 먼저 쓰고, 배포는 통제하는 방식을 선택했어요. 신뢰 파트너로만 제한하면 사실상 국가 단위 접근이 되는 건데, AI 무기화 논의가 실제 제품 정책으로 올라오는 순간이에요.

실무적으로는 API 비용을 줄이려는 팀이라면 3.6 Flash를 한번 평가해볼 만해요. 출력 토큰이 줄어들면 장기 에이전트 작업에서 비용 차이가 꽤 나거든요. Flash-Lite는 속도 우선 파이프라인에 넣어두면 좋아요.

Breaking Down the Three Models

Gemini 3.6 Flash — Built for Agents

The headline improvement is doing the same work with fewer tokens. Artificial Analysis Index measured a 17% reduction in output tokens versus 3.5 Flash; some tasks saw up to 65% efficiency gains per Datacurve. In practice: fewer unnecessary code edits and shorter execution loops.

Agent benchmarks moved up across the board:

Benchmark3.6 Flash3.5 Flash
DeepSWE49%37%
MLE Bench63.9%49.7%
OSWorld-Verified83.0%78.4%
GDPval-AA v214211349

Priced at $1.50/$7.50/1M tokens. Harvey reported 12% faster document review versus the prior model. Already shipping with Figma, Harvey, Hebbia, and JetBrains.

Gemini 3.5 Flash-Lite — When Speed Is the Bottleneck

The fastest in the 3.5 series at 350 output tokens per second. $0.30/$2.50/1M tokens — low end of the competitive range.

Benchmark3.5 Flash-LitePrevious
Terminal-Bench 2.154%3.1 Flash-Lite: 31%
SWE-Bench Pro54.2%3 Flash: 49.6%
OSWorld-Verified74.0%3 Flash: 65.1%
GDPval-AA v21140Prior: 642

Good fit for throughput-constrained pipelines where latency is the real constraint.

Gemini 3.5 Flash Cyber — Vulnerability Work Only

The most distinctive model in this drop. Fine-tuned from 3.5 Flash for vulnerability detection, validation, and patching — not available on the standard API. Access is limited to government and trusted partners through CodeMender.

Already in production: deployed across Chrome, Android, Cloud, Ads, and YouTube for internal vulnerability work. Found 55 unique issues in V8 JS Engine testing vs. 47 for 3.5 Flash and 36 for Claude Opus 4.6 (Google DeepMind blog). Pricing not disclosed.

Why It Matters

The numbers are good, but the strategy is the real story. Google is no longer trying to serve every use case with one Flash model — it’s segmenting by purpose.

Releasing a cybersecurity model with restricted access is also a deliberate governance statement. As AI capability for offensive work increases, “who gets what model and under what conditions” is becoming a real policy question. Flash Cyber is Google’s first explicit product-level answer to that question.

Meanwhile, Gemini 3.5 Pro is coming soon and Gemini 4 is in pre-training. The three Flash models read as Google filling out the API layer before the upper tier arrives.

Key Insight

Flash family segmentation signals a shift in Google’s API strategy — from “one capable model” to “right model for the job.”

In a price-competitive API market, purpose-specific models let customers pay only for the performance they need. Flash-Lite drops to $0.30/1M. Flash Cyber doesn’t list a price at all. Google is running value differentiation and scarcity simultaneously.

My Take

Flash Cyber is the one worth watching.

The debate over AI in cybersecurity has run for years. Google’s answer is to use it for defense first, and control distribution tightly. Limiting access to trusted partners makes it effectively a government-adjacent product — which means AI capability governance is now a product design decision, not just a policy paper.

For API teams: if you run long-horizon agent workflows, 3.6 Flash’s token reduction is worth a real evaluation. Compounding across many steps, the cost difference adds up. Flash-Lite is a straightforward swap for throughput-first pipelines.

댓글Comments