AI Research NotesGemini 3.8 Flash · Model Card · Sep 2026
Google · Gemini 3.8 Flash Model Card · Published Sep 2026

더 빠른 모델이 아니라
생산형 Agent의 균형점을 겨냥한다

Gemini 3.8 Flash advances software engineering and agentic knowledge workflows while retaining configurable effort, multimodal 1M-token context, and production-oriented cost/latency controls.

TEXTIMAGEAUDIO / VIDEOUP TO 1M TOKENSGemini 3.8 Flashbased on Gemini 3.7 FlashCONFIGURABLE EFFORTqualitycostlatency64K text outputSOFTWARE ENGINEERINGlong-horizon codingAGENTIC KNOWLEDGE WORKfinance · legal · terminalMULTIMODAL / SCIENCEcharts · video · biologyPRODUCTION CHANNELSGemini app / APIAI StudioEnterprise AgentAI Mode / Antigravityone model card, three tensions: capability × operating cost × safety
Executive Reading

Gemini 3.8 Flash의 핵심 포지셔닝은 최고 단일 benchmark 점수를 모두 가져가는 모델이 아니라, software engineering과 agentic knowledge workflow를 강화하면서 quality·cost·latency의 조절 가능성을 유지하는 생산형 범용 Agent 모델이다.

모델카드는 Gemini 3.8 Flash를 Gemini 3.7 Flash의 다음 iteration으로 정의한다. 입력은 text, image, audio, video이며 context window는 최대 1M tokens, 출력은 text 최대 64K tokens이다. 모델은 configurable effort level을 유지해 사용자가 성능, 비용, 지연시간의 조합을 조절하도록 한다.

1Mmax input context
64Kmax text output
$0.75intro input / 1M tokens
$3.75intro output / 1M tokens
페이지 5의 benchmark table은 한 가지 사실을 선명하게 보여준다. Gemini 3.8 Flash는 모든 영역에서 1위가 아니지만, finance agent·legal workflow·terminal coding·chart reasoning·long video·multidisciplinary reasoning·biology workflow에서 폭넓게 경쟁력 있는 프로파일을 형성한다. Flash라는 이름의 의미가 단순 속도보다 “폭넓은 production utility”에 가까워진 셈이다.
Part I · Model Identity

3.7 Flash 위에 software engineering과 agentic knowledge work를 더 밀어 올린다

모델카드는 구조적 계보를 명확히 하면서도, 이번 release의 차별점을 capability positioning에 둔다.

§1 · Description and dependency

Gemini 3.8 Flash는 Gemini 3 model family의 다음 iteration이며 Gemini 3.7 Flash를 기반으로 한다. 모델카드가 직접 강조하는 성능 진전 영역은 software engineering과 agentic knowledge workflows이다. 동시에 configurable effort levels를 그대로 지원해 quality, cost, latency 사이의 운영점을 사용자가 조절할 수 있다.

여기서 중요한 점은 architecture 자체의 상세를 새로 공개하지 않는다는 사실이다. architecture 설명은 Gemini 3.7 Flash model card로 위임된다. 즉 이번 문서는 새 architecture disclosure보다 capability, deployment, evaluation, usage, safety에 초점을 맞춘 업데이트형 model card다.

§2 · Model data and implementation disclosure

Training Dataset

3.8 Flash 전용 dataset detail은 본 카드에서 새로 설명하지 않고 Gemini 3.7 Flash model card를 참조하도록 한다.

Data Processing

training-data processing 역시 3.7 Flash model card를 참조한다.

Hardware

hardware 및 sustainability 관련 설명은 3.7 Flash model card에 의존한다.

Software

software implementation 세부도 3.7 Flash model card를 참조한다.

Transparency boundary

Gemini 3.8 Flash model card만으로는 새로운 architecture, training corpus 구성, preprocessing pipeline, training hardware의 구체적인 delta를 재구성할 수 없다. 이 문서가 지원하는 해석은 3.7 Flash 기반이라는 계보와, 3.8에서 보고된 capability·evaluation·safety 변화까지다.

Part II · Inputs, Outputs & Distribution

1M-token multimodal input을 다양한 production surface에 배포한다

모델 자체뿐 아니라 어디에서 어떤 입력으로 사용할 수 있는지가 model card의 핵심 정보다.

§3 · Input and output contract

Inputs

Text strings, document(s), image, audio, video.

Context

최대 1M tokens의 context window.

Output

Text, 최대 64K tokens.

이 조합은 한 번의 요청 안에서 긴 문서, 시각정보, 음성, 영상을 함께 취급하는 general-purpose multimodal agent의 입력계약을 형성한다. 다만 출력 modality는 이 model card 기준 text로 한정된다.

§4 · Distribution channels

모델은 API를 통해 downstream provider에도 제공되며, 사용을 위해 별도의 필수 hardware/software가 요구되지는 않는다고 카드가 설명한다. 배포 채널은 다음과 같다.

Gemini appGemini Enterprise Agent PlatformGoogle AI StudioGemini APIGoogle AI ModeGoogle Antigravity

consumer app, developer API, enterprise agent platform, search/AI mode, agentic development surface까지 하나의 모델이 여러 운영층에 배치된다는 점은 3.8 Flash의 production-oriented positioning을 강화한다.

Part III · Evaluation Landscape

페이지 5의 표는 coding에서 biology까지 하나의 agent capability surface를 그린다

평가범위는 coding, knowledge work, multimodal, long-context, computer use, scientific reasoning이다. 아래 값은 2026년 9월 model card에 보고된 결과를 그대로 옮긴다.

§5 · Pricing and benchmark table
Metric / BenchmarkGemini 3.8 FlashGemini 3.7 FlashClaude Opus 5Claude Sonnet 5GPT-5.6 SolGPT-5.6 Terra
Input price · $/1M tokens, no caching$0.75
$1.50 regular
$0.75
$1.50 regular
$5.00$2.00$4.00$2.00
Output price · $/1M tokens$3.75
$7.50 regular
$3.75
$7.50 regular
$25.00$10.00$20.00$12.00
DeepSWE v1.1
Long-horizon software engineering
73.7%65.3%74.0%53.8%72.7%69.6%
GDPval-AA v2
Knowledge work · Elo
154514821824158417101528
Vals Finance Agent v2
Financial analyst tasks
61.4%59.0%58.6%53.9%53.8%54.4%
Harvey's Legal Agent Benchmark
Complex legal workflows · all-pass rate
10.0%8.8%6.7%5.0%2.5%0.8%
Terminal-bench 2.1
Agentic terminal coding
89.4%85.8%89.1%80.4%88.8%87.4%
Terminal-bench 4.0
General agent capabilities
19.1%11.2%51.8%12.4%37.3%23.6%
GDP.PDF
Expert PDF document comprehension · all-pass rate
35.0%34.0%37.0%28.0%40.0%29.0%
CharXiv Reasoning
Synthesis from complex charts · no tools
86.2%84.5%83.7%70.1%85.8%85.9%
LVBench
Long video understanding
87.8% agentic
87.1% static
85.4%75.4%68.5%82.1%78.9%
HLE-Verified
Multidisciplinary expert reasoning
54.9%53.6%54.4%31.0%54.5%51.1%
OSWorld-2.0
Agentic computer use · partial score / batch tool enabled
59.0%50.6%75.4%42.6%62.6%50.2%
BioMysteryBench · Human Solvable
Bioinformatics research workflows
88.8%87.1%90.1%87.5%79.5%83.8%
BioMysteryBench · Human Difficult56.5%43.5%49.4%34.1%44.7%49.4%
LABBench2
Biology real-world research tasks
86.2%82.1%84.2%80.1%82.1%81.2%

Model Card p.5. 각 benchmark는 metric과 조건이 다르므로 서로 다른 row의 숫자를 하나의 통합점수처럼 비교해서는 안 된다. 표의 굵은 값은 해당 row에서 수치상 가장 높은 값을 읽기 쉽게 표시한 편집적 강조다.

§6 · What the table actually says

Clear gains vs 3.7

DeepSWE 65.3→73.7, Finance Agent 59.0→61.4, Legal 8.8→10.0, Terminal-bench 2.1 85.8→89.4, LABBench2 82.1→86.2처럼 3.7 대비 폭넓은 상승이 보인다.

Not universal best

GDPval-AA, Terminal-bench 4.0, GDP.PDF, OSWorld-2.0, BioMysteryBench Human Solvable에서는 다른 비교모델이 더 높은 값을 기록한다.

Broad profile

legal·finance·terminal·chart·video·expert reasoning·biology를 모두 한 모델카드에서 평가한다는 점 자체가 agentic general-purpose positioning을 보여준다.

3.8 Flash의 강점은 “모든 benchmark 1위”가 아니라, 3.7 대비 software/agentic/scientific workflow에서 일관된 전진을 보이면서 낮은 가격대를 유지하는 조합에 있다.

Part IV · Agentic & Scientific Knowledge Work

Finance·legal·terminal·PDF·chart·video·biology가 하나의 agent workflow 관점으로 묶인다

이 model card의 benchmark 선택은 production agent가 실제로 만나게 되는 복합 업무형태를 넓게 반영한다.

§7 · Software engineering and terminal agents

DeepSWE v1.1에서 73.7%를 기록해 3.7 Flash의 65.3%에서 상승하고, Terminal-bench 2.1에서는 89.4%로 비교표에서 가장 높은 값을 기록한다. 반면 Terminal-bench 4.0에서는 19.1%로 상승했지만 Claude Opus 5의 51.8%, GPT-5.6 Sol의 37.3%보다 낮다.

이는 “coding capability”를 단일 축으로 부르기 어렵다는 점을 보여준다. long-horizon software engineering, terminal coding, general agent capability는 서로 다른 operating demands를 측정한다.

§8 · Knowledge work and document reasoning

Finance Agent v2는 61.4%, Harvey's Legal Agent Benchmark는 10.0%로 비교표 내 최고값이다. GDP.PDF는 35.0%로 GPT-5.6 Sol의 40.0%보다 낮지만 3.7 Flash의 34.0%보다 상승한다. CharXiv Reasoning은 no-tools 조건에서 86.2%로 가장 높다.

이 조합은 agentic knowledge work가 단순 QA가 아니라 문서 이해, 복잡한 표/차트 reasoning, 다단계 전문업무 flow를 포함한다는 model-card 관점을 드러낸다.

§9 · Multimodal and scientific reasoning

LVBench에서 Gemini 3.8 Flash는 agentic 설정 87.8%, static 설정 87.1%를 보고한다. HLE-Verified는 54.9%다. BioMysteryBench는 Human Solvable 88.8%, Human Difficult 56.5%이며, LABBench2는 86.2%다.

Scientific reasoning 관련 결과를 “AI Scientist 완성”으로 읽는 것은 과도하다. 모델카드가 보여주는 것은 biology research task와 expert reasoning benchmark에서의 성능이며, autonomous experiment loop나 실험실 실행능력을 직접 입증하는 것은 아니다.
Part V · Cost, Intended Use & Limitations

Production-ready라는 말은 capability뿐 아니라 가격·지연·실패모드까지 포함한다

모델카드는 3.8 Flash를 cost-effective scaling of general-purpose, production-ready agents에 적합한 모델로 설명한다.

§10 · Pricing and effort

페이지 5의 표에서 3.8 Flash의 introductory price는 caching 없는 input 기준 $0.75 / 1M tokens, output 기준 $3.75 / 1M tokens이다. 각 cell에는 regular price로 input $1.50, output $7.50가 병기되어 있다.

Introductory pricing footnote

3.7 Flash와 3.8 Flash의 introductory price는 2026년 12월 31일 종료되고, 2027년 1월 1일부터 input $1.50 / 1M tokens, output $7.50 / 1M tokens가 적용된다고 페이지 5 footnote가 명시한다.

동시에 higher effort level에서는 성능을 최대화하기 위해 token을 더 많이 사용할 수 있다고 limitations section이 밝힌다. 따라서 configurable effort는 단순 “quality knob”가 아니라 token consumption, latency, cost와 연동되는 production control surface로 이해하는 편이 정확하다.

§11 · Intended usage

권장 대상은 users, developers, enterprises이며 핵심 use case는 software engineering, agent tasks, complex knowledge workflows다. 모델카드는 이를 general-purpose production-ready agents를 cost-effective하게 scale하는 용도와 연결한다.

§12 · Known limitations

Hallucinations

foundation model의 일반적 limitation으로 hallucination 가능성을 명시한다.

Jailbreak resistance

지속적인 개선영역이며 Frontier Safety mitigation을 최근 강화했다고 설명한다.

Latency / timeout

간헐적인 slowness나 timeout 문제가 발생할 수 있다.

Token usage

higher effort에서 성능 극대화를 위해 더 많은 token을 사용할 수 있다.

knowledge cutoff는 March 2026로 제시된다. 다만 일부 domain은 더 최신 정보를 기대할 수 있는 반면, 다른 domain에서는 Gemini 3 family와 동일하게 January 2025 수준으로 제한될 수 있다고 카드가 경고한다. “cutoff = 모든 지식이 그 날짜까지 균일하게 최신”이라는 해석은 model card가 직접 부정하는 셈이다.

Acceptable Usage, broader known limitations, safety policies의 상세는 Gemini 3.7 Flash model card를 참조하도록 되어 있다.

Part VI · Safety, Red Teaming & Frontier Assessment

안전성은 “대체로 유사”하지만 multilingual safety와 unjustified refusal 지표에는 regression이 보고된다

페이지 7의 개발단계 safety table은 자동평가이며 human evaluation이나 red teaming과 구분해야 한다.

§13 · Automated development evaluations
EvaluationDescriptionGemini 3.8 Flash vs 3.7 FlashDirection
Text to Text SafetyAutomated content safety evaluation measuring safety policies-0.4ppLower is better
Multilingual SafetyAutomated safety policy evaluation across multiple languages+5.4ppLower is better
Image to Text SafetyAutomated content safety evaluation measuring safety policies0.0ppLower is better
ToneAutomated evaluation measuring objective tone of responses+0.2ppHigher is better
Unjustified-refusalsAbility to respond to borderline prompts while remaining safe+1.1ppLower is better

모델카드는 전체적으로 safety와 tone이 3.7 Flash와 유사하고 unjustified refusal은 낮다고 평가한다. 그러나 non-English language safety는 3.7 Flash 대비 소폭 regression했다고 직접 명시한다.

Comparability warning

이 결과는 개발단계의 automated evaluations이며 human evaluation이나 red teaming이 아니다. 또한 evaluation 자체가 개선되었기 때문에 이전 Gemini model card의 결과와 직접 비교할 수 없다고 카드가 경고한다. 자동평가 variation을 고려해 flagged content를 수동검토했으며, losses의 압도적 다수는 false positive이거나 egregious하지 않은 사례였다고 보고한다.

§14 · Human red teaming

수동 red teaming은 model development team 밖의 specialist team이 수행하며 high-level finding이 다시 model team에 전달된다. child safety 평가에서는 launch threshold를 충족했고, 일반적인 content safety policy에서도 3.7 Flash 대비 유사하거나 개선된 안전성능을 관찰했다고 카드가 설명한다.

red-team scope는 strict policy 밖의 잠재적 issue까지 포함했고 Gemini 3.1 Pro와 비교했으며, egregious concern을 발견하지 않았다고 보고한다.

§15 · Frontier Safety Assessment

Gemini 3.7 Flash는 April-2026 Frontier Safety Framework에 따라 평가되었고 Tracked 또는 Critical Capability Levels(T/CCLs)에 도달하지 않았다. 3.8 Flash에 대해서는 3.7 대비 해당 domain에서 meaningful new capability 또는 material performance increase가 없다는 평가를 바탕으로, 3.8 역시 T/CCL에 도달할 가능성이 낮다고 판단한다.

이 문구는 “3.8 Flash가 독립적인 새 frontier threshold를 넘지 않았다는 완전한 새 평가표”라기보다, 3.7의 평가결과와 3.8의 capability delta를 함께 사용한 판단으로 읽는 것이 정확하다.
Part VII · Synthesis

Gemini 3.8 Flash가 보여주는 방향은 “작은 모델”보다 “운영 가능한 Agent 모델”에 가깝다

이 결론은 model card의 capability, price, deployment, limitation, safety evidence를 함께 읽은 해석이다.

§16 · Three operating tensions

Capability

software engineering, finance/legal workflows, terminal coding, chart/video reasoning, biology research tasks에서 3.7 대비 진전.

Economics

introductory low price와 configurable effort로 token use, quality, latency 사이의 운영점을 조절.

Safety

overall safety는 유사하지만 multilingual safety와 unjustified-refusal 자동지표 regression을 숨기지 않고 공개.

따라서 Gemini 3.8 Flash의 가장 설득력 있는 해석은 “최고성능 frontier model의 축소판”이 아니라, agentic workload를 넓게 처리하면서 deployment economics와 safety envelope를 함께 관리하는 production-tier foundation model이라는 것이다.

§17 · What remains unknown

Evidence boundary

이 model card만으로는 3.8 Flash의 상세 architecture 변화, 새 training dataset composition, 새로운 training-data processing, hardware/software stack의 차이를 알 수 없다. benchmark는 각기 다른 methodology와 metric을 사용하므로 row 간 수치를 직접 합산하거나 “종합 1위”를 계산하는 것도 근거가 없다. scientific-reasoning benchmark의 높은 점수 역시 autonomous science execution이나 AGI를 입증하지 않는다.

§18 · Final takeaway

Gemini 3.8 Flash가 실용적으로 중요한 이유는 1M multimodal context, agentic workflow 성능, 64K output, configurable effort, 넓은 distribution surface가 하나의 운영계약으로 묶여 있기 때문이다.

Model Card 전체를 한 문장으로 압축하면 이렇다. Gemini 3.8 Flash는 더 많은 일을 할 수 있게 된 3.7 Flash이면서, 그 능력을 실제 agent product에 배치할 때 필요한 가격·지연·안전의 경계조건을 함께 명시한 모델이다.
Primary Source & Methodology

References

01
Gemini 3.8 Flash — Model Card
Google · Published September 2026 · 8 pages

본 게시물의 모든 model specification, benchmark, pricing, intended use, limitation, safety, red-team, frontier-safety 서술의 1차 근거다.

02
Gemini 3.8 Flash Evaluation Methodology
Google DeepMind · methodology reference named in the model card
deepmind.google/models/evals-methodology/gemini-3-8-flash

Source boundary

본 게시물은 첨부된 8쪽 model card 전체와 특히 페이지 5 benchmark/pricing 표, 페이지 7 safety-evaluation 표를 직접 검토해 재구성했다. 외부 웹검색으로 추가 benchmark나 spec을 보강하지 않았으며, 카드가 Gemini 3.7 Flash model card로 위임한 architecture·training data·processing·hardware·software·acceptable use·policy detail은 추측하지 않았다.