Google, released Sep 2, 2026

Gemini 3.8 Flashprice, context, benchmarks and release details

100% confidence 100 percent, Full confidence, 7 of 7 expected sources in
68.8
SI Score
#17 of 142 ranked models
Input, per 1M tokens
$0.75Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 9, 2026 · CC-BY-4.0 factual citation
Open source ↗
Output, per 1M tokens
$3.75Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 9, 2026 · CC-BY-4.0 factual citation
Open source ↗
Context window
1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
65.5Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Sep 2, 2026Google Gemini release notesPublished source fact Retrieved Oct 9, 2026 · CC-BY-4.0 factual citation
Open source ↗

How the score breaks down

Coding (weight 40 percent) 56.4
Math (weight 15 percent) 77.5
Preference (weight 15 percent) 81.7
Reasoning (weight 30 percent) 74.5

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 05:13 UTC.

Around it on the leaderboard

  1. 15 Claude Sonnet 5.5 69.4
  2. 16 GPT-6 Sol 68.9
  3. 17 Gemini 3.8 Flash 68.8
  4. 18 Qwen3.8 Max 68.7
  5. 19 Gemini 3.5 Flash 68.3

Full leaderboard

Benchmark results

30 benchmarks, 41 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

ARC-AGI-1 (public eval)reasoning 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 3 settings
  • low effort 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • medium effort 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (public eval)
ARC-AGI-1 (semi-private)reasoning 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 3 settings
  • low effort 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • medium effort 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (semi-private)
ARC-AGI-2 (public eval)reasoning 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 3 settings
  • low effort 77.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • medium effort 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (public eval)
ARC-AGI-2 (semi-private)reasoning 89.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.2 3 settings
  • low effort 77.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • medium effort 82.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 89.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (semi-private)
ARC-AGI-3 (semi-private)reasoning 10.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
10.4 3 settings
  • low effort 6.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • medium effort 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 10.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-3 (semi-private)
GPQA Diamondreasoning 95.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
95.4
Humanity's Last Exam (1,811 verified items)reasoning 54.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] 1811 verified and revised items Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
54.9
Humanity's Last Exam (Scale AI)reasoning 44.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Sep 9, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
44.5
LiveBench Reasoning: connectionsreasoning 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
100.0
LiveBench Reasoning: consecutive eventsreasoning 20.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
20.6
LiveBench Reasoning: logic with navigationreasoning 82%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
82.0
LiveBench Reasoning: spatialreasoning 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
100.0
LiveBench Reasoning: theory of mindreasoning 76.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
76.9
LiveBench Reasoning: zebra puzzlesreasoning 98.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
98.3
FrontierMath Tier 4 (v2)math 22.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
22.0
FrontierMath Tiers 1–3 (v2)math 68.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
68.4
LiveBench Math: AMPS Hardmath 99%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
99.0
LiveBench Math: competition mathmath 95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
95.1
LiveBench Math: integralsmath 80%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
80.0
LiveBench Math: olympiadmath 92.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
92.2
LiveBench Math: simplifymath 75.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
75.4
OTIS Mock AIME 2024–2025math 98.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
98.9
LiveBench Coding: code completioncoding 71.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
71.7
LiveBench Coding: code generationcoding 73.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
73.2
LiveBench Coding: JavaScriptcoding 72.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
72.7
LiveBench Coding: Pythoncoding 50%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
50.0
LiveBench Coding: TypeScriptcoding 40%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
40.0
Terminal-Bench 2.1coding 89.4%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
89.4
Terminal-Bench 4.0coding 19.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
19.1 2 settings
  • Setting 1 19.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
  • Setting 2 19.1%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] mini-SWE-agent; highPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗
About Terminal-Bench 4.0
LMArena Textpreference 1499.0 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
81.7

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Nomodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
not yet reported
Input modalities
text, image, video, audio, pdfmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
100% of expected source weight

Reported (7)

  • ARC PrizeOct 8, 2026
  • Epoch AI BenchmarkingOct 8, 2026
  • Humanity’s Last ExamOct 8, 2026
  • Official model cards via models.devOct 8, 2026
  • LiveBenchOct 8, 2026
  • LMArena / ArenaOct 8, 2026
  • Terminal-BenchOct 8, 2026

Awaiting (0)

Every expected source has reported for this model.

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed