Google, released May 19, 2026

Gemini 3.5 Flashprice, context, benchmarks and release details

100% confidence 100 percent, Full confidence, 5 of 7 expected sources in
68.3
SI Score
#19 of 142 ranked models
Input, per 1M tokens
$1.50Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 9, 2026 · CC-BY-4.0 factual citation
Open source ↗
Output, per 1M tokens
$9.00Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 9, 2026 · CC-BY-4.0 factual citation
Open source ↗
Context window
1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
65.5Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
May 19, 2026Google Gemini release notesPublished source fact Retrieved Oct 9, 2026 · CC-BY-4.0 factual citation
Open source ↗

How the score breaks down

Coding (weight 40 percent) 64.0
Math (weight 15 percent) 74.1
Preference (weight 15 percent) 80.0
Reasoning (weight 30 percent) 79.7

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 05:13 UTC.

Around it on the leaderboard

  1. 17 Gemini 3.8 Flash 68.8
  2. 18 Qwen3.8 Max 68.7
  3. 19 Gemini 3.5 Flash 68.3
  4. 20 Claude Sonnet 4.6 68.3
  5. 21 GPT-5.6 Terra 68.2

Full leaderboard

Benchmark results

30 benchmarks, 34 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

ARC-AGI-1 (public eval)reasoning 96%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.0
ARC-AGI-1 (semi-private)reasoning 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5
ARC-AGI-2reasoning 72.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
72.1
ARC-AGI-2 (public eval)reasoning 72.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.1
ARC-AGI-2 (semi-private)reasoning 72.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.1
GPQA Diamondreasoning 92.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished May 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
92.8 3 settings
  • minimal effort 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 15, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • low effort 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • high effort 92.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished May 22, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About GPQA Diamond
Humanity's Last Exam (full set, text + multimodal)reasoning 40.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] full set, text + MMPublished May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
40.2
LiveBench Reasoning: connectionsreasoning 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
100.0
LiveBench Reasoning: consecutive eventsreasoning 48.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
48.0
LiveBench Reasoning: logic with navigationreasoning 74%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
74.0
LiveBench Reasoning: spatialreasoning 96%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
96.0
LiveBench Reasoning: theory of mindreasoning 80.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
80.8
LiveBench Reasoning: zebra puzzlesreasoning 77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
77.3
FrontierMath Tier 4 (v2)math 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
26.8
FrontierMath Tiers 1–3 (v2)math 62.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
62.8
LiveBench Math: AMPS Hardmath 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
98.0
LiveBench Math: competition mathmath 95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
95.1
LiveBench Math: integralsmath 68%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
68.0
LiveBench Math: olympiadmath 91.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
91.9
LiveBench Math: simplifymath 69.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
69.1
OTIS Mock AIME 2024–2025math 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished May 25, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
95.6 3 settings
  • minimal effort 80%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 15, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • low effort 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • high effort 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished May 25, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025
LiveBench Coding: code completioncoding 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
76.1
LiveBench Coding: code generationcoding 80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
80.3
LiveBench Coding: JavaScriptcoding 63.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
63.6
LiveBench Coding: Pythoncoding 50%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
50.0
LiveBench Coding: TypeScriptcoding 33.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
33.3
SWE-bench Pro (public)coding 55.1%Official model cards via models.devLab-reported; metric resolve rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] single attempt; publicPublished May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
55.1
SWE-bench Verifiedcoding 79.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 1, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
79.3
Terminal-Bench 2.1coding 76.2%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
76.2
LMArena Textpreference 1477.5 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
80.0

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Nomodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
not yet reported
Input modalities
text, image, video, audio, pdfmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
81% of expected source weight

Reported (5)

  • ARC PrizeOct 8, 2026
  • Epoch AI BenchmarkingOct 8, 2026
  • Official model cards via models.devOct 8, 2026
  • LiveBenchOct 8, 2026
  • LMArena / ArenaOct 8, 2026

Awaiting (2)

  • Humanity’s Last Exam13% of weight
  • Terminal-Bench6% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed