GPT-6 AstravsGPT-6.1 Sol
SI Score, benchmarks, price and context compared, with every number sourced.
77.5
SI Score
#3 of 142 ranked models
73.0
SI Score
#6 of 142 ranked models
Pillar by pillar
GPT-6 AstraGPT-6.1 Sol
87.1 Reasoning30% of score 87.5
92.9 Math15% of score 93.0
64.8 Coding40% of score 65.0
76.9 Preference15% of score 77.4
The basics
| Attribute | GPT-6 Astra | GPT-6.1 Sol |
|---|---|---|
| Input price, per 1M tokens | $10.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $2.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Output price, per 1M tokens | $50.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $10.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Context window | 1.1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ | 1.1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ |
| Released | Sep 3, 2026OpenAI API changelogPublished source fact
Retrieved Oct 9, 2026 · factual citation Open source ↗ | Sep 29, 2026OpenAI API changelogPublished source fact
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Weights | Closed | Closed |
Highlighted values are the lower price or the larger context window.
Shared benchmarks
27 in common. Best result ahead: GPT-6 Astra on 11, GPT-6.1 Sol on 6Each side shows its best published result. Results can come from different settings or harnesses, so open a row to compare like with like before reading much into a small gap.
ARC-AGI-1 (public eval) 99% 98.5%
GPT-6 Astra
- low effort98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
GPT-6.1 Sol
- low effort96.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
ARC-AGI-1 (semi-private) 98.5% 98.5%
GPT-6 Astra
- low effort96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
GPT-6.1 Sol
- low effort93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
ARC-AGI-2 (public eval) 97.9% 97.1%
GPT-6 Astra
- low effort94.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort96.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
GPT-6.1 Sol
- low effort77.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort90.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort97.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
ARC-AGI-2 (semi-private) 95% 94.2%
GPT-6 Astra
- low effort85.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
GPT-6.1 Sol
- low effort76.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort94.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
ARC-AGI-3 (semi-private) 62.7% 52.7%
GPT-6 Astra
- low effort17.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort38.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort54.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort59.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort62.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
GPT-6.1 Sol
- low effort3.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - medium effort10.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - high effort26.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - xhigh effort39.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - max effort52.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
FrontierMath Tier 4 (v2) 97.6% 100%
GPT-6 Astra
- no reasoning82.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - low effort87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - medium effort97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - high effort97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - xhigh effort97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - max effort97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
GPT-6.1 Sol
- max effort100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
FrontierMath Tiers 1–3 (v2) 93.7% 93.7%
GPT-6 Astra
- max effort93.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
GPT-6.1 Sol
- max effort93.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
GPQA Diamond 96% 95.4%
GPT-6 Astra
- max effort95.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - Setting 296%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
GPT-6.1 Sol
- max effort95.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
LiveBench Coding: code completion 80.4% 82.6%
GPT-6 Astra
- Published result80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Coding: code generation 80.3% 78.9%
GPT-6 Astra
- Published result80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result78.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Coding: JavaScript 63.6% 63.6%
GPT-6 Astra
- Published result63.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result63.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Coding: Python 65% 60%
GPT-6 Astra
- Published result65%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result60%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Coding: TypeScript 43.3% 46.7%
GPT-6 Astra
- Published result43.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result46.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: AMPS Hard 98% 98%
GPT-6 Astra
- Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: competition math 97.1% 97.1%
GPT-6 Astra
- Published result97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: integrals 100% 99%
GPT-6 Astra
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result99%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: olympiad 92.2% 91.9%
GPT-6 Astra
- Published result92.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result91.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: simplify 70.5% 68.2%
GPT-6 Astra
- Published result70.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result68.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: connections 100% 100%
GPT-6 Astra
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: consecutive events 90.6% 91.3%
GPT-6 Astra
- Published result90.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result91.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: logic with navigation 88% 82%
GPT-6 Astra
- Published result88%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result82%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: spatial 98% 98%
GPT-6 Astra
- Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: theory of mind 84.6% 86.5%
GPT-6 Astra
- Published result84.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result86.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: zebra puzzles 100% 100%
GPT-6 Astra
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
GPT-6.1 Sol
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LMArena Text 1440.5 elo 1446.5 elo
GPT-6 Astra
- Published result1440.5 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
GPT-6.1 Sol
- Published result1446.5 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
OTIS Mock AIME 2024–2025 100% 100%
GPT-6 Astra
- max effort100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
GPT-6.1 Sol
- max effort100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
Terminal-Bench 4.0 58.2% 58.2%
GPT-6 Astra
- Setting 158.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗ - Setting 257.9%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗ - Setting 357.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; highPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗ - Setting 457.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; xhighPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗ - Setting 554.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; mediumPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗ - Setting 650.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; lowPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
GPT-6.1 Sol
- Published result58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
Which should you choose?
- GPT-6.1 Sol costs less per input token ($2.00 vs $10.00 per 1M).
- Confidence is 100% for GPT-6 Astra and 100% for GPT-6.1 Sol; sources still to report can move either score.
These follow from the numbers above. They're not a verdict on your use case.