NVIDIA, released Aug 11, 2026

Nemotron 3.5 Lightning 30B A3Bprice, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 100% confidence 100 percent, Full confidence, 1 of 1 expected sources in
46.1
SI Score
#116 of 142 ranked models
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
262Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
262Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Aug 11, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How the score breaks down

Coding (weight 40 percent) 38.1
Math (weight 15 percent) —
Preference (weight 15 percent) —
Reasoning (weight 30 percent) 45.2

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 05:13 UTC.

Around it on the leaderboard

  1. 114 GLM-4.6 47.8
  2. 115 Claude Sonnet 4 46.6
  3. 116 Nemotron 3.5 Lightning 30B A3B 46.1
  4. 117 GPT-5 Pro 45.8
  5. 118 Llama-3.1-70B-Instruct 45.7

Full leaderboard

Benchmark results

5 benchmarks, 5 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

GPQA Diamondreasoning 75.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; no tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
75.4
Humanity's Last Exam (text only)reasoning 11.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; no tools; text-only Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
11.7
MMLU-Proreasoning 81.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; NeMo Gym / NeMo Evaluator SDK Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
81.9
SWE-bench Verifiedcoding 51.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; NeMo Evaluator Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
51.6
Terminal-Bench 2.1coding 24.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; NeMo Evaluator; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
24.6

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Yesmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
not yet reported
Input modalities
textmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
100% of expected source weight

Reported (1)

  • Official model cards via models.devOct 8, 2026

Awaiting (0)

Every expected source has reported for this model.

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed