The best overall starting point: Claude Fable 5.1
If you want one general shortlist candidate, Claude Fable 5.1 is the current overall leader on SuperIndex. It scores 80.2 under our composite of published evaluations, with 100% confidence. That makes it an evidence-based place to start, rather than a promise that it will be the best tool for every person. Your goal, budget and preferred interface still matter.
An AI model is the system producing answers; the app around it supplies tools, memory, file handling and the interface. SuperIndex compares model evidence, not the quality of every app that offers that model. A subscription app’s pricing can also differ from the API token rates shown here. Before paying, check that the product exposes the exact version you intended to try.
For coding: Fugu Ultra
Fugu Ultra leads the coding pillar among models eligible for an overall rank. It is a sensible candidate for implementation and debugging. Coding evaluations cover particular repositories, tools and harnesses; they do not establish universal success in your stack. Try a small representative change, run your existing tests and inspect the diff. The coding board and model sources show which evaluations support the recommendation.
For writing: Claude Opus 5.5
Claude Opus 5.5 has the highest human-preference pillar in the ranked set. This is a conversational preference proxy, not a dedicated writing benchmark. Use it to choose a trial candidate, then test your own tone, editing constraints and factual grounding. A model favored in broad comparisons may still need several revisions to fit a particular audience. We do not label that proxy as proof of the best prose.
For math: GPT-6.1 Sol
GPT-6.1 Sol leads the reported math pillar. Consider it when a problem needs multi-step mathematical reasoning, then check the final result independently. Competition math, code-based calculations and classroom explanations are distinct tasks. A leaderboard can help choose a candidate; it does not replace checking assumptions, units and arithmetic. Examine the math evaluations for the actual problem families measured.
For a tighter budget: Gemma 4 26B A4B IT
Gemma 4 26B A4B IT leads our value heuristic among ranked models with SI Score at least 50 and known input and output rates: $0.00 input and $0.00 output per million tokens. The calculation uses three input tokens per output token and a small price floor. It is a screening rule, not a prediction of your bill. Long generated answers or repeated tool calls change your costs. Missing price means unknown, not free.
For downloadable weights: Kimi K3
Kimi K3 is the highest overall rank with confirmed open weights. Downloadable weights can give you more control over hosting, but hardware, operations and licensing become part of the decision. Review the license and context facts on the profile. An API can be simpler for a small workload even when open weights are available.
Make the final choice with a short trial
Use the same realistic prompts on two or three candidates, check the output against known answers and record the time and cost in the app you will actually use. Start with the task winner and the overall leader, rather than testing the entire catalog. Confidence on SuperIndex is source completeness; it does not mean the answers are 100% accurate. Read the FAQ for the terminology, or compare shared results before deciding.