Ask 'what's the best AI model right now' and you'll get a different answer depending on who's selling. So let's use a referee instead. Artificial Analysis runs an independent Intelligence Index that scores the major models on the same battery of tests, and as of its August 2026 update the ranking is clear enough to settle a few arguments — including one surprise: Google's Gemini, long assumed to be in the top three, has slipped well behind. This is the honest leaderboard, what each top model actually costs, and — more useful than any single number — which one you should reach for depending on what you're doing.
One thing to hold onto as you read: 'best' and 'best for you' are different questions. The highest-scoring model isn't automatically the right pick, because price, speed and the specific job matter enormously. We'll rank them straight, then translate the ranking into real choices.
Quick summary: On Artificial Analysis's independent Intelligence Index (August 2026), the order at the top is Claude Opus 5 (63), Claude Fable 5 (62), GPT-5.6 Sol (61), Moonshot's Kimi K3 (60), Alibaba's Qwen 3.8-Max (58) and xAI's Grok 4.5 (56). Google's Gemini 3.1 Pro sits far lower at 48 - it hasn't shipped a new flagship, only a faster Flash model. Prices vary hugely: Claude Opus 5 is $5/$25 per million tokens in-out, GPT-5.6 Sol is $5/$30, while Qwen 3.8-Max and Grok 4.5 undercut them at $2/$6. The takeaway: Anthropic and OpenAI hold the top, Chinese models offer near-frontier quality far cheaper, and Gemini has fallen behind on raw intelligence.
The Honest Leaderboard
Here's the top of the independent Intelligence Index as it stands in August 2026. Claude Opus 5, released July 24, leads at 63. Anthropic's even pricier Claude Fable 5 is right behind at 62. OpenAI's GPT-5.6 Sol, the flagship of its Sol/Terra/Luna family, sits at 61 — close enough that on any given task the top three are effectively a three-way race. Just behind them, Moonshot's open-weight Kimi K3 lands at 60 — an open model matching the closed flagships — then Alibaba's Qwen 3.8-Max at 58 and xAI's Grok 4.5 at 56. Further down come the open-weight workhorses GLM-5.2 (53) and DeepSeek V4-Flash (52). And Gemini? Google's current flagship, Gemini 3.1 Pro, scores 48 on the latest index — outside the top tier entirely.
A note on reading these: the scores are from Artificial Analysis's August 2026 index, and benchmark numbers shift over time as testing methods are refined and models are updated. Treat them as a current snapshot of how these models compare with each other, not a permanent, fixed ranking.
What Happened to Gemini?
The Gemini slip deserves a word, because it surprises people. Google hasn't shipped a new top-tier Gemini in a while — as of August 2026 the flagship remains Gemini 3.1 Pro from earlier in the year, and the newest release was Gemini 3.6 Flash, a fast, cheap model built for speed rather than a frontier-pushing brain. On the current index, 3.1 Pro scores 48. That doesn't make Gemini useless — its Flash models are excellent value, and its integration with Google Search, Docs and Drive is genuinely convenient — but on raw reasoning power it has fallen behind Anthropic, OpenAI, and even several Chinese models. If your mental model still has Gemini locked in the top three, it's out of date.
The Price Twist
Raw intelligence isn't the whole story, because the top models cost very different amounts. Claude Opus 5 is $5 per million input tokens and $25 per million output; GPT-5.6 Sol is $5 and $30. Meanwhile Qwen 3.8-Max and Grok 4.5 deliver a score in the high 50s for just $2 in and $6 out — roughly a quarter of the flagships' output price. That's the real shape of the 2026 market: you pay a steep premium for the last few points of intelligence at the very top, while 'near-frontier for a fraction of the cost' is now a crowded, competitive tier. For a lot of real work, the value models are the smarter buy, and only the hardest tasks justify paying flagship rates.
| Rank | Model | Score (AA) | Price /M (in-out) |
|---|---|---|---|
| 1 | Claude Opus 5 | 63 | $5 / $25 |
| 2 | Claude Fable 5 | 62 | $10 / $50 |
| 3 | GPT-5.6 Sol | 61 | $5 / $30 |
| 4 | Kimi K3 | 60 | Open weights |
| 5 | Qwen 3.8-Max | 58 | $2 / $6 |
| 6 | Grok 4.5 | 56 | $2 / $6 |
| - | Gemini 3.1 Pro | 48 | $2 / $12 |
Which One Should You Actually Use?
Translate the leaderboard into decisions. For the hardest reasoning, best writing, and serious coding or agent work, Claude Opus 5 is the current standard-setter, with GPT-5.6 Sol a near-equal alternative — pick by ecosystem preference. For everyday chat, research and writing where you don't need the absolute top, GPT-5.6's cheaper Luna tier or Claude's Sonnet tier give you most of the quality for far less. For high-volume or cost-sensitive work, Qwen 3.8-Max and Grok 4.5 offer near-frontier scores at a quarter of the price, and Grok specifically wins when you need live, real-time information from X. For anything you want to run yourself or self-host, the open-weight models Kimi K3, GLM-5.2 and DeepSeek V4-Flash are the picks. And Gemini earns its place mainly if you live in Google's apps and value that integration over raw horsepower.
- Top of the independent index (Aug 2026): Claude Opus 5 (63), Fable 5 (62), GPT-5.6 Sol (61).
- Value tier: Kimi K3 (60), Qwen 3.8-Max (58), Grok 4.5 (56) - near-frontier for far less money.
- Gemini 3.1 Pro sits at 48 - Google hasn't shipped a new flagship, only a faster Flash model.
- Price gap is huge: top models are $5/$25-$30; value models are $2/$6 per million tokens.
- Best overall: Claude Opus 5 or GPT-5.6 Sol. Best value: Qwen 3.8-Max or Grok 4.5. Best real-time: Grok 4.5.
- Scores are from Artificial Analysis's August 2026 index and shift as testing methods evolve.
01What is the best AI model right now?
By Artificial Analysis's independent August 2026 index, Claude Opus 5 is the highest-scoring model at 63, with Claude Fable 5 (62) and GPT-5.6 Sol (61) close behind. For most people the practical choice is between Claude Opus 5 and GPT-5.6 - they're effectively neck and neck.
02Is Gemini still one of the best AI models?
Not on raw intelligence. Google's flagship Gemini 3.1 Pro scores 48 on the latest index, well behind the top tier, because Google hasn't shipped a new flagship - only the faster Gemini 3.6 Flash. Gemini remains great value and convenient inside Google's apps, but it has slipped on pure reasoning power.
03What's the best cheap AI model?
Qwen 3.8-Max and Grok 4.5 both score in the high 50s on the independent index at about $2 in / $6 out per million tokens - roughly a quarter of the flagship output price. Open-weight models like Kimi K3 are also excellent value if you can self-host.
04Is Claude better than ChatGPT in 2026?
On the independent index Claude Opus 5 (63) edges out GPT-5.6 Sol (61), but the gap is small enough that they're effectively tied for most uses. Claude is often preferred for writing and coding; GPT-5.6 for its broader ecosystem. Choose by preference and price, not the one-point difference.
05Why do the scores differ from older rankings?
Benchmark scores shift over time as testing methods are refined and models are updated. These come from Artificial Analysis's August 2026 index, so treat them as a current snapshot of how the models compare with each other rather than a fixed, permanent ranking.
The 2026 leaderboard rewards a simple habit: don't assume, check — the 'obvious best' model changes every few months, and the gap between the top and the value tier is now small enough that the right choice is genuinely about your task and budget, not brand loyalty. The only way to know which model fits your work is to run the same prompt through a few of them. LumiChats makes that easy, putting many of these leading models under one login at a pay-per-day price, so you can compare the top of this leaderboard head-to-head and pay only for the days you use it.
