August 2026 – Top AI Models and LLM Leaderboards
Updated Friday, August 21, 2026: There is no honest one-model answer this month. Claude Opus 5 is the best all-around model in the latest Artificial Analysis index. GPT-5.6 Sol has the highest verified ARC-AGI-2 score. Claude Opus 5 is far ahead on ARC-AGI-3. Claude Fable 5 still leads ARC-AGI-1. Gemini 3.7 Flash is the price-performance shocker.
- Best overall frontier model: Claude Opus 5. It scores 63 on the current Artificial Analysis Intelligence Index, ahead of Claude Fable 5 at 62 and GPT-5.6 Sol and Grok 4.6 at 61. Source: Artificial Analysis.
- Best on ARC-AGI-2: GPT-5.6 Sol at max reasoning, with a verified score of 92.5%. Source: ARC Prize.
- Best on ARC-AGI-3: Claude Opus 5 at high reasoning, with a verified score of 30.16%. Source: ARC Prize.
- Best value among the frontier models: Gemini 3.7 Flash scores 84.6% on ARC-AGI-2 at just $0.25 per task. Source: ARC Prize.
- Best low-cost open-weight result: DeepSeek V4 Flash 0731 scores 61.4% on ARC-AGI-2 at $0.04 per task. Source: ARC Prize.
My read: Claude Opus 5 gets the overall crown this month because it leads the broad Artificial Analysis index and dominates the newer interactive ARC-AGI-3 benchmark. GPT-5.6 Sol is the better answer when the job is coding-heavy or when raw ARC-AGI-2 performance matters most. Gemini 3.7 Flash is the model I would test first for high-volume production work where speed and cost matter. Grok 4.6 has returned xAI to the frontier, but it does not lead the verified ARC tables.
August 2026 Most Powerful Models According to ARC-AGI
The table below uses one top-performing reasoning variant per model or model family and is ranked by verified ARC-AGI-2 score. ARC-AGI-1 and ARC-AGI-2 test static abstract reasoning. ARC-AGI-3 is a much newer interactive benchmark that tests whether an agent can explore, learn and adapt inside unfamiliar environments. These are useful signals, not a complete measurement of writing quality, speed, product experience, tool reliability or real-world usefulness. Source: ARC-AGI-2 methodology and ARC Prize leaderboard.
| Rank | AI system | Company | Reasoning level | ARC-AGI-1 | ARC-AGI-2 | ARC-AGI-3 | ARC-AGI-2 cost/task |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | OpenAI | Max | 96.5% | 92.5% | 7.78% | Not stated |
| 2 | Claude Opus 5 | Anthropic | Max* | 97.5% | 90.4% | 30.16%* | Not stated |
| 3 | Claude Fable 5 | Anthropic | Max | 98.5% | 89.2% | Not tested | $5.45 |
| 4 | Gemini 3.7 Flash | High | 95.5% | 84.6% | Not tested | $0.25 | |
| 5 | GPT-5.6 Terra | OpenAI | Max | 96.5% | 83.9% | 0.80% | Not stated |
| 6 | Grok 4.6 | SpaceXAI | XHigh | 87.0% | 67.1% | 2.11% | $0.76 |
| 7 | DeepSeek V4 Flash 0731 | DeepSeek | Max | 89.0% | 61.4% | Not tested | $0.04 |
| 8 | DeepSeek V4 Pro 0813 | DeepSeek | Max | 90.0% | 61.3% | Not tested | $0.60 |
| 9 | Kimi K3 | Moonshot AI | Max | 94.5% | 60.4% | Testing | $1.59 |
| 10 | GPT-5.6 Luna | OpenAI | Max | 90.7% | 59.6% | Not tested | $0.18 |
*ARC Prize tested Opus 5 at Max effort on ARC-AGI-1 and ARC-AGI-2, but only at High effort on ARC-AGI-3 because of the short testing window. Costs marked “Not stated” were not provided in the corresponding ARC Prize result summary, so I have not estimated them.
August 2026 AI Model Releases
- August 6: OpenAI updated GPT-5.6 Sol in ChatGPT for more focused responses and more reliable factual answers. GPT-5.6 Luna began rolling out as the default model for Free and Go users, with unlimited text chats and a Think button. Source: OpenAI.
- August 12: SpaceXAI released Grok 4.6 for coding, agentic work and long-running tasks. The API launch price is $2 per million input tokens and $6 per million output tokens for prompts below 200,000 tokens. Source: SpaceXAI.
- August 13: Google released Gemini 3.7 Flash as a stable production model with a 1,048,576-token input limit, 65,536-token output limit and low, medium and high thinking levels. Source: Google AI for Developers.
- August 13: DeepSeek released V4 Pro 0813. Its official model card says the new checkpoint supersedes the preview and improves agentic performance, especially in production environments. Source: DeepSeek model card.
Claude Opus 5 Is the Best Overall Model in August
Claude Opus 5 is the cleanest overall winner this month. Artificial Analysis scores Opus 5 Max at 63 on its Intelligence Index, with Claude Fable 5 at 62 and GPT-5.6 Sol and Grok 4.6 tied at 61. That index combines nine evaluations covering professional work, tool use, coding, science, general knowledge and reasoning. Source: Artificial Analysis.
The more interesting result is ARC-AGI-3. Opus 5 scored 30.16%, while GPT-5.6 Sol scored 7.78% and Grok 4.6 scored 2.11%. ARC-AGI-3 is still new, but the gap is too large to ignore. It suggests Opus 5 is unusually good at learning the rules of unfamiliar interactive environments instead of only solving static puzzles. Source: ARC Prize.
Anthropic positions Opus 5 as an everyday frontier model that approaches Fable 5 capability at half the price. The company says it improves coding, long-running agents and professional knowledge work while using the same list price as Opus 4.8. Source: Anthropic.
GPT-5.6 Sol Is the ARC-AGI-2 and Coding Leader
GPT-5.6 Sol has the strongest verified ARC-AGI-2 result at 92.5% with Max reasoning. It also became the first model to win an ARC-AGI-3 public game and scored 7.78% across the semi-private ARC-AGI-3 set. Source: ARC Prize.
OpenAI launched the GPT-5.6 family on July 9 with Sol as the flagship, Terra as the balanced middle option and Luna as the fast, lower-cost model. Sol supports reasoning levels up to Max, while OpenAI’s Ultra setting coordinates multiple agents for especially demanding work. The standard API price for Sol is $5 per million input tokens and $30 per million output tokens. Source: OpenAI.
For coding, Sol remains the strongest public result I found. Artificial Analysis reported a score of 80 on its Coding Agent Index at launch, ahead of Claude Fable 5, while using fewer output tokens and less time. Source: Artificial Analysis.
Gemini 3.7 Flash Is the Price-Performance Winner
Gemini 3.7 Flash is the result that should make every model buyer pay attention. At High reasoning it scores 84.6% on ARC-AGI-2 for $0.25 per task. That puts it only 7.9 points behind GPT-5.6 Sol Max on this benchmark while costing far less per verified task than the other frontier entries with published ARC costs. Source: ARC Prize.
Google lists Gemini 3.7 Flash as a stable model with text, image, video, audio and PDF input. It supports code execution, function calling, search grounding, structured output and a 1,048,576-token input window. Source: Google AI for Developers.
The practical takeaway is straightforward: if a workflow runs thousands of times, test Gemini 3.7 Flash before paying flagship-model prices. The raw ARC result is strong enough that the old assumption that “Flash” automatically means second-tier intelligence no longer holds.
Grok 4.6 Puts SpaceXAI Back in the Frontier Group
Grok 4.6 is a major improvement over Grok 4.5. Artificial Analysis gives it a score of 61, tied with GPT-5.6 Sol and one point behind Claude Fable 5. Its $2 input and $6 output pricing per million tokens is substantially below Opus 5 and GPT-5.6 Sol. Source: Artificial Analysis.
ARC Prize verified Grok 4.6 at 67.1% on ARC-AGI-2 and 2.11% on ARC-AGI-3 at XHigh reasoning. That is a real step forward, but it also shows why a broad leaderboard and a single reasoning benchmark can tell different stories. Source: ARC Prize.
SpaceXAI says the model was trained for longer-running agents, interactive visual work, research, codebase analysis and polished application development. It is available through the SpaceXAI API, Cursor, Grok Build and other partners. Source: SpaceXAI.
DeepSeek Owns the Low-Cost Open-Weight Lane
DeepSeek V4 Flash 0731 scores 61.4% on ARC-AGI-2 for only $0.04 per task. DeepSeek V4 Pro 0813 scores 61.3% for $0.60. The Flash result is slightly higher on ARC-AGI-2, while the newer Pro checkpoint is positioned for stronger agentic and production work. Source: ARC Prize V4 Flash result and ARC Prize V4 Pro result.
DeepSeek’s official V4 Pro 0813 model card reports improvements over its preview model and publishes the model through Hugging Face. This makes DeepSeek one of the most important options for teams that want open weights, local control or a lower-cost path than the closed frontier APIs. Source: DeepSeek model card.
What Happened to Claude Fable 5?
Fable 5 did not disappear. It still leads ARC-AGI-1 at 98.5% and scores 89.2% on ARC-AGI-2. Anthropic describes Fable as a Mythos-class model made safe for general use, with broader capability than its Opus tier. Source: ARC Prize and Anthropic.
Why give Opus 5 the overall crown? Opus 5 now edges Fable 5 on the broad Artificial Analysis index, costs less and is designed as the everyday model for Claude users. Fable 5 remains a top-end option for the hardest work, but Opus 5 has the stronger all-around August case. Source: Anthropic and Artificial Analysis.
Which AI Model Should You Use?
- For the best all-around frontier performance: Claude Opus 5.
- For coding and the hardest static reasoning tasks: GPT-5.6 Sol.
- For high-volume production agents: Gemini 3.7 Flash.
- For a lower-cost frontier alternative: Grok 4.6.
- For open weights and very low cost: DeepSeek V4 Flash or V4 Pro.
- For free everyday ChatGPT use: GPT-5.6 Luna, which OpenAI is rolling out as the default for Free and Go users. Source: OpenAI.
The model race is no longer a two-company horse race. Anthropic owns the broad overall lead and the strongest ARC-AGI-3 result. OpenAI leads ARC-AGI-2 and coding. Google has the most surprising price-performance result. SpaceXAI is back in the frontier pack. DeepSeek remains a serious open-weight and cost competitor.
Follow us on Twitter/X for updates to this post, or subscribe to the newsletter below. I will keep updating the rankings as new verified results land.
Sources checked August 21, 2026: ARC Prize verified leaderboard; Artificial Analysis Grok 4.6 analysis; OpenAI GPT-5.6 release; Anthropic Opus 5 release; Google Gemini 3.7 Flash documentation; SpaceXAI Grok 4.6 release; and DeepSeek V4 Pro 0813 model card.
April 2026 Top Model Announcement: Thursday April 23, 2026: Chat GPT 5.5 released at approximately 2:30pm PDT 5:30 pm EDT
OpenAI released ChatGPT GPT-5.5 on April 23, 2026 around 2:30pm PDT. From the release:
“GPT‑5.5 understands what you’re trying to do faster and can carry more of the work itself. It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Instead of carefully managing every step, you can give GPT‑5.5 a messy, multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going.”
ChatGPT 5.5 is the new most powerful model, beating Opus 4.7 across most benchmarks.
This table was shared by OpenAI:
GPT-5.5 | GPT-5.4 | GPT-5.5 Pro | GPT-5.4 Pro | Claude Opus 4.7 | Gemini 3.1 Pro | |
| Terminal-Bench 2.0 | 82.7% | 75.1% | – | – | 69.4% | 68.5% |
| Expert-SWE (Internal) | 73.1% | 68.5% | – | – | – | – |
| GDPval (wins or ties) | 84.9% | 83.0% | 82.3% | 82.0% | 80.3% | 67.3% |
| OSWorld-Verified | 78.7% | 75.0% | – | – | 78.0% | – |
| Toolathlon | 55.6% | 54.6% | – | – | – | 48.8% |
| BrowseComp | 84.4% | 82.7% | 90.1% | 89.3% | 79.3% | 85.9% |
| FrontierMath Tier 1–3 | 51.7% | 47.6% | 52.4% | 50.0% | 43.8% | 36.9% |
| FrontierMath Tier 4 | 35.4% | 27.1% | 39.6% | 38.0% | 22.9% | 16.7% |
| CyberGym | 81.8% | 79.0% | – | – | 73.1% | – |
OpenAI GPT-5.5 Artificial Analysis Intelligence Index Chart
OpenAI also released this chart from ArtificialAnalysis and noted on the capabilities of the model:
“Across these domains, GPT‑5.5 is not just more intelligent; it is more efficient in how it works through problems, often reaching higher-quality outputs with fewer tokens and fewer retries. On Artificial Analysis’s Coding Index, GPT‑5.5 delivers state-of-the-art intelligence at half the cost of competitive frontier coding models.”

The Artificial Analysis Intelligence Index is a weighted average of 10 evals ran by an external party: AA-LCR, AA-Omniscience, CritPt, GDPval-AA, GPQA Diamond, Humanity’s Last Exam, IFBench, SciCode, Terminal-Bench Hard, τ²-Bench Telecom.
Follow us on Twitter/X to get the latest updates to this post, or subscribe in our newsletter below.
April 2026 – Top AI Models and LLM Leaderboards
- According to the latest benchmarks, ChatGPT GPT-5.5 is the most powerful generally available model.
April 2026 Most Powerful Models According to Arc-2 Leaderboard
We’ve been checking AI model leaderboards on Arc-2 so far, even though Arc-3 has been released, so we will continue to track on Arc-2 for the time being for lookback comparisons. At present GPT-5.5 is at the top of the board.
| AI System | Author | Date | System Type | ARC-AGI-1 | ARC-AGI-2 | Cost/Task |
|---|---|---|---|---|---|---|
| GPT-5.5 Pro (High) | OpenAI | 2026-04-23 | CoT | 96.50% | 84.60% | $10.51 |
| GPT-5.5 Pro (xHigh) | OpenAI | 2026-04-23 | CoT | 95.00% | 84.20% | $10.76 |
| GPT-5.5 (Low) | OpenAI | 2026-04-22 | CoT | 76.20% | 33.30% | $0.35 |
| GPT-5.5 (Medium) | OpenAI | 2026-04-22 | CoT | 92.20% | 70.40% | $0.86 |
| GPT-5.5 (High) | OpenAI | 2026-04-22 | CoT | 94.50% | 83.30% | $1.45 |
| GPT-5.5 (xHigh) | OpenAI | 2026-04-22 | CoT | 95.00% | 85.00% | $1.87 |
| Claude 4.7 (Low) | Anthropic | 2026-04-16 | CoT | 91.00% | 62.10% | $2.38 |
| Claude 4.7 (Medium) | Anthropic | 2026-04-16 | CoT | 91.00% | 67.50% | $2.96 |
| Claude 4.7 (High) | Anthropic | 2026-04-16 | CoT | 93.50% | 68.30% | $3.17 |
| Claude 4.7 (Max) | Anthropic | 2026-04-16 | CoT | 92.00% | 75.80% | $7.43 |
| GPT-5.4 Mini (xHigh) | OpenAI | 2026-03-17 | CoT | 63.70% | 18.90% | $0.75 |
| GPT-5.4 Mini (High) | OpenAI | 2026-03-17 | CoT | 58.00% | 13.20% | $0.56 |
| GPT-5.4 Mini (Medium) | OpenAI | 2026-03-17 | CoT | 40.80% | 4.40% | $0.33 |
| GPT-5.4 Mini (Low) | OpenAI | 2026-03-17 | CoT | 13.00% | 1.10% | $0.06 |
| GPT-5.4 Nano (xHigh) | OpenAI | 2026-03-17 | CoT | 51.50% | 5.70% | $0.16 |
| GPT-5.4 Nano (High) | OpenAI | 2026-03-17 | CoT | 38.20% | 3.60% | $0.13 |
| GPT-5.4 Nano (Medium) | OpenAI | 2026-03-17 | CoT | 33.00% | 1.90% | $0.06 |
| GPT-5.4 Nano (Low) | OpenAI | 2026-03-17 | CoT | 18.30% | 1.50% | $0.01 |
| Grok 4.20 (Reasoning) | xAI | 2026-03-09 | CoT | 89.50% | 65.10% | $0.92 |
| Grok 4.20 (Beta Reasoning) | xAI | 2026-03-05 | CoT | N/A | N/A | N/A |
April 2026 AI Model Releases
- Thursday April 23, 2026: Chat GPT 5.5 released at approximately 2:30pm PDT 5:30 pm EDT
- ChatGPT 5.5 is the new most powerful model, beating Opus 4.7 across most benchmarks
- April 16, 2026: Anthropic Claude Opus 4.7 released. , this is a general-purpose agentic model designed for complex, long-running workflows.
- Google Gemini 3.1 Pro: This model was released in early 2026. It features 1M-token context windows and leads in multimodal reasoning across voice, vision, and video
Claude Opus 4.7 Release Highlights
“Opus 4.7 is a notable improvement on Opus 4.6 in advanced software engineering, with particular gains on the most difficult tasks.”
Anthropic included this chart as part of their release, showing Opus 4.7 beating Opus 4.6, ChatGPT GPT-5.4 and Google Gemini 3.1 Pro across benchmarks like Agentic coding, Multidisciplinary reasoning (Humanity’s Last Exam”, Agentic search, Scaled tool use and more. They also hinted at the power of Mythos Preview which is not generally available.

My personal experience with Claude Opus 4.7 has generally been good. For me, Opus 4.6 was a watershed moment that was personally transformative, and I’m not exaggerating. I have shifted most workflows from ChatGPT to Claude. There have been complaints around Opus 4.7 Adaptive Thinking as that part has been unclear to people.

For example, on this Reddit thread, the poster asks about how Adaptive Thinking actually works, posting:
“My question: if I toggle Adaptive thinking ‘Off’, does that mean that Opus 4.7 is defaulted to max thinking regardless of the prompt? Or if Adaptive thinking is toggled ‘Off’, does that mean I am just getting a non-thinking version?
This is confusing since the toggle used to be “Extended thinking”.”
With the consensus being:
- Adaptive Thinking is the new Extended Thinking, but they can choose to ignore it, “Extended thinking you could leave on to ensure heavier reasoning. There’s no way to ensure it now.” (Reddit)
- If you toggle it off? Looks like that means don’t think at all, or think fast. One commentor phrased it as “amazingly, it looks more and more like your concerns are right and it’s either “never think hard”(off) or “guess whether or not you should think hard”(on) which is a terrible set of options.” (Reddit)
What’s the argument for Adaptive Thinking? This commentor brings up a good point:
“Might not be a bad thing. There was that paper a while ago that showed reasoning models actually performs worse for simple tasks, vs better for complex tasks. If we assume the software can adequately distinguish, it might be better than forcing reasoning for everything.”
Wharton Professor Ethan Molick, renown thought leader in the world of AI models and business economics, brings a balanced take to this:
Note that Ethan thinks Adaptive has gotten better in the 2 weeks since release, as on April 13, 2026 he noted Adaptive was disliked because there’s no manual override:
Will we adapt to Adaptive Thinking? Or will it be like the Apple Touch Bar and be universally despised until they rolled it back?
March 2026 – Top AI Models and LLM Leaderboards
What are the top AI models? Here is the consensus on top models for March 2026 as well as new releases.
- Models seem to be converging in general, with a general triopoly forming with OpenAI, Anthropic, and Gemini taking strong leads and market share. xAI is trying hard to make it a quadopoly, but they and Meta’s LLM seem to be falling behind in this horse race. But as they say, don’t be against Elon.
- Consensus zeitgeist read: 5.4 Codex is now the most powerful coding model, and has more generous limits than Claude Code.
- Users are finding they hit limits in Claude Opus 4.6 much faster than equivalent models, but that seems to be resolving more and more.
- Personal opinion and preference: I’m using Claude Opus 4.6 as my daily driver in Claude Desktop with CoWork and Chat being 80% of my usage. I will still constantly test with ChatGPT about 20% of the time, and Gemini and Grok occasionally when I really want to generate fast across 4 models at once.
March 2026 Most Powerful Models According to Arc-2 Leaderboard
At present, Gemini 3 Deep Think (2/26) and GPT-5.4 Pro (xHigh) are the most powerful models according to the ARC-AGI-2 test. While this test isn’t fully relevant to everyday users, who value UI and experience, it shows which labs are currently on the cutting edge, with Gemini & OpenAI neck-and-neck in the race for the most powerful raw model. Claude/Anthropic seems to be pivoting to focusing on having the best coding model, rolling out CoWork further, and integrating into workflows with excellent UIs.
| AI System | Author | System Type | ARC-AGI-1 | ARC-AGI-2 | Cost/Task | Code / Paper |
|---|---|---|---|---|---|---|
| Human Panel | Human | N/A | 98.0% | 100.0% | $17.00 | — |
| Gemini 3 Deep Think (2/26) | CoT | 96.0% | 84.6% | $13.62 | — | |
| GPT-5.4 Pro (xHigh) | OpenAI | CoT | 94.5% | 83.3% | $16.41 | — |
| Gemini 3.1 Pro (Preview) | CoT | 98.0% | 77.1% | $0.962 | 📄 | |
| GPT-5.4 (xHigh) | OpenAI | CoT | 93.7% | 74.0% | $1.52 | — |
| GPT-5.2 (Refine.) | Johan Land | Refinement | 94.5% | 72.9% | $38.99 | 💻 |
| Claude Opus 4.6 (120K, High) | Anthropic | CoT | 94.0% | 69.2% | $3.47 | — |
| Claude Opus 4.6 (120K, Max) | Anthropic | CoT | 93.0% | 68.8% | $3.64 | — |
| GPT-5.4 (High) | OpenAI | CoT | 92.7% | 67.5% | $1.02 | — |
| Claude Opus 4.6 (120K, Medium) | Anthropic | CoT | 92.0% | 66.3% | $2.72 | — |
March 2026 Releases:
- OpenAI Release March 6, 2026: ChatGPT 5.4 released – OpenAI’s most powerful model
Ethan Mollick, author of Co-Intelligence and professor at Wharton, sees GPT-5.4 Pro as in a class on its own. He uses Claude Opus 4.6 as his daily driver (as do I), but 5.4 Pro when he needs to crack something hairy:
February 2026 Top Models:
- Claude Opus 4.6 is a sea-change model for Anthropic, seeing unprecedented adoption.
- Consensus zeitgeist read: 5.3 Codex is now the most powerful coding model, and has more generous limits than Claude Code.
- Users are finding they hit limits in Claude Opus 4.6 much faster than GPT-5.2 (even w/Thinking).
February 2026 Releases:
- OpenAI Release February 5, 2026: GPT-5.3-Codex – OpenAI’s new powerful model for Codex. Developers are loving this and switching from Claude Code.
- OpenAI Release February 12, 2026: GPT-5.3-Codex-Spark – 1,000 tokens per second. Game is changed. Move is on Claude Code to catch up.
- Anthropic Release February 6, 2026: Claude Opus 4.6 released to much fanfare with a 1m token context window in beta.
- Anthropic Release February 17, 2026: Claude Sonnet 4.6 launch
- Google Gemini Release Feb 12, 2026: Gemini 3 Deep Think – specialized reasoning mode – pushing the frontier and establishing new benchmarks.
- xAI Grok 4.20 released February 16, 2026: See Grok Models here
- xAI acquired by SpaceX in trillion dollar+ combined valuation. X employees ecstatic at SpaceX stock conversion.
New Year 2026 AI Model Rankings
January 2026 Highlights & Releases:
- ChatGPT 5.2 is the most popular and powerful general-purpose model with generous limits
- Claude Opus 4.5 in Claude Code is seen as another creature, as the penultimate AI coding experience. Some are saying that “coding is solved”
- Claude is rolling out more and more Claude Code improvements with Skills being loved by all
- ClaudeBot is going to the stratosphere. Mac Minis are flying off the shelf for the iMessage capabilities. Things may never be the same.
- xAI remains a favorite for being more uncensored. Grokipedia is a fascinating experiment and a challenger to Wikipedia, powered by Grok. xAI is the dark horse to win it all.
- Some say that Google’s Gemini is inevitable and will win it all based on Google’s distribution.
Arc Prize 2026 Ranking Update
There are a lot of different ranking systems, but the Arc Prize is a great one to start with as a definitive source of LLM leaderboard rankings. See our post on AI ranking factors for more intel.
As of January 6, 2026, these are the top-ranked LLM models:
| Rank | AI System | Author | System Type | ARC-AGI-1 | ARC-AGI-2 | Cost/Task |
| 1 | Human Panel | Human | N/A | 98.00% | 100.00% | $17.00 |
| 2 | GPT-5.2 Pro (High) | OpenAI | CoT | 85.70% | 54.20% | $15.72 |
| 3 | Gemini 3 Pro (Refine.) | Poetiq | Refinement | N/A | 54.00% | $30.57 |
| 4 | GPT-5.2 (X-High) | OpenAI | CoT | 86.20% | 52.90% | $1.90 |
| 5 | Gemini 3 Deep Think (Preview) ² | CoT | 87.50% | 45.10% | $77.16 | |
| 6 | GPT-5.2 (High) | OpenAI | CoT | 78.70% | 43.30% | $1.39 |
| 7 | GPT-5.2 Pro (Medium) | OpenAI | CoT | 81.20% | 38.50% | $8.99 |
| 8 | Opus 4.5 (Thinking, 64K) | Anthropic | CoT | 80.00% | 37.60% | $2.40 |
| 9 | Gemini 3 Flash Preview (High) | CoT | 84.70% | 33.60% | $0.23 | |
| 10 | Gemini 3 Pro | CoT | 75.00% | 31.10% | $0.81 |
The ARC Prize is a solid ranking, however there are some downsides:
- Does not measure the speed of these models
- Does not rank by a combination of score, speed, and cost
- Does not show the limitations of requests for each of these models
For both personal and business users, these three factors are crucial for understanding their role within business process pipelines.
For example, GPT-5.2 (X-High) is not a model widely accessible without extra configuration. This Reddit user complained:
“Unless you have 10-30 minutes for each task you give it, this model is useless.
I would rather use less smart model like Gemini 3 pro that can do things like 10 times faster.
The only use case i can think of either doing things on background. Like walking outside or going to the gym and typing what the model should do, and then when you come back you look at the results.”
Orca’s 2025 AI Models Report
A new report by Orca ranked the “10 most popular AI models of 2025 based on analysis of billions of cloud assets by the Orca Research Pod.”

One issue I may have with the report is any lack of GPT-5 mention or adoption. Perhaps enterprises are reluctant to implement the latest model in production, but given that it was released August 7, 2025 I would have expected to see it rank somewhere on the list. It’s possible that it was #11, and just didn’t make the top 10.
Opinions & Predictions for the Best AI Models by the End of 2026
Editorial opinions by Joe Robison, founder of Green Flag Digital, for 2026:
- We will see a race to real-time virtual personal assistants.
- OpenAI will retain the crown for the best general-purpose AI at wide distribution. Others like Claude Code and
- OpenAI will become a mega-search crawler on their own, as large as Google’s crawler using different techniques.
- OpenAI will purchase crawling, indexing, ranking companies, including Exa.ai.
- Apple or Google will buy Perplexity.
- Lovable will be acquired in a huge deal by Google or Anthropic.
- Replit will be on the IPO path for 2027.
- The AI model companies will end up like the big cloud computing companies, carving up 4-5 distinct domains
- Prediction markets will roar between 2026-2030 and grow massively. 2026 will be an explosive year of growth.
- AI models will continue to be embedded in the real world more and more, as seen recently in drive-thru order windows such as Carl’s Jr.
Polymarket AI model predictions for 2026:
Polymarket, Kalshi, and other prediction markets are emerging as massive “truth machines” with potentially transformative, Minority-Report meets Jason Borne-esque ramifications for 2026-2030.
As it stands on Monday Jan 5, 2026 at 6:18pm PST, Kalshi gamblers predict that Google will have the best AI model at the end of January 2026. See this bet here: What will be the top AI model this month?

Are prediction markets reliable for predicting the best AI models? They’re not perfect. However, they do incentivize betters to reveal the truth with their stakes. If you feel strongly about a prediction – such as what the best model is – you can place your bet and earn a monetary reward.
Just like sports betting, it’s a gamble. But they may point us in the right direction and show what the consensus view is.
You can then decide to bet with the consensus, or against it. And the truth will be revealed in the fullness of time.
And right now, the consensus is that Google will have the best model by January 31, 2026. Personally, I don’t buy it. I think OpenAI is cooking up the next GPT model – they have to. They’re likely releasing GPT 5.3
Previous AI Model Rankings – and Changes over Time
October 26, 2025 Rankings
Just eyballing it, but GPT-5 Pro is the current top AI model in production, only beaten by custom AIs and a human panel.
Grok 4 (Thinking) is the #2 production model, with a very solid cost/task of $2.17 compared to GPT-5 Pro’s $7.14 a task for pretty close scores. Should likely be used in production a lot, and more often soon.
Claude Sonnet 4.5 (Thinking 32K) is right behind the other two, with an astonishingly low $0.76 a task, lower than Grok 4 (Thinking) by a factor of 3. This should be used even more frequently in production as a strong default AI.
Fresh pull of the rankings today, here are the top 20 as of now:
| AI System | Organization | System Type | ARC-AGI-1 | ARC-AGI-2 | Cost/Task |
| Human Panel | Human | N/A | 98.00% | 100.00% | $17.00 |
| J. Berman (2025) | Bespoke | CoT + Synthesis | 79.60% | 29.40% | $30.40 |
| E. Pang (2025) | Bespoke | CoT + Synthesis | 77.10% | 26.00% | $3.97 |
| GPT-5 Pro | OpenAI | CoT | 70.20% | 18.30% | $7.14 |
| Grok 4 (Thinking) | xAI | CoT | 66.70% | 16.00% | $2.17 |
| Claude Sonnet 4.5 (Thinking 32K) | Anthropic | CoT | 63.70% | 13.60% | $0.76 |
| GPT-5 (High) | OpenAI | CoT | 65.70% | 9.90% | $0.73 |
| Claude Opus 4 (Thinking 16K) | Anthropic | CoT | 35.70% | 8.60% | $1.93 |
| GPT-5 (Medium) | OpenAI | CoT | 56.20% | 7.50% | $0.45 |
| Claude Sonnet 4.5 (Thinking 8K) | Anthropic | CoT | 46.50% | 6.90% | $0.24 |
| Claude Sonnet 4.5 (Thinking 16K) | Anthropic | CoT | 48.30% | 6.90% | $0.35 |
| o3 (High) | OpenAI | CoT | 60.80% | 6.50% | $0.83 |
| Tiny Recursion Model (TRM) | Bespoke | N/A | 40.00% | 6.30% | $2.10 |
| o4-mini (High) | OpenAI | CoT | 58.70% | 6.10% | $0.86 |
| Claude Sonnet 4 (Thinking 16K) | Anthropic | CoT | 40.00% | 5.90% | $0.49 |
| Claude Sonnet 4.5 (Thinking 1K) | Anthropic | CoT | 31.00% | 5.80% | $0.14 |
| Grok 4 (Fast Reasoning) | xAI | CoT | 48.50% | 5.30% | $0.06 |
| o3-Pro (High) | OpenAI | CoT + Synthesis | 59.30% | 4.90% | $7.55 |
| Gemini 2.5 Pro (Thinking 32K) | CoT | 37.00% | 4.90% | $0.76 | |
| Claude Opus 4 (Thinking 8K) | Anthropic | CoT | 30.70% | 4.50% | $1.16 |
View the entire leaderboard here at ARCprize.
Sep 18, 2025 AI Leaderboard Rankings
Table recreated courtesy of ARC Prize, a nonprofit.
This table shows the latest rankings following ARC 1 and 2 tests.
| Rank | AI System | Organization | System Type | ARC-AGI-1 | ARC-AGI-2 |
| 1 | Human Panel | Human | N/A | 98.00% | 100.00% |
| 2 | J. Berman (2025) | Bespoke | CoT + Synthesis | 79.60% | 29.40% |
| 3 | E. Pang (2025) | Bespoke | CoT + Synthesis | 77.10% | 26.00% |
| 4 | Grok 4 (Thinking) | xAI | CoT | 66.70% | 16.00% |
| 5 | GPT-5 (High) | OpenAI | CoT | 65.70% | 9.90% |
| 6 | Claude Opus 4 (Thinking 16K) | Anthropic | CoT | 35.70% | 8.60% |
| 7 | GPT-5 (Medium) | OpenAI | CoT | 56.20% | 7.50% |
| 8 | o3 (High) | OpenAI | CoT | 60.80% | 6.50% |
| 9 | o4-mini (High) | OpenAI | CoT | 58.70% | 6.10% |
| 10 | Claude Sonnet 4 (Thinking 16K) | Anthropic | CoT | 40.00% | 5.90% |
| 11 | o3-Pro (High) | OpenAI | CoT + Synthesis | 59.30% | 4.90% |
| 12 | Gemini 2.5 Pro (Thinking 32K) | CoT | 37.00% | 4.90% | |
| 13 | Claude Opus 4 (Thinking 8K) | Anthropic | CoT | 30.70% | 4.50% |
| 14 | GPT-5 Mini (High) | OpenAI | CoT | 54.30% | 4.40% |
| 15 | Gemini 2.5 Pro (Thinking 16K) | CoT | 41.00% | 4.00% | |
| 16 | GPT-5 Mini (Medium) | OpenAI | CoT | 37.30% | 4.00% |
| 17 | o3-preview (Low)* | OpenAI | CoT + Synthesis | 75.70% | 4.00% |
| 18 | Gemini 2.5 Pro (Preview) | CoT | 33.00% | 3.80% | |
| 19 | Gemini 2.5 Pro (Preview, Thinking 1K) | CoT | 31.30% | 3.40% | |
| 20 | o3-mini (High) | OpenAI | CoT | 34.50% | 3.00% |
See full table here: ARC Leaderboard
How ARC-AGI-1 Works
“ARC-AGI-1 consists of 800 puzzle-like tasks, designed as grid-based visual reasoning problems. These tasks, trivial for humans but challenging for machines, typically provide only a small number of example input-output pairs (usually around three). This requires the test taker (human or AI) to deduce underlying rules through abstraction, inference, and prior knowledge rather than brute-force or extensive training.”
ARC-AGI-2 Explained:
Here’s a direct quote:
“ARC-AGI-1 was created in 2019 (before LLMs even existed). It endured 5 years of global competitions, over 50,000x of AI scaling, and saw little progress until late 2024 with test-time adaptation methods pioneered by ARC Prize 2024 and OpenAI.
ARC-AGI-2 – the next iteration of the benchmark – is designed to stress test the efficiency and capability of state-of-the-art AI reasoning systems, provide useful signal towards AGI, and re-inspire researchers to work on new ideas.
Pure LLMs score 0%, AI reasoning systems score only single-digit percentages, yet extensive testing shows that humans can solve every task.
Can you create a system that can reach 85% accuracy?”
July 10, 2025 AI Leaderboard Rankings
As of July 10, 2025 Grok 4 is the best AI model, according to ARC Prize’s ARC-AGI Leadersboard.

According to their X announcement:
“Grok 4 (Thinking) achieves new SOTA on ARC-AGI-2 with 15.9% This nearly doubles the previous commercial SOTA and tops the current Kaggle competition SOTA.”
-ARC on X
View the full ARC-AGI Leaderboard page for real-time updates.
According to their team:
“ARC-AGI has evolved from its first version (ARC-AGI-1) which measured basic fluid intelligence, to ARC-AGI-2 which challenges systems to demonstrate both high adaptability and high efficiency.
The scatter plot above visualizes the critical relationship between cost-per-task and performance – a key measure of intelligence efficiency. True intelligence isn’t just about solving problems, but solving them efficiently with minimal resources.”
Other Leaderboards
Kearney Leaderboard: Out of Date
We don’t recommend referencing this one by Kearny, as it mentions o1 as an “up and coming” model, so it’s already out of date.
Updates Log
- April 23, 2026: Full updates for April across the board, with a new leaderboard table. GPT-5.5 announcement featured.
- March 23, 2026: Update with notes and Tweet from Ethan Mollick
- March 19, 2026: Updated with the latest ARC-AGI-2 leaderboard, as well as March releases and notes.
- February 17, 2026: Updated with new releases
- February 6, 2026: Early February updates
- January 6, 2026: January updates
Subscribe for Updates!
Follow us on Twitter/X for the latest updates:
Subscribe to our newsletter below for other updates and news across AI and marketing from Green Flag Digital.
Leave a Reply