Toolradar Research
AI Model Head to Head 2026: Benchmarks vs Adoption
We put the capability leaderboard next to our own adoption data. Claude Opus 5 leads the Artificial Analysis Intelligence Index at 61, but ChatGPT holds 71.9% of the public review corpus and roughly 17 times Claude's review volume. Satisfaction is flat across the field, so reach, not raw intelligence, is deciding this market.

Founder, Toolradar & Dupple
Key findings
What the data shows.
- 01
Benchmarks and adoption disagree. On the Artificial Analysis Intelligence Index (v4.1, August 2026), Claude Opus 5 leads at 61, with GPT-5.6 just behind at 59. Yet in our data, ChatGPT has roughly 17 times more user reviews than Claude (2,204 versus 129). The smartest model on the tests is not the one the market has chosen.
- 02
Intelligence is a photo finish, adoption is a landslide. The top models cluster within a few points on the benchmark. On real-world footprint they are worlds apart: ChatGPT alone carries more review volume than every other assistant combined, and 22% of all software press coverage (from our media analysis).
- 03
One product holds 71.9% of the review corpus. Across the seven assistants we track, ChatGPT accounts for 71.9% of all public reviews. The remaining six split the other 28.1%, and the largest of them holds 8.3%.
- 04
Satisfaction barely separates the leaders. ChatGPT and Claude both average 4.6 out of 5; Gemini, Perplexity, and Copilot all sit at 4.5. The star rating does not pick a winner. Distribution, ecosystem, and fit do.
- 05
Open weights are four points off the frontier. Kimi K3 scores 57 against the leader's 61. On the capability axis that is a smaller gap than most buyers assume they are paying to close.
- 06
Hype is not adoption, and Grok is the proof. Despite relentless coverage, Grok has the lowest user rating of the group (4.2) and a tiny review corpus. A model can dominate the timeline and barely register in real use.
About the research
How we built this report.
Toolradar tool database. Editorial review with weekly pricing verification.
2026. Snapshot taken August 4, 2026. Refresh due Oct 4, 2026.
Public scoring rubric. See how we rate for the full criteria.
Creative Commons BY 4.0. Quote, link, and reuse with attribution.
Every few weeks a new model tops a benchmark and the internet declares a new king. Then you look at what people actually use, and the crown never moves. We put the two scoreboards side by side: the independent capability benchmarks that rank raw intelligence, and our own data on which models people actually review, rate, and adopt. They tell completely different stories. The benchmark leader and the market leader are not the same model, and the gap between them is enormous.
The single most striking number is the ratio. On the Artificial Analysis Intelligence Index (v4.1, August 2026), Claude Opus 5 leads at 61 and GPT-5.6 Sol sits two points behind at 59. In our public review corpus, ChatGPT carries 2,204 reviews against Claude's 129, roughly 17 times more. A two point capability lead coexists with a seventeen fold adoption deficit. Whatever the leaderboard is measuring, it is not what decides which assistant ends up on people's screens.
How we built this report
This report joins two data sources that are usually read in isolation, which is exactly why they contradict each other so cleanly.
Capability figures come from the Artificial Analysis Intelligence Index v4.1 as of August 2026, an independent aggregate of standardized benchmarks. We did not run these tests ourselves and we do not adjust them. Leaderboards move constantly, so treat the numbers as a snapshot of that date rather than a permanent ranking, and check the live index before quoting them as current.
Adoption figures are Toolradar's first-party data: the aggregated public review corpus behind each product, meaning real user ratings and review counts pulled from G2, Capterra, Trustpilot, and the app stores. We use each product's main consumer offering, and ratings are normalized to a 0 to 5 scale so that sources with different native scales stay comparable.
Two limits are worth stating plainly. First, our catalog counts are a lower bound, not a census. They cover the review venues we aggregate, so any product with a large footprint on a platform outside that set is undercounted, and the true absolute volumes are higher than what appears here. What the counts are good for is relative scale, and the relative scale in this data is not subtle enough for that caveat to change any conclusion. A 17 to 1 ratio does not flip because of coverage gaps at the margin. Second, review volume is a lagging signal. It rewards products that have been available longer and to more people, which is precisely the property that makes it a useful counterweight to a benchmark measured under ideal conditions on the day of release.
Every share figure in this report is computed from the review counts shown in the breakdown table, rounded to one decimal place. This report combines a third-party benchmark with our first-party adoption data and is released under CC BY 4.0. Reuse it with a link back to this page.
Scoreboard one: the capability race is nearly over
Independent benchmarks measure one thing: how well a model answers hard test questions under controlled conditions. As of August 2026, the frontier looks like this.
Independent aggregate of standardized benchmarks (v4.1). Higher is better.Artificial Analysis Intelligence Index, August 2026
The chart's most important feature is how flat it is. The entire proprietary frontier fits inside a two point band: 61 for Claude Opus 5 at max effort, 60 for the same model at xhigh effort, 59 for GPT-5.6 Sol. Note what that first comparison implies. Running the same model at a different effort setting moves the score by a point, which is half the distance between the two leading labs. When an internal configuration knob is worth half the gap between vendors, the vendor gap is not a meaningful decision input for most work.
The second feature is where the open-weights models land. Kimi K3 sits at 57, four points behind the leader, and DeepSeek V4 Flash at 50, eleven points back. The four point gap is the one to watch. It means the best open model is closer to the closed frontier than the frontier's own effort settings sometimes are to each other across a full configuration range, and it means the capability argument for paying frontier prices is thinner than it was a year ago.
What this scoreboard cannot tell you is anything about deployment. It has no view of price, latency, rate limits, context handling in long sessions, tool reliability, or whether the model is available inside the software your team already has open. Those are the variables that decide real usage, and the second scoreboard is what they add up to.
Scoreboard two: adoption is not close
Now the same market measured by what people actually signed up for and reviewed.
Aggregated review counts from G2, Capterra, Trustpilot, and the app storesPublic reviews per AI assistant
Everything inverts. Claude, the benchmark leader, ranks fifth of seven on review volume with 129. ChatGPT, which does not top the intelligence index, has 2,204, more than every other assistant on the list combined. The three way cluster in the middle is its own story: Perplexity at 253, Microsoft Copilot at 227, and Google Gemini at 222 are separated by 31 reviews across the whole group, which is inside the noise of how any one of them gets counted. There is no clear second place in this market. There is a first place and then a pack.
Rendering the same distribution as share makes the concentration harder to look away from.
Each product's public reviews as a percentage of the seven product totalShare of the AI assistant review corpus
ChatGPT holds 71.9% of the corpus. That is not a lead, it is a category definition. The gap between first and second (71.9% against 8.3%) is larger than the entire rest of the market put together, and it lines up with the other footprint signal we track independently: ChatGPT accounts for 22% of all software press coverage in our media analysis. Two different measurement systems, one built on user reviews and one on published articles, both put the same product far outside the distribution of its peers.
The tail matters too. Grok at 0.7% and DeepSeek at 0.3% are rounding errors in the review corpus despite both being household names in tech coverage. For DeepSeek there is a benign explanation: it is an open-weights model whose real usage happens through APIs and self-hosting, where nobody writes a Capterra review. For Grok, that explanation does not apply.
Satisfaction is flat, so it cannot be the tiebreaker
If adoption were tracking quality as users experience it, ratings would spread out. They do not.
Normalized public review ratings across G2, Capterra, Trustpilot, and the app storesAverage user rating, out of 5
Six of the seven products sit between 4.5 and 4.6. That band is narrower than the rounding on most review platforms, and it holds across products with wildly different review volumes: ChatGPT's 4.6 is averaged over 2,204 reviews, Claude's 4.6 over 129, DeepSeek's 4.5 over 8. When a metric returns the same answer for a product with 2,204 data points and one with 8, it is telling you something about the metric as much as about the products.
What it is telling you is that these tools have converged on a satisfaction plateau. Users who choose an AI assistant are, with one exception, happy with the one they chose. The rating is a floor check, not a ranking. It confirms that no product on this list is broken, and then it stops being useful.
The exception is Grok at 4.2. That is the only reading in the set that breaks the plateau, and it is worth taking seriously precisely because the rest of the column is so uniform. In a distribution this tight, a 0.3 to 0.4 point deficit is a real signal rather than sampling noise, though it rests on only 21 reviews and should be read as directional.
Why the two scoreboards diverge
A benchmark measures how well a model answers hard test questions under ideal conditions. Adoption measures everything else: when the product launched, how easy it is to reach, whether it is bundled into tools people already use, how much distribution muscle is behind it, and how good the free tier is. ChatGPT won the market by being first, being a verb, and being everywhere, not by topping a leaderboard.
The structure of our data shows why those two things drift apart rather than converge. Capability is a rapidly depreciating asset. A model that leads the index in August may be third by November, and the lead when it exists is worth one or two points. Distribution compounds in the opposite direction: every month a product spends as the default answer to "which AI should I use" adds reviews, integrations, tutorials, workplace habits, and press, none of which reset when a competitor ships a better model. The 71.9% share is the accumulated interest on being early, not a claim about current model quality.
There is a second mechanism visible in the middle of the pack. Microsoft Copilot and Google Gemini reach essentially the same review volume as Perplexity despite radically different go to market strategies, because bundling into an existing suite and winning a standalone audience converge on similar footprints at this stage. Distribution has more than one shape, and the benchmark is blind to all of them.
This is why chasing benchmark headlines is a poor way to choose a tool. The difference between a 61 and a 59 on an intelligence index is invisible in almost every real task. The difference between a tool your whole team already uses and one they have never opened is enormous. For the vast majority of users, the top four assistants are all more than capable, and the right choice is decided by ecosystem, price, and the specific job, not by two points on a benchmark.
The market map
The chart below plots the seven assistants on two axes derived from our own catalog signals: adoption footprint on the horizontal axis and user satisfaction on the vertical axis. Both are relative positions we computed to make the shape of the market legible. They are not published scores, and no vendor reports them. Horizontal position is derived from the review counts in the breakdown table, compressed so that the tail stays readable next to a product holding 71.9% of the corpus. Vertical position is derived from the 4.2 to 4.6 rating band, stretched so that a spread too small to see on a 0 to 5 axis becomes visible. Read the positions as ordering and grouping, never as measurements.
Relative positions computed from Toolradar catalog signals, not published vendor scoresAI assistant market map, August 2026
Four groupings fall out. ChatGPT occupies the top right corner alone, high on both axes, with no product close enough on the footprint axis to be described as competition. Claude sits at the same satisfaction height and far to its left: identical user verdict, a fraction of the reach. That pairing is the entire report in two dots.
The challenger cluster (Perplexity, Copilot, Gemini) sits mid axis and tightly packed, which is what a genuine three way tie looks like when it is measured rather than asserted. DeepSeek lands far left and at the plateau height, the classic emerging position for a product whose usage is real but happens in places review platforms do not observe. Grok is the only product low on the satisfaction axis, and it sits to the left of the challengers on footprint as well, which is a difficult combination: a product can survive being niche and loved, or broad and merely adequate, but small and comparatively less liked is the hardest quadrant to grow out of.
The full breakdown
| Model | User rating | Public reviews | Share of corpus |
|---|---|---|---|
| ChatGPT | 4.6 | 2,204 | 71.9% |
| Perplexity | 4.5 | 253 | 8.3% |
| Microsoft Copilot | 4.5 | 227 | 7.4% |
| Google Gemini | 4.5 | 222 | 7.2% |
| Claude | 4.6 | 129 | 4.2% |
| Grok | 4.2 | 21 | 0.7% |
| DeepSeek | 4.5 | 8 | 0.3% |
Read the columns against each other and the pattern is consistent: the rating column is nearly constant, the volume column spans three orders of magnitude, and the share column shows a single product defining the category. Claude and ChatGPT are tied on the only quality signal our data contains and separated by 67.7 percentage points of share. Perplexity, Copilot, and Gemini are indistinguishable on both quality and share. The two products at the bottom, Grok and DeepSeek, arrive there for opposite reasons: one has thin usage relative to its coverage, the other has usage that this measurement method structurally cannot see.
What this means
For buyers
Do not pick a model off a leaderboard. The top proprietary models are within a rounding error of each other on capability. Choose on the things that actually differ: which ecosystem you live in, what the free tier gives you, and whether it is strong at your specific task (coding, research, writing, images). The effort setting comparison inside the index makes the point cleanly, since how you run a model can matter as much as which lab built it.
Weight the open-weights option seriously. With Kimi K3 near the closed frontier and DeepSeek close behind, the case for an open or self-hosted model on cost and control is stronger than it has ever been. For high-volume or privacy-sensitive use, the capability gap may no longer justify the price gap. Four index points is the price of admission you are being asked to pay, and for many workloads that is not what the bill reflects.
Separate coverage from adoption. Grok is a live example: heavy press, thin real-world use. If you are evaluating a model because you keep seeing it in headlines, check its actual review corpus first. Attention is not evidence.
Run your own evaluation on your own tasks. Since public ratings cluster at 4.5 to 4.6 and index scores cluster within two points, neither external signal can resolve a choice between the leaders. The only data that will is a short structured trial on the work you actually do, with the same prompts across two or three candidates.
For builders
Distribution beats the leaderboard, and the gap is not two points wide. Nothing in this data suggests a better model wins a market. The product with 71.9% of the review corpus is not the product at the top of the index, and the product at the top of the index sits at 4.2% share with an identical user rating. If you are building on top of models, your defensibility lives in workflow, data, and placement, not in which API you called this quarter.
Model choice is now a swappable decision. A frontier separated by one to two index points, with an open-weights option four points back, is a frontier you should design around rather than bet on. Abstract the provider, benchmark on your own task set, and treat the specific model as a configuration value.
Satisfaction is table stakes, not a differentiator. Six of seven products sit at 4.5 or above. Shipping a tool users rate well no longer distinguishes you, because everyone in this category already does. The variance that remains is in reach.
What we will be watching
The interesting question for the next snapshot is not who leads the index. It is whether the adoption curve responds to capability at all. If Claude's review volume climbs materially while its index lead holds, that is evidence the market does eventually reward the better model with a lag. If the 71.9% share is stable while the top of the leaderboard changes hands again, the honest conclusion is that in this category capability and market position are close to independent variables, and every "new king" headline is describing a race that the market stopped watching.
We will also be watching the open-weights line. Four points is a gap that can close in one release cycle, and if it does, the pricing argument for the closed frontier gets harder to make in exactly the high volume segment where the money is. Nothing in the review corpus will show that shift, because the people making it do not write reviews. That is a measurement problem we would rather name than paper over.
Frequently asked questions
Which AI model is the best in 2026?
It depends what you mean by best. On the Artificial Analysis Intelligence Index, Claude Opus 5 leads on raw capability. On real-world adoption, ChatGPT leads by a wide margin. Because the top models are within a point or two on capability and satisfaction, the best model for you is decided by ecosystem, price, and your specific task, not by the leaderboard.
Is Claude better than ChatGPT?
On the capability benchmark, Claude Opus 5 currently scores highest. On adoption and review volume, ChatGPT is far ahead. Both average 4.6 out of 5 in user ratings, so users are highly satisfied with each. Claude tends to be preferred for writing and coding, ChatGPT for breadth and ecosystem, but the capability gap is small.
Are open-source AI models good enough now?
Closer than ever. The leading open-weights model in this snapshot (Kimi K3) scores within a few points of the top proprietary models, with DeepSeek close behind. For cost-sensitive, high-volume, or privacy-focused use, an open or self-hosted model is now a serious option rather than a compromise.
Why does Grok get so much attention but rank low here?
Grok has heavy press coverage but a small user-review corpus and the lowest rating in this group (4.2). It is a clear example of the gap between attention and adoption: being talked about constantly does not translate into being widely used or highly rated.
Why does ChatGPT have so many more reviews than the others?
Review volume accumulates with time and reach rather than with model quality. ChatGPT launched into the category first, became the default consumer entry point, and has been collecting public reviews for longer and from a wider audience than anyone else. Its 71.9% share of our review corpus reflects that history, and its 22% share of software press coverage in our media analysis points the same way. It is a measure of footprint, not of current capability.
Cite this report
Use the data, credit the source.
Released under Creative Commons BY 4.0. You may quote, link, and reuse the data with attribution.
More research
Two AI Labs Took 75% of 2026's Software Funding
Software companies raised $272B across 449 rounds in 2026, but OpenAI ($110B) and Anthropic ($95B) took 75% of it. The other 414 companies split a quarter.
How Software Is Priced in 2026: The $18 Median and the $178 Mean
We read the pricing pages of 5,194 software tools that publish at least two plans. The median starts at $18 a month, the mean at $178, and that ten-fold gap explains almost everything about how the market prices itself: 60% of tools start under $25, 51% ship a free tier, and 43% publish exactly three plans.
AI for Marketing 2026: AI Writes, It Does Not Yet Send
First-party data on where AI concentrates in the marketing stack. AI is deepest in generative tasks (copywriting 91%, content marketing 63%) and thinnest in operational ones (email marketing 32%): in marketing, AI writes, it does not yet send.