Toolradar Research
AI Assistants Name the Right SaaS Price Only 15% of the Time
We tested GPT-4o, Claude Sonnet 4, and Gemini 2.5 Flash on the current price of 79 well-known SaaS tools. They were right 15% of the time; of the answers that named a price, 74% named the wrong one.
Key findings
What the data shows.
- 01
Correct only 15% of the time. Across 237 questions (79 tools x 3 assistants), the models stated the correct current starting paid price in just 36 answers (15.2%).
- 02
When they commit, they are usually wrong. Of the 139 answers that named a specific price, 74% named the wrong one. The rest (41.4%) declined or hedged without giving a usable current figure.
- 03
An independent numeric check agrees. Ignoring the grader entirely, only 41.8% of answers contained any dollar figure within 12% of the verified price.
- 04
Confidence does not equal accuracy. The assistant that hedged least (Gemini 2.5 Flash, 25.3% hedged) was wrong most (58.2%) the worst failure mode, a wrong price stated with confidence.
- 05
Errors run one way: stale and high. Assistants overwhelmingly quote an older, more expensive tier and miss recent price cuts and new entry plans.
About the research
How we built this report.
Toolradar tool database. Editorial review with weekly pricing verification.
2026. Snapshot taken August 21, 2026.
Public scoring rubric. See how we rate for the full criteria.
Creative Commons BY 4.0. Quote, link, and reuse with attribution.
Ask a leading AI assistant what a piece of software costs and there is a 6-in-7 chance it will not give you the current price. We tested three frontier assistants on the starting paid price of 79 widely used SaaS tools, checked our own answer against each vendor within the last two weeks, and let an independent model grade the results. The assistants named the correct current price 15% of the time. When they did commit to a number, they were wrong far more often than right.
This is the single clearest reason software data belongs in a live tool, not a model's memory: prices move, plans get renamed, cheaper tiers appear, and training data freezes.
Why this matters
An assistant that invents a plausible-but-wrong price is worse than one that says "I do not know," because the user acts on it. A buyer comparing tools, an agent building a shortlist, or a spreadsheet of stack costs inherits the error silently. As AI assistants take on more purchasing research, the cost of stale pricing knowledge compounds.
The fix is not a better model. It is giving the model a live, verified source at answer time.
What we measured
We sampled 79 well-known SaaS products (editorial score 89 or higher) where Toolradar had verified the pricing within the previous ~10 days. For each, we recorded the cheapest paid monthly plan as ground truth. We then asked three frontier assistants, from memory and with no web access or tools, for the current cheapest paid plan and its price. A separate model graded each answer against our verified figure as correct, wrong (a specific price that materially differs), or hedged (declines or gives no current price).
Results by model
| Assistant | Correct | Wrong | Hedged |
|---|---|---|---|
| OpenAI GPT-4o | 12.7% | 39.2% | 48.1% |
| Anthropic Claude Sonnet 4 | 16.5% | 32.9% | 50.6% |
| Google Gemini 2.5 Flash | 16.5% | 58.2% | 25.3% |
| All (237 answers) | 15.2% | 43.5% | 41.4% |
The three assistants disagree on style, not accuracy. Claude Sonnet 4 is the most calibrated: it hedges the most and is wrong the least. Gemini 2.5 Flash is the opposite, committing to a price four times out of five and getting most of them wrong. GPT-4o, with the oldest training cutoff, hedges nearly half the time.
The failure mode: confidently wrong
These are real answers from the test. The verified column is the current cheapest paid plan as of the check date; the assistant column is what the model stated from memory.
| Tool | Assistant said | Verified now | Model |
|---|---|---|---|
| Tableau | $75 | $15 | Gemini 2.5 Flash |
| Zendesk | $69 | $19 | Gemini 2.5 Flash |
| HubSpot | $50 | $15 | GPT-4o |
| Shopify | $39 | $29 | GPT-4o |
| Docker | $5 | $9 | GPT-4o |
| DocuSign | $15 | $10 | Gemini 2.5 Flash |
| Postman | $15 | $9 | Gemini 2.5 Flash |
| Refersion | $99 | $39 | GPT-4o |
| Grafana | $49 | $19 | GPT-4o |
| Odoo | $6 | $31 | Claude Sonnet 4 |
The direction is consistent: the model quotes a higher, older number. It learned a past price list and cannot see that the vendor added a cheaper tier or cut the entry price. No amount of reasoning fixes a stale fact.
What fixes it
Toolradar verifies pricing continuously and exposes it to AI systems through the Toolradar MCP server. Instead of recalling a price from training data, an assistant calls get_pricing or recommend_tools and receives the figure we verified this week, with the plan name and the date we last checked it. On the 79 tools in this test, that turns a 15% hit rate into the verified number, every time, with a citation.
The same server also lets an agent flag a discrepancy back to us with report_issue, so the data gets sharper the more it is used.
Methodology
Sample. 79 published SaaS products with a Toolradar editorial score of 89 or higher and a clean, monthly, per-seat starting paid price between $3 and $100, each with pricing verified within roughly 10 days of the test (verification dates span August 11-21, 2026). We deliberately restricted the set to well-known tools and simple monthly pricing so the question is fair and the ground truth is unambiguous; usage-based, contact-sales, one-time, and enterprise-only products were excluded.
Ground truth. The cheapest paid monthly plan recorded in Toolradar's database, each field verified against the vendor's own pricing page within the check window.
Baseline. OpenAI GPT-4o, Anthropic Claude Sonnet 4, and Google Gemini 2.5 Flash, queried through OpenRouter with no tools and no web access, at temperature 0, so answers reflect parametric (training) knowledge only. Reasoning-tier models were excluded because they returned empty or truncated content over this API path.
Grading. A separate model (GPT-4o mini) classified each answer against the verified figure as correct, wrong, or hedged, using a fixed rubric. As an independent cross-check we also ran a rule-based numeric match (any dollar figure within 12% of the verified price); it agreed with the grader on the overall order of magnitude (41.8% loose match vs 15.2% strict-correct).
Limitations. Some "wrong" verdicts reflect a billing-basis gap (a monthly figure vs our annual-equivalent) or a plan-scope gap (the model quotes a mid-tier when a cheaper tier exists); we count these as misses because a user would be misled either way. Prices change constantly, so this snapshot reflects mid-August 2026. Grading with a single model introduces some noise; we spot-checked a sample of verdicts by hand and found them consistent.
Cite this report
Use the data, credit the source.
Released under Creative Commons BY 4.0. You may quote, link, and reuse the data with attribution.
More research
One Company Now Owns the Review Record for 96% of B2B Software
After G2 acquired Capterra, Software Advice and GetApp from Gartner, a single owner covers 6,232 of the 6,499 reviewed products in our catalog. Measured across 4.3 million aggregated reviews.
Two AI Labs Took 75% of 2026's Software Funding
Software companies raised $272B across 449 rounds in 2026, but OpenAI ($110B) and Anthropic ($95B) took 75% of it. The other 414 companies split a quarter.
How Software Is Priced in 2026: The $18 Median and the $178 Mean
We read the pricing pages of 5,194 software tools that publish at least two plans. The median starts at $18 a month, the mean at $178, and that ten-fold gap explains almost everything about how the market prices itself: 60% of tools start under $25, 51% ship a free tier, and 43% publish exactly three plans.
