Best LLM Monitoring Tools in 2026
Short answer: Langfuse suits cost and error alerts without a per-seat fee. Core is $29/mo, with 100,000 units a month and 20 alerts, so the shared chart has no seat line. Datadog suits a page that already fires there: the first 100,000 LLM spans are $160/mo on an annual bill, and 40,000 spans are free. LangSmith suits a LangChain app, at $39 per seat on Plus, because the chart outlives a 14-day trace. Helicone suits a base-URL change, at $79/mo on Pro. Replaying one bad prompt is a different purchase.
Page on cost, latency, and errors, and keep the chart after the prompt is gone.
Buy the page, not the replay, because a trace viewer that cannot fire an alert is a debugger you open after the incident. An LLM monitoring tool should tell you that cost, latency, errors, or a quality score moved. It should still say that after the prompt text has been deleted.
Toolradar data: of the 43 AI observability tools in the catalog, 58% offer a free or freemium plan, while 17 (40%) are paid-only.
That split matters because almost every vendor below will chart a prototype for free, then bill production by the unit you forgot to count. A Langfuse unit, a Datadog span, a LangSmith trace, a Helicone request, or an Opik span will each show up as a different line. A free plan that stops recording, while the app keeps answering, is a monitoring outage with a green status page. If the job is to open one failed request and score it, use the LLM observability guide. This page is the threshold, the invoice, and what the dashboard still knows next month.
Start with Langfuse when several people need the chart and you refuse a seat fee, because Core is the entry paid plan and alerts are a numbered allowance. Move to Datadog when the on-call rotation already lives there, because Agent Observability is its own span price and not a host. Use LangSmith when the agent is LangChain or LangGraph and you need the chart to survive the short trace. Use Helicone when you can change the base URL and cannot ship an SDK this week.
How we ranked: these ten were picked from the 43 AI observability tools in the catalog for production monitoring of LLM traffic (alerts, cost, latency, and what remains after retention), every price was read on the vendor's own page in September 2026, and nobody paid for a slot. WhyLabs' own site says the company is discontinuing operations, so it is not a pick.
Top Picks
Picked by editorial review, informed by G2 and Capterra review volume and rating and by media mentions, the signals behind our category rankings. How we rate
| Tool | Starting price | Rating | Best for |
|---|---|---|---|
| Langfuse | From $29/mo | n/a | Teams that want cost and error alerts without paying per seat |
| Datadog | From $160/mo annual | 4.4726 reviews | On-call teams that already page from Datadog and want an LLM span price |
| LangSmith | From $39/seat | n/a | LangChain and LangGraph apps that need a chart after the prompt expires |
| Helicone | From $79/mo | n/a | Apps that can change a base URL and need request cost and latency alerts |
| Arize AX | From $50/mo | 4.348 reviews | Teams that want hosted monitors and a free local Phoenix install beside them |
| Portkey | From $49/mo | 4.618 reviews | Gateway users who need alerts, and metrics that outlive the request log |
| Opik | From $19/mo | n/a | Production agents that should webhook when a trace errors inside a span budget |
| Galileo | $100/mo billed yearly | n/a | Teams that want eval scores turned into a production guardrail |
| Braintrust | From $249/mo | 4.354 reviews | Teams whose monitor is a quality score chart, not an infrastructure page |
| Fiddler | $0.002 per trace | n/a | Security teams that want a per-trace meter and a fast free guardrail |
Teams that want cost and error alerts without paying per seat
Langfuse's pricing is quite generous, especially with a robust free Hobby tier offering 50k observations.
Watch out
Overage fees not explicitly stated for Pro tier
On-call teams that already page from Datadog and want an LLM span price
Datadog is the most comprehensive observability platform, infrastructure monitoring, APM, logs, RUM, security, and 750+ integrations in one place.
Watch out
High-water mark billing: Datadog measures host count hourly and bills based on the highest usage (minus top 1%). Auto-scaling environments that spike to 200 hosts for 2 hours but normally run 50 will be billed at ~200 hosts for the entire month
LangChain and LangGraph apps that need a chart after the prompt expires
Plus is $39/seat/month with 10,000 traces and one free small serverless deploy.
Watch out
Traces over the included 5k/10k are pay-as-you-go. Seats are only the floor.
Apps that can change a base URL and need request cost and latency alerts
It's best for developers and small to medium-sized businesses building and scaling AI applications.
Watch out
Potential overage fees if exceeding Pro tier limits
Teams that want hosted monitors and a free local Phoenix install beside them
Arize AI offers a generous free tier (AX Free) and an open-source option (Phoenix), making it highly accessible for individual developers and small teams.
Watch out
Self-hosting is an Enterprise add-on
Gateway users who need alerts, and metrics that outlive the request log
Portkey's pricing is fair, offering a generous free Developer tier with 10k requests/month.
Production agents that should webhook when a trace errors inside a span budget
Opik by Comet's pricing is extremely generous, offering a fully-featured Free tier that includes advanced capabilities like automated agent optimization and built-in guardrails.
Watch out
Potential overage fees for high usage
Teams that want eval scores turned into a production guardrail
Teams whose monitor is a quality score chart, not an infrastructure page
Braintrust's pricing is quite generous, offering a robust Free tier with unlimited evals and 3 users, which is excellent for individual developers or small teams.
Watch out
No clear overage fees mentioned for usage
Security teams that want a per-trace meter and a fast free guardrail
The pricing for Fiddler AI is generally fair, with a generous free tier for basic guardrails.
What LLM monitoring actually is
LLM monitoring is software that charts cost, latency, errors, and quality on live model traffic, then pages someone when a threshold breaks. If it cannot fire without a human opening a trace, it is a log, not a monitor.
The chart is the product, and the prompt text is evidence you may not be allowed to keep. Langfuse sells custom dashboards on every cloud plan and then meters alerts: 2 on Hobby, 20 on Core, 50 on Pro, 100 on Enterprise, so the page is capped by plan. Datadog prices Agent Observability by LLM span and, in its SDK docs, keeps token and cost metrics on all instrumented traffic even when you sample traces. A cheaper sample does not hide the spend chart. LangSmith's monitoring tab keeps working after a base trace's 14-day retention ends, because it reads metadata kept for more than 30 days. Helicone's alert docs cover error rate, cost, latency, token counts, and request count, with email or Slack.
A gateway monitor sees the request you were willing to proxy. Portkey and Helicone sit on that path, which is fast to turn on and blind to a tool call that never left your process. Skip the proxy when that failure lives inside the agent loop. An SDK monitor sees the steps you instrumented, which is what you want for an agent and what you will not get from a base-URL change.
Quality monitors are a third shape, for a team that pages on a wrong answer rather than a slow one. Galileo wants eval scores to become guardrails, Braintrust wants custom charts of scores, and Fiddler sells a per-trace meter next to a free guardrail. None of those is a substitute for a latency page if the score never runs.
The wrong adjacent purchase is a generic ingest bill with no LLM span. New Relic prices data at $0.40/GB after 100 GB a month, which can hold model telemetry and still does not name an LLM span. The head-to-head on the rest of the stack is Datadog vs New Relic, and prompt-injection blocking belongs in LLM security.
Why the chart outlives the prompt
The expensive mistake is paying for 14 days of prompt text and assuming the page will still explain last month's bill. LangSmith base traces are kept 14 days, and extended traces are 180 days for new traces after the September 14, 2026 change in LangSmith's billing docs. The monitoring tab is the part that stays when the review is a chart rather than the prompt. If your on-call playbook says "open the prompt," you needed the longer tier, and that upgrade is a second charge on the same trace.
Langfuse makes the same split with a clearer clock. Hobby keeps 30 days of data and 2 alerts. Core keeps 90 days and 20 alerts, enough to cover a monthly review. Pro keeps 3 years, 50 alerts, and the SOC 2 and ISO 27001 reports, at $199/mo, with the same 100,000 included units as Core. You are paying for history and the page, not for more included traffic. Enterprise is $2,499/mo and raises the alert cap to 100.
Portkey splits the clock inside one plan. Developer keeps logs for 3 days and metrics for 30. Production keeps logs for 30 days and metrics for 90. The chart answers what spiked, and the log answers what the model said. Buying the short log and expecting a postmortem is how teams discover the gap during the incident. Helicone's Hobby plan keeps 7 days and Pro keeps 1 month, so a review the following week needs the paid retention. Team, at $799/mo, keeps 3 months and is the first tier that lists SOC 2 and HIPAA.
Datadog's default on the first span block is 15-day trace retention. Longer trace retention is a higher rate per 10,000 spans, up to $7.50 per 10,000 on an annual bill for 90-day traces. Sample the traces if you must, and read the SDK note first. Token and cost metrics stay on 100% of instrumented traffic, while the bill follows the spans you send, and the default sample rate is 1.0. A sample that was meant to save money does not shrink the cost chart.
Key Features to Look For
An alert with a number on the plan (Essential)
Langfuse prints the alert cap on each tier, and Helicone's docs alert on error rate, cost, latency, tokens, or request count. A dashboard you have to remember to open is not the monitor.
A billing unit you can count (Essential)
Langfuse units are traces plus observations plus scores. Datadog bills LLM spans. Opik bills spans at $5 per extra 100,000 on Pro. If you cannot count the unit, you cannot forecast the month.
Cost in dollars, not only tokens (Essential)
Langfuse tracks token and cost on every cloud plan. Datadog shows estimated cost in US dollars when the provider is OpenAI, Azure OpenAI, or Anthropic. A token chart without a dollar still surprises finance.
A path that matches where the call happens (Important)
Helicone and Portkey monitor what passes the proxy. Langfuse, LangSmith, Arize, and Opik monitor what the SDK or OpenTelemetry emits. A base-URL change will not see a local tool call.
A quality signal you can page on (Important)
Galileo wants evals distilled into guardrails. Braintrust puts custom charts on Pro. Fiddler's free guardrail is aimed at harmful output. A latency page will not catch a confident wrong answer.
A place the prompts are allowed to sit (Important)
Langfuse Cloud offers US, EU, or Japan, with a HIPAA region on Enterprise. Arize AX lists US, EU, or CA. Fiddler Enterprise adds VPC or on-prem, and that tier publishes no list price.
A spend ceiling you can set yourself (Nice to have)
Langfuse's pricing FAQ says you can email yourself when spend crosses a threshold, and LangSmith lets a workspace set a spend limit that becomes a trace cap. Without that, the monitor's own bill is the next incident.
What to settle before the demo
Write down whether the missing signal is a proxy request, an in-process agent step, or a quality score. Helicone and Portkey start at the URL. Langfuse, LangSmith, Arize, and Opik start at instrumentation. Galileo, Braintrust, and Fiddler start at a score or a guardrail.
Decide how long the prompt text has to exist before you compare stickers. Fourteen days on a LangSmith base trace, seven on Helicone Hobby, and three on Portkey Developer are all real clocks. If legal wants the text for a quarter, you are on a higher tier before the logos matter.
Ask whether the chart survives sampling and deletion, because a vendor that deletes both is a log with a timer. Datadog's cost metrics stay on full traffic when traces are sampled, and LangSmith's monitoring tab stays after the base trace is gone.
Separate the LLM line from the rest of the observability contract, or a host quote will look like this tool's price. Datadog Infrastructure Pro is $15 per host per month on an annual bill, and Agent Observability is a different row.
Evaluation Checklist
On Langfuse, count units as traces plus observations plus scores for one busy day, then see whether you still fit Hobby's 50,000 units and 2 alerts or you need Core's 20 alerts.
On Datadog, confirm the order is Agent Observability spans, note that 40,000 spans are marked free, and check whether you are on the annual, month-to-month, or on-demand column.
On LangSmith, set base retention and confirm the monitoring tab still has last month's chart, then see whether online evaluators are upgrading traces to 180 days.
On Helicone, send enough traffic to pass 10 logs per minute on Hobby, and confirm whether alerts are included or an add-on before you treat Pro as the paging plan.
On Portkey, exceed the Developer log cap on purpose in a test project and confirm requests still succeed while new logs stop, which is the failure mode of that free tier.
On Opik, decide whether you are on the open-source install, Free Cloud's 25,000 spans, or Pro's 100,000 spans, because those three are not the same allowance.
On Fiddler, price a week of traces at the published per-trace rate before you assume the free guardrail is the production monitor.
Pricing Overview
Alert platforms with a seat-free fee
Monthly plan plus a usage unit
Per span block, per seat, or per trace
Yearly cards and score charts
Galileo Pro on a yearly card, Braintrust Pro, and Helicone Team when you need the compliance line.
Annual display price, or a platform fee
Pricing Comparison
| Tool | Published price | What that price buys | Billing |
|---|---|---|---|
Langfuse | $29/mo Core | 100,000 units, 90-day access, 20 alerts. Hobby is free. | Monthly, plus units |
Datadog | $160/mo annual | First 100,000 LLM spans with 15-day retention, and 40,000 spans are free. | Annual span block |
$39/seat/mo | Plus includes 10,000 base traces, then pay as you go, and one Developer seat has no seat fee. | Per seat, plus usage | |
Helicone | $79/mo Pro | Unlimited seats, 10,000 requests included, 1-month retention. | Monthly, then usage |
$50/mo Pro | 50,000 spans, 10 GB, 30-day retention, 25 Signal issues. | Monthly | |
Portkey | $49/mo Production | 100,000 logs, alerts, 30-day logs and 90-day metrics. | Monthly, plus logs |
Opik | $19/mo Pro | 100,000 spans, 60-day retention, up to 50 members. | Monthly, plus spans |
Galileo | $100/mo yearly | Pro lists 50,000 traces, and the month-to-month rate is not printed. | Billed yearly |
Braintrust | $249/mo Pro | 5 GB included, 50,000 scores, custom charts, 30-day retention. | Monthly, plus usage |
Fiddler | $0.002 per trace | Developer observability on SaaS, with RBAC and SSO, and the free guardrail is separate. | Per trace |
Prices were read on September 24, 2026 from each vendor's pricing page. Datadog's annual column is the first span block in this table, and Galileo's Pro card is the yearly price. LangSmith does not print a per-trace dollar, only seat fees and a storage unit. See AI observability for the wider catalog, and LLM observability when the job is a trace replay.
Mistakes to Avoid
- ×
Buying Langfuse Pro because it sounds like more traffic does not raise the included units, which stay at 100,000. The jump pays for 3-year access, 50 alerts, higher ingest, and the compliance reports. If you only needed 20 alerts and 90 days, Core was the plan.
- ×
Treating a Datadog host price as the LLM monitor prices the wrong row. Infrastructure Pro is a per-host row, and Agent Observability is the span row, with its own annual, month-to-month, and on-demand columns. A quote that never says spans did not price this job.
- ×
Leaving LangSmith evaluators on the retention upgrade moves a trace from 14 days to 180 days. The monitoring chart did not need that text, and the invoice did, so turn the upgrade off if the page is all you open.
- ×
Shipping Helicone Hobby into production will miss the burst and forget it before the review. Ten logs a minute and 7 days of retention are the reason, and Pro is the first plan with a month of retention and a much higher ingest cap.
- ×
Assuming Portkey Developer keeps a production log fails past the recorded-log cap. Past 10,000 recorded logs, requests continue and new logs do not. The app looks healthy while the monitor goes dark.
- ×
Reading Fiddler's free guardrail as the monitoring price skips the meter. The guardrail card has no trace meter, so Developer is the per-trace line, and Enterprise is a quote for VPC or on-prem.
Expert Tips
- →
Count Langfuse units on one real trace before you pick a tier. Add the trace, every observation, and every score, because a single agent turn with a retrieval step and a judge is several units, which is why the 50,000-unit Hobby plan disappears on a small production app.
- →
Set the LangSmith spend limit before you enable online scores. The limit becomes a trace cap, and the score can extend retention. If you only needed the chart, keep the trace on the 14-day tier and let the monitoring tab hold the history.
- →
Sample Datadog traces only after you accept that the cost chart stays whole. The SDK keeps token and cost metrics on all instrumented traffic. The savings, if any, are the spans you no longer store, at the per-10,000 rate in the table.
- →
Put Helicone alerts through checkout, not only the docs, because the docs describe cost, latency, and error alerts while the upgrade call still lists alerts as an add-on flag. Confirm the invoice line before you promise the on-call a Slack channel.
- →
Split the gateway from the agent step. Portkey and Helicone see the proxied model call. An in-process tool call needs an SDK, which is the job described in AI agent observability. The platform layer around both is observability platforms.
- →
Ask Galileo for the month-to-month number in writing. The public Pro card is the yearly price and says the rate scales with traces. If procurement cannot sign a year, that card is not a budget.
Red Flags to Watch For
- !
A Langfuse order that prices Pro for the compliance reports and then assumes the included units grew with the fee. They did not.
- !
A Datadog quote that sells Infrastructure hosts and describes LLM spans as if they were included in the host price.
- !
A LangSmith workspace with no spend limit, and online evaluators left on the default that extends retention, so the chart quietly buys the long tier.
- !
A Helicone Hobby rollout on production traffic: 7-day retention and 10 logs per minute will drop the incident you meant to page on.
- !
A Portkey Developer plan treated as production logging. Past 10,000 recorded logs, the pricing page says requests keep working and further logs are not recorded.
- !
The Bottom Line
Langfuse when the monitor is a shared chart and a numbered alert cap, and you will not pay per person to look at it. Skip it when you need SSO tomorrow, because that is the Teams add-on on top of Pro, or when you need the prompt text for years and will not pay for that retention.
Datadog when the page has to land in the same place as the rest of the on-call, and you will buy LLM spans as their own line. Skip it when you wanted a seat-free LLM tool and do not already run Datadog.
LangSmith when the stack is LangChain or LangGraph and the chart has to survive a 14-day trace. Helicone when a base URL is the integration you can actually ship, after you confirm alerts are on the invoice. Arize AX when you want hosted monitors at the Pro span cap and a free local Phoenix install for the notebook.
Portkey when the gateway should alert and the metrics must outlive the log, and you have read the line that renames the enterprise product. Opik when the budget is spans and a webhook on a trace error. Galileo when the monitor is a guardrail distilled from an eval, and you can live with a yearly card. Braintrust when the chart is a score and the Pro fee is acceptable before overages. Fiddler when you want a published per-trace rate and a free guardrail with a stated latency, and you will not confuse those two cards.
Cite this: Toolradar, "Best LLM Monitoring Tools in 2026", September 2026. Prices checked on vendor pages on September 24, 2026. No paid placement. Compared with the 43 AI observability tools we track.
Frequently Asked Questions
What is the best LLM monitoring tool in 2026?
Langfuse, if you want cost and error alerts without a per-seat fee. Core includes 100,000 units, 90 days of access, unlimited users, and 20 alerts, at the monthly price in the comparison table. Hobby is free and stops at 2 alerts and 50,000 units. Choose Datadog if the page has to land in an existing on-call tool, LangSmith if the app is LangChain, and Helicone if you can only change a base URL.
How much does an LLM monitoring tool cost in 2026?
As of September 24, 2026, Langfuse Core is $29/mo and Pro is $199/mo, then $8 per extra 100,000 units at the first overage step. Datadog's first 100,000 LLM spans are $160/mo on an annual bill, $200 month-to-month, or $240 on demand. LangSmith Plus is $39 per seat with 10,000 base traces included. Helicone Pro is $79/mo, Arize AX Pro is $50/mo, and Portkey Production is $49/mo. Opik Pro is $19/mo, Galileo Pro is $100/mo when billed yearly, Braintrust Pro is $249/mo, and Fiddler Developer is $0.002 per trace. Enterprise tiers on these products publish no list price.
Is there a free LLM monitoring tool in 2026?
Yes, and every free option below stops on a hard cap. Langfuse Hobby includes 50,000 units, 30 days, and 2 alerts. Datadog's price list marks 40,000 LLM spans free. LangSmith Developer includes 5,000 base traces on one seat until you add a card. Helicone Hobby includes 10,000 requests and 7 days. Arize AX Free includes 25,000 spans. Portkey Developer includes 10,000 logs and then stops recording. Opik's open-source build is free, and Free Cloud includes 25,000 spans. Galileo Free includes 5,000 traces. Braintrust Starter includes 1 GB and 10,000 scores. Fiddler's free plan is the guardrail, not the per-trace monitor.
How does Langfuse compare with Datadog for LLM monitoring?
Langfuse is the seat-free LLM chart: Core is the first paid row in the table, alerts are a printed cap, and a unit is a trace plus observations plus scores. You can self-host the MIT build and pay no unit fee, which suits a team that will run the upgrades itself. Datadog is the span line on an existing on-call stack: the annual block in the table covers the first 100,000 spans, and token and cost metrics stay on full traffic when you sample traces. Choose Langfuse when the LLM app is the thing you monitor. Choose Datadog when the page has to sit next to the rest of the infrastructure alerts, and keep host pricing off that comparison.
What is the difference between LLM monitoring and LLM observability?
Monitoring is the chart and the page: cost, latency, errors, and a threshold that notifies someone. Observability is the replay: the prompt, the tool calls, and the score on one request. LangSmith's monitoring tab still works after a 14-day base trace is deleted, which is the split in one product. If you need the replay, use the LLM observability guide. If you need the page, stay on this list.
Does the monitoring chart survive after the prompt is deleted?
On LangSmith, yes for the monitoring tab: base traces last 14 days, and the tab reads metadata kept for more than 30 days. On Portkey Production, metrics last 90 days and logs last 30. On Langfuse, the data-access clock is the chart clock: 30 days on Hobby, 90 on Core, and 3 years on Pro. If the review needs the prompt text itself, buy the longer retention. The chart is not a copy of the text.
Which LLM monitoring tool fits a LangChain app?
LangSmith Plus, at the per-seat price in the table, includes 10,000 base traces. The monitoring chart remains after the 14-day base trace expires, which is the production question. Langfuse is the better fit when the same app should not pay per seat and you will add the SDK yourself. Helicone fits when you will only change the base URL and can accept that in-process tool calls may never hit the proxy.
Cite this page: Toolradar, "Best LLM Monitoring Tools in 2026", updated September 2026, https://toolradar.com/guides/best-llm-monitoring-tools
Sources
Prices and plan details on this page come from each vendor's own pricing page, re-checked by the Toolradar pricing tracker:
- Langfuse pricing, checked
- Datadog pricing, checked
- LangSmith pricing, checked
- Helicone pricing, checked
- Arize AX pricing, checked
- Portkey pricing, checked
- Opik pricing, checked
- Galileo pricing
- Braintrust pricing, checked
- Fiddler pricing, checked
