Skip to content

Best LLM Monitoring Tools in 2026

TL;DR

Short answer: Langfuse suits cost and error alerts without a per-seat fee. Core is $29/mo, with 100,000 units a month and 20 alerts, so the shared chart has no seat line. Datadog suits a page that already fires there: the first 100,000 LLM spans are $160/mo on an annual bill, and 40,000 spans are free. LangSmith suits a LangChain app, at $39 per seat on Plus, because the chart outlives a 14-day trace. Helicone suits a base-URL change, at $79/mo on Pro. Replaying one bad prompt is a different purchase.

Page on cost, latency, and errors, and keep the chart after the prompt is gone.

As featured in
  • TechCrunch
  • Forbes
  • Bloomberg
  • Business Insider
  • The Verge
43 AI Observability tools tracked

Buy the page, not the replay, because a trace viewer that cannot fire an alert is a debugger you open after the incident. An LLM monitoring tool should tell you that cost, latency, errors, or a quality score moved. It should still say that after the prompt text has been deleted.

Toolradar data: of the 43 AI observability tools in the catalog, 58% offer a free or freemium plan, while 17 (40%) are paid-only.

That split matters because almost every vendor below will chart a prototype for free, then bill production by the unit you forgot to count. A Langfuse unit, a Datadog span, a LangSmith trace, a Helicone request, or an Opik span will each show up as a different line. A free plan that stops recording, while the app keeps answering, is a monitoring outage with a green status page. If the job is to open one failed request and score it, use the LLM observability guide. This page is the threshold, the invoice, and what the dashboard still knows next month.

Start with Langfuse when several people need the chart and you refuse a seat fee, because Core is the entry paid plan and alerts are a numbered allowance. Move to Datadog when the on-call rotation already lives there, because Agent Observability is its own span price and not a host. Use LangSmith when the agent is LangChain or LangGraph and you need the chart to survive the short trace. Use Helicone when you can change the base URL and cannot ship an SDK this week.

How we ranked: these ten were picked from the 43 AI observability tools in the catalog for production monitoring of LLM traffic (alerts, cost, latency, and what remains after retention), every price was read on the vendor's own page in September 2026, and nobody paid for a slot. WhyLabs' own site says the company is discontinuing operations, so it is not a pick.

Top Picks

Picked by editorial review, informed by G2 and Capterra review volume and rating and by media mentions, the signals behind our category rankings. How we rate

Best LLM Monitoring Tools in 2026 compared: starting price, rating and best use, as of September 2026
ToolStarting priceRatingBest for
LangfuseFrom $29/mon/aTeams that want cost and error alerts without paying per seat
DatadogFrom $160/mo annual4.4726 reviewsOn-call teams that already page from Datadog and want an LLM span price
LangSmithFrom $39/seatn/aLangChain and LangGraph apps that need a chart after the prompt expires
HeliconeFrom $79/mon/aApps that can change a base URL and need request cost and latency alerts
Arize AXFrom $50/mo4.348 reviewsTeams that want hosted monitors and a free local Phoenix install beside them
PortkeyFrom $49/mo4.618 reviewsGateway users who need alerts, and metrics that outlive the request log
OpikFrom $19/mon/aProduction agents that should webhook when a trace errors inside a span budget
Galileo$100/mo billed yearlyn/aTeams that want eval scores turned into a production guardrail
BraintrustFrom $249/mo4.354 reviewsTeams whose monitor is a quality score chart, not an infrastructure page
Fiddler$0.002 per tracen/aSecurity teams that want a per-trace meter and a fast free guardrail

Teams that want cost and error alerts without paying per seat

+Core includes 100,000 units, 90 days of data access, unlimited users, and 20 alerts, at the monthly price in the table, so the team shares one chart. Hobby is free for 50,000 units, 30 days, 2 users, and 2 alerts, with no card.
+A unit is a trace, an observation, or a score, and token and cost tracking is on every cloud plan. Their own example puts 200,000 units on Core at $37 for the month: the plan fee plus $8 for the second 100,000.
+Custom dashboards are on every cloud tier, and the pricing FAQ says you can set spend alerts by email. Past the included units, paid plans step from $8 per 100,000 to $7, then $6.50, then $6.
−Pro is $199/mo for the same 100,000 included units. You buy 3-year access, 50 alerts, 20,000 requests per minute, and SOC 2, ISO 27001, and a HIPAA BAA, not more traffic. SSO is a Teams add-on at $300/mo, not part of Pro.
−Hobby does not list an overage rate, so the free chart stops being a plan you can burst. Self-hosting the MIT build avoids the unit meter and leaves you the upgrades, backups, and storage.
Good value

Langfuse's pricing is quite generous, especially with a robust free Hobby tier offering 50k observations.

Watch out

Overage fees not explicitly stated for Pro tier

2
Datadog logo

Datadog

  • 4.4 on G2 (726 reviews)

On-call teams that already page from Datadog and want an LLM span price

+The annual price in the table covers the first 100,000 LLM spans, with 15-day retention by default. Month-to-month is $200 and on-demand is $240, and the same list marks 40,000 LLM spans free.
+Extra 15-day spans are $3.50 per 10,000 on the annual column. The SDK docs say sampling does not shrink token and cost metrics: those stay on all instrumented traffic, and the default sample rate is 1.0.
+Estimated cost displays in US dollars when the model provider is OpenAI, Azure OpenAI, or Anthropic. That is the dollar a finance partner can read without a token converter.
−Ninety-day trace retention is $7.50 per 10,000 spans on an annual bill, so the long postmortem is a different rate from the default 15-day block. On-demand experiment retention on the 15-day tier falls to 15 days unless you are on an annual or month-to-month commitment.
−Infrastructure Pro is $15 per host per month billed annually, and it is not the LLM line. A quote that never says Agent Observability sold you hosts. The wider host comparison is Datadog alternatives.
Fair value

Datadog is the most comprehensive observability platform, infrastructure monitoring, APM, logs, RUM, security, and 750+ integrations in one place.

Watch out

High-water mark billing: Datadog measures host count hourly and bills based on the highest usage (minus top 1%). Auto-scaling environments that spike to 200 hosts for 2 hours but normally run 50 will be billed at ~200 hosts for the entire month

LangChain and LangGraph apps that need a chart after the prompt expires

LangSmith screenshot
+Plus is the per-seat price in the table, with 10,000 base traces included and more seats at the same seat price, so the bill grows with logins. Developer is one seat with 5,000 base traces a month until a card is on file.
+Base traces last 14 days and extended traces last 180 days for new traces, per the billing docs' September 14, 2026 change. The monitoring tab keeps working after a base trace expires, on metadata kept for more than 30 days.
+A workspace spend limit turns into a trace cap. The pricing page prices storage at $1.00 per LSU and compute at $1.50 per LCU. Traces consume LSUs, and the page does not print how many traces sit in one LSU.
−Without a card, Developer stops at 5,000 traces a month. With a card, excess traces are pay as you go, and the pricing page still does not print a per-trace dollar. Online evaluators can upgrade a trace to extended retention unless you opt out, which raises the bill.
−Plus allows 500,000 trace events and 5 GB of trace data per hour. A single trace stops at 25,000 runs. Enterprise, including hybrid and self-hosted, publishes no list price.
Good value

Plus is $39/seat/month with 10,000 traces and one free small serverless deploy.

Watch out

Traces over the included 5k/10k are pay-as-you-go. Seats are only the floor.

4
Helicone logo

Helicone

  • 4.5 on G2 (2 reviews)

Apps that can change a base URL and need request cost and latency alerts

+Pro, at the monthly price in the table, includes unlimited seats, a 7-day trial, 10,000 requests, 1-month retention, and 1,000 logs per minute. Hobby is 10,000 requests, 1 GB, 1 seat, 7-day retention, and 10 logs per minute.
+The alert docs let you threshold error rate, cost, latency, token counts, or request count, and notify email or Slack. Grouping can split the same alert by user, model, or provider.
+On the pricing page, the calculator's sample of 10,000 requests and 0.30 GB showed $0.00 for the requests and $0.97 for the storage. The request allotment was not the line that cost money in that sample.
−The Pro upgrade API still has an alerts flag under add-ons, so do not treat the Pro card as proof that paging is included until the checkout says so. Team is $799/mo for 5 organizations, SOC 2 and HIPAA, 3-month retention, and 15,000 logs per minute.
−Hobby at 10 logs per minute will not absorb a production burst, and 7-day retention will not support a postmortem the following week. Enterprise, with forever retention, publishes no list price. Gateway routing is covered in LLM gateways.
Good value

It's best for developers and small to medium-sized businesses building and scaling AI applications.

Watch out

Potential overage fees if exceeding Pro tier limits

5
Arize AX logo

Arize AX

  • 4.3 on G2 (48 reviews)

Teams that want hosted monitors and a free local Phoenix install beside them

Arize AX screenshot
+AX Pro is $50/mo for 50,000 spans, 10 GB, 30-day retention, unlimited users, unlimited evals, and 25 Signal issues a month. AX Free is 25,000 spans, 1 GB, 15 days, and 10 Signal issues, with no card.
+The pricing comparison lists custom monitors plus token, latency, and cost tracking. Data Fabric is marked off Free and Pro, so that row is an Enterprise conversation.
+Phoenix, the open-source local install, is the path that stays on your machine, and the pricing page tells you to move to AX when you need the hosted product. The directory page for that install is Phoenix.
−Pro's retention is 30 days and its span cap is 50,000, so a busy agent is an Enterprise quote rather than a printed overage. Free and Pro are SaaS only. Self-hosted deployment, SSO, and HIPAA are Enterprise, which publishes no list price.
−Free is one space, and Pro is one organization with two spaces. A second production environment past that is not a toggle on the Pro plan.
Good value

Arize AI offers a generous free tier (AX Free) and an open-source option (Phoenix), making it highly accessible for individual developers and small teams.

Watch out

Self-hosting is an Enterprise add-on

6
Portkey logo

Portkey

  • 4.6 on G2 (18 reviews)

Gateway users who need alerts, and metrics that outlive the request log

+Production is $49/mo for 100,000 recorded logs, alerts, 30-day logs, and 90-day metrics, so the chart outlives the text. Extra logs are $9 per 100,000, and the comparison table applies that published rate up to 3 million requests.
+Developer is free for 10,000 recorded logs, with 3-day logs and 30-day metrics. The pricing page says exceeding that cap does not affect requests: only the extra logs are dropped.
+The open-source column lists a local gateway with retries, fallbacks, and a basic dashboard, which is the install when prompts cannot sit in their cloud. Virtual keys on the paid plans include budgeting.
−The pricing page says Portkey is now Prisma AIRS AI Gateway for enterprises. The Production plan is still listed, and the enterprise deployment (private cloud, HIPAA, SSO) publishes no list price.
−Developer will keep serving traffic after it stops recording, which is a silent monitoring failure. Alerts are listed on Production, not on the free tier's feature list.
Good value

Portkey's pricing is fair, offering a generous free Developer tier with 10k requests/month.

Production agents that should webhook when a trace errors inside a span budget

Opik screenshot
+Pro Cloud is $19/mo for up to 50 members, 100,000 spans, and 60-day retention. Additional spans are $5 per 100,000. Free Cloud is 25,000 spans, 60 days, and up to 10 members.
+The reliability table lists production-scale monitoring, online evaluation, and alerts as webhook notifications for trace errors, new feedback scores, or prompt changes. Token and cost tracking is on the observability grid.
+The open-source column is $0 with unlimited spans in that table, and Comet says the OSS build carries the same core feature set as the hosted product. Guardrails in that grid are marked for self-hosted, not for the cloud columns.
−Stretching retention from 60 days to 400 days is a second per-span meter on top of the Pro fee. A long postmortem is not included in the sticker.
−The plan cards cap Free Cloud at 10 members and Pro at 50, while one FAQ on the same page says all plans include unlimited members. Use the card until Comet reconciles it. Enterprise (SSO, HIPAA, custom retention) publishes no list price.
Great value

Opik by Comet's pricing is extremely generous, offering a fully-featured Free tier that includes advanced capabilities like automated agent optimization and built-in guardrails.

Watch out

Potential overage fees for high usage

8
Galileo logo

Galileo

  • 3.0 on G2 (1 reviews)

Teams that want eval scores turned into a production guardrail

Galileo screenshot
+Free is $0 for 5,000 traces a month, unlimited users, and unlimited custom evals. Pro, on the yearly card in the table, includes 50,000 traces, standard RBAC, advanced analytics, and Slack support.
+The homepage says Luna models monitor 100% of traffic at 96% lower cost than the LLM-as-a-judge they were distilled from. Treat that percentage as Galileo's claim until you measure it on your traffic.
+Enterprise, which is a sales conversation, adds unlimited traces, hosted, VPC, or on-prem deployment, SSO, and real-time guardrails. The product story is that an offline eval becomes the thing that can block a live action.
−The Pro card says the price scales with traces, and it does not print the month-to-month rate next to the yearly card. A team that cannot pay yearly cannot budget from the card alone.
−This is a score and guardrail product. A latency or provider-outage page is a different tool, and Enterprise publishes no list price for the real-time guardrail tier.
9
Braintrust logo

Braintrust

  • 4.3 on G2 (54 reviews)

Teams whose monitor is a quality score chart, not an infrastructure page

+Pro is $249/mo and includes $100 of model credits, 5 GB of processed data, 50,000 scores, custom charts, and 30-day retention. Starter is $0 with $10 of credits, 1 GB, 10,000 scores, and 14-day retention.
+Past the included scores, Pro charges $1.50 per 1,000 and Starter charges $2.50 per 1,000. Past the included data, Pro is $3/GB and Starter is $4/GB. Retention past 30 days on Pro is $0.50 per GB per month.
+Custom charts are listed on the Pro card, built from metrics, score aggregations, and usage. That is the monitor when the thing you page on is a score, not a host CPU.
−Starter's 14-day retention and 1 GB will not hold a month of production traces, and the Pro fee does not remove the usage meters. A quiet month is still the platform fee before the overage.
−SSO and a BAA are Enterprise, which publishes no list price, so a healthcare or SSO requirement moves you off the card. If the chart you wanted was provider latency, Langfuse or Datadog is the closer fit, and the score-and-trace workflow sits in LLM observability.
Good value

Braintrust's pricing is quite generous, offering a robust Free tier with unlimited evals and 3 users, which is excellent for individual developers or small teams.

Watch out

No clear overage fees mentioned for usage

10
Fiddler logo

Fiddler

  • 4.4 on G2 (4 reviews)

Security teams that want a per-trace meter and a fast free guardrail

Fiddler screenshot
+Developer is $0.002 per trace and adds unified observability for agentic and predictive systems, custom evaluators, RBAC, and SSO, on SaaS. There is no separate seat line on that card, so the bill follows traces rather than seats.
+The free plan is a real-time guardrail for hallucinations, toxicity, PII, prompt injection, and jailbreaks, and Fiddler states latency under 80 ms, powered by its Centor models.
+Enterprise adds a VPC or on-prem deployment, a named customer success manager, and enterprise-grade guardrails. Use it when the free latency claim is the feature and the data still cannot leave the tenant.
−The published Developer price is a straight per-trace rate, with no monthly minimum and no volume discount on the card. The bill moves with trace count, and you have to multiply it yourself before the call.
−Enterprise publishes no list price. The free guardrail is not the observability line, so a demo of the under-80-ms check does not include the trace meter.
Good value

The pricing for Fiddler AI is generally fair, with a generous free tier for basic guardrails.

What LLM monitoring actually is

LLM monitoring is software that charts cost, latency, errors, and quality on live model traffic, then pages someone when a threshold breaks. If it cannot fire without a human opening a trace, it is a log, not a monitor.

The chart is the product, and the prompt text is evidence you may not be allowed to keep. Langfuse sells custom dashboards on every cloud plan and then meters alerts: 2 on Hobby, 20 on Core, 50 on Pro, 100 on Enterprise, so the page is capped by plan. Datadog prices Agent Observability by LLM span and, in its SDK docs, keeps token and cost metrics on all instrumented traffic even when you sample traces. A cheaper sample does not hide the spend chart. LangSmith's monitoring tab keeps working after a base trace's 14-day retention ends, because it reads metadata kept for more than 30 days. Helicone's alert docs cover error rate, cost, latency, token counts, and request count, with email or Slack.

A gateway monitor sees the request you were willing to proxy. Portkey and Helicone sit on that path, which is fast to turn on and blind to a tool call that never left your process. Skip the proxy when that failure lives inside the agent loop. An SDK monitor sees the steps you instrumented, which is what you want for an agent and what you will not get from a base-URL change.

Quality monitors are a third shape, for a team that pages on a wrong answer rather than a slow one. Galileo wants eval scores to become guardrails, Braintrust wants custom charts of scores, and Fiddler sells a per-trace meter next to a free guardrail. None of those is a substitute for a latency page if the score never runs.

The wrong adjacent purchase is a generic ingest bill with no LLM span. New Relic prices data at $0.40/GB after 100 GB a month, which can hold model telemetry and still does not name an LLM span. The head-to-head on the rest of the stack is Datadog vs New Relic, and prompt-injection blocking belongs in LLM security.

Why the chart outlives the prompt

The expensive mistake is paying for 14 days of prompt text and assuming the page will still explain last month's bill. LangSmith base traces are kept 14 days, and extended traces are 180 days for new traces after the September 14, 2026 change in LangSmith's billing docs. The monitoring tab is the part that stays when the review is a chart rather than the prompt. If your on-call playbook says "open the prompt," you needed the longer tier, and that upgrade is a second charge on the same trace.

Langfuse makes the same split with a clearer clock. Hobby keeps 30 days of data and 2 alerts. Core keeps 90 days and 20 alerts, enough to cover a monthly review. Pro keeps 3 years, 50 alerts, and the SOC 2 and ISO 27001 reports, at $199/mo, with the same 100,000 included units as Core. You are paying for history and the page, not for more included traffic. Enterprise is $2,499/mo and raises the alert cap to 100.

Portkey splits the clock inside one plan. Developer keeps logs for 3 days and metrics for 30. Production keeps logs for 30 days and metrics for 90. The chart answers what spiked, and the log answers what the model said. Buying the short log and expecting a postmortem is how teams discover the gap during the incident. Helicone's Hobby plan keeps 7 days and Pro keeps 1 month, so a review the following week needs the paid retention. Team, at $799/mo, keeps 3 months and is the first tier that lists SOC 2 and HIPAA.

Datadog's default on the first span block is 15-day trace retention. Longer trace retention is a higher rate per 10,000 spans, up to $7.50 per 10,000 on an annual bill for 90-day traces. Sample the traces if you must, and read the SDK note first. Token and cost metrics stay on 100% of instrumented traffic, while the bill follows the spans you send, and the default sample rate is 1.0. A sample that was meant to save money does not shrink the cost chart.

Key Features to Look For

  • An alert with a number on the plan (Essential)

    Langfuse prints the alert cap on each tier, and Helicone's docs alert on error rate, cost, latency, tokens, or request count. A dashboard you have to remember to open is not the monitor.

  • A chart that outlives the prompt (Essential)

    LangSmith's monitoring tab survives a base trace, and Portkey keeps Production metrics longer than the log. If the postmortem needs the text, buy the longer retention on purpose, because the chart is not a copy of the prompt.

  • A billing unit you can count (Essential)

    Langfuse units are traces plus observations plus scores. Datadog bills LLM spans. Opik bills spans at $5 per extra 100,000 on Pro. If you cannot count the unit, you cannot forecast the month.

  • Cost in dollars, not only tokens (Essential)

    Langfuse tracks token and cost on every cloud plan. Datadog shows estimated cost in US dollars when the provider is OpenAI, Azure OpenAI, or Anthropic. A token chart without a dollar still surprises finance.

  • A path that matches where the call happens (Important)

    Helicone and Portkey monitor what passes the proxy. Langfuse, LangSmith, Arize, and Opik monitor what the SDK or OpenTelemetry emits. A base-URL change will not see a local tool call.

  • A quality signal you can page on (Important)

    Galileo wants evals distilled into guardrails. Braintrust puts custom charts on Pro. Fiddler's free guardrail is aimed at harmful output. A latency page will not catch a confident wrong answer.

  • A place the prompts are allowed to sit (Important)

    Langfuse Cloud offers US, EU, or Japan, with a HIPAA region on Enterprise. Arize AX lists US, EU, or CA. Fiddler Enterprise adds VPC or on-prem, and that tier publishes no list price.

  • A spend ceiling you can set yourself (Nice to have)

    Langfuse's pricing FAQ says you can email yourself when spend crosses a threshold, and LangSmith lets a workspace set a spend limit that becomes a trace cap. Without that, the monitor's own bill is the next incident.

What to settle before the demo

  1. Write down whether the missing signal is a proxy request, an in-process agent step, or a quality score. Helicone and Portkey start at the URL. Langfuse, LangSmith, Arize, and Opik start at instrumentation. Galileo, Braintrust, and Fiddler start at a score or a guardrail.

  2. Decide how long the prompt text has to exist before you compare stickers. Fourteen days on a LangSmith base trace, seven on Helicone Hobby, and three on Portkey Developer are all real clocks. If legal wants the text for a quarter, you are on a higher tier before the logos matter.

  3. Count the billing unit on one real day of traffic. A Langfuse unit is a trace, an observation, or a score, so one user turn can be many units. A Datadog span and an Opik span are not that unit.

  4. Ask whether the chart survives sampling and deletion, because a vendor that deletes both is a log with a timer. Datadog's cost metrics stay on full traffic when traces are sampled, and LangSmith's monitoring tab stays after the base trace is gone.

  5. Separate the LLM line from the rest of the observability contract, or a host quote will look like this tool's price. Datadog Infrastructure Pro is $15 per host per month on an annual bill, and Agent Observability is a different row.

Evaluation Checklist

  • On Langfuse, count units as traces plus observations plus scores for one busy day, then see whether you still fit Hobby's 50,000 units and 2 alerts or you need Core's 20 alerts.

  • On Datadog, confirm the order is Agent Observability spans, note that 40,000 spans are marked free, and check whether you are on the annual, month-to-month, or on-demand column.

  • On LangSmith, set base retention and confirm the monitoring tab still has last month's chart, then see whether online evaluators are upgrading traces to 180 days.

  • On Helicone, send enough traffic to pass 10 logs per minute on Hobby, and confirm whether alerts are included or an add-on before you treat Pro as the paging plan.

  • On Portkey, exceed the Developer log cap on purpose in a test project and confirm requests still succeed while new logs stop, which is the failure mode of that free tier.

  • On Opik, decide whether you are on the open-source install, Free Cloud's 25,000 spans, or Pro's 100,000 spans, because those three are not the same allowance.

  • On Fiddler, price a week of traces at the published per-trace rate before you assume the free guardrail is the production monitor.

Pricing Overview

Alert platforms with a seat-free fee

Langfuse Core and Pro, Arize AX Pro, Opik Pro, and Portkey Production.

Monthly plan plus a usage unit

Span and trace meters

Datadog Agent Observability, LangSmith Plus, and Fiddler Developer.

Per span block, per seat, or per trace

Yearly cards and score charts

Galileo Pro on a yearly card, Braintrust Pro, and Helicone Team when you need the compliance line.

Annual display price, or a platform fee

Pricing Comparison

Best LLM Monitoring Tools in 2026 pricing comparison, as of September 2026
ToolPublished priceWhat that price buysBilling

Langfuse

$29/mo Core

100,000 units, 90-day access, 20 alerts. Hobby is free.

Monthly, plus units

Datadog

$160/mo annual

First 100,000 LLM spans with 15-day retention, and 40,000 spans are free.

Annual span block

$39/seat/mo

Plus includes 10,000 base traces, then pay as you go, and one Developer seat has no seat fee.

Per seat, plus usage

Helicone

$79/mo Pro

Unlimited seats, 10,000 requests included, 1-month retention.

Monthly, then usage

$50/mo Pro

50,000 spans, 10 GB, 30-day retention, 25 Signal issues.

Monthly

Portkey

$49/mo Production

100,000 logs, alerts, 30-day logs and 90-day metrics.

Monthly, plus logs

Opik

$19/mo Pro

100,000 spans, 60-day retention, up to 50 members.

Monthly, plus spans

Galileo

$100/mo yearly

Pro lists 50,000 traces, and the month-to-month rate is not printed.

Billed yearly

Braintrust

$249/mo Pro

5 GB included, 50,000 scores, custom charts, 30-day retention.

Monthly, plus usage

Fiddler

$0.002 per trace

Developer observability on SaaS, with RBAC and SSO, and the free guardrail is separate.

Per trace

Prices were read on September 24, 2026 from each vendor's pricing page. Datadog's annual column is the first span block in this table, and Galileo's Pro card is the yearly price. LangSmith does not print a per-trace dollar, only seat fees and a storage unit. See AI observability for the wider catalog, and LLM observability when the job is a trace replay.

Mistakes to Avoid

  • ×

    Buying Langfuse Pro because it sounds like more traffic does not raise the included units, which stay at 100,000. The jump pays for 3-year access, 50 alerts, higher ingest, and the compliance reports. If you only needed 20 alerts and 90 days, Core was the plan.

  • ×

    Treating a Datadog host price as the LLM monitor prices the wrong row. Infrastructure Pro is a per-host row, and Agent Observability is the span row, with its own annual, month-to-month, and on-demand columns. A quote that never says spans did not price this job.

  • ×

    Leaving LangSmith evaluators on the retention upgrade moves a trace from 14 days to 180 days. The monitoring chart did not need that text, and the invoice did, so turn the upgrade off if the page is all you open.

  • ×

    Shipping Helicone Hobby into production will miss the burst and forget it before the review. Ten logs a minute and 7 days of retention are the reason, and Pro is the first plan with a month of retention and a much higher ingest cap.

  • ×

    Assuming Portkey Developer keeps a production log fails past the recorded-log cap. Past 10,000 recorded logs, requests continue and new logs do not. The app looks healthy while the monitor goes dark.

  • ×

    Reading Fiddler's free guardrail as the monitoring price skips the meter. The guardrail card has no trace meter, so Developer is the per-trace line, and Enterprise is a quote for VPC or on-prem.

Expert Tips

  • →

    Count Langfuse units on one real trace before you pick a tier. Add the trace, every observation, and every score, because a single agent turn with a retrieval step and a judge is several units, which is why the 50,000-unit Hobby plan disappears on a small production app.

  • →

    Set the LangSmith spend limit before you enable online scores. The limit becomes a trace cap, and the score can extend retention. If you only needed the chart, keep the trace on the 14-day tier and let the monitoring tab hold the history.

  • →

    Sample Datadog traces only after you accept that the cost chart stays whole. The SDK keeps token and cost metrics on all instrumented traffic. The savings, if any, are the spans you no longer store, at the per-10,000 rate in the table.

  • →

    Put Helicone alerts through checkout, not only the docs, because the docs describe cost, latency, and error alerts while the upgrade call still lists alerts as an add-on flag. Confirm the invoice line before you promise the on-call a Slack channel.

  • →

    Split the gateway from the agent step. Portkey and Helicone see the proxied model call. An in-process tool call needs an SDK, which is the job described in AI agent observability. The platform layer around both is observability platforms.

  • →

    Ask Galileo for the month-to-month number in writing. The public Pro card is the yearly price and says the rate scales with traces. If procurement cannot sign a year, that card is not a budget.

Red Flags to Watch For

  • !

    A Langfuse order that prices Pro for the compliance reports and then assumes the included units grew with the fee. They did not.

  • !

    A Datadog quote that sells Infrastructure hosts and describes LLM spans as if they were included in the host price.

  • !

    A LangSmith workspace with no spend limit, and online evaluators left on the default that extends retention, so the chart quietly buys the long tier.

  • !

    A Helicone Hobby rollout on production traffic: 7-day retention and 10 logs per minute will drop the incident you meant to page on.

  • !

    A Portkey Developer plan treated as production logging. Past 10,000 recorded logs, the pricing page says requests keep working and further logs are not recorded.

  • !

    A Fiddler or Galileo enterprise conversation that copies a free guardrail demo onto a contract with no unit price.

The Bottom Line

Langfuse when the monitor is a shared chart and a numbered alert cap, and you will not pay per person to look at it. Skip it when you need SSO tomorrow, because that is the Teams add-on on top of Pro, or when you need the prompt text for years and will not pay for that retention.

Datadog when the page has to land in the same place as the rest of the on-call, and you will buy LLM spans as their own line. Skip it when you wanted a seat-free LLM tool and do not already run Datadog.

LangSmith when the stack is LangChain or LangGraph and the chart has to survive a 14-day trace. Helicone when a base URL is the integration you can actually ship, after you confirm alerts are on the invoice. Arize AX when you want hosted monitors at the Pro span cap and a free local Phoenix install for the notebook.

Portkey when the gateway should alert and the metrics must outlive the log, and you have read the line that renames the enterprise product. Opik when the budget is spans and a webhook on a trace error. Galileo when the monitor is a guardrail distilled from an eval, and you can live with a yearly card. Braintrust when the chart is a score and the Pro fee is acceptable before overages. Fiddler when you want a published per-trace rate and a free guardrail with a stated latency, and you will not confuse those two cards.

Cite this: Toolradar, "Best LLM Monitoring Tools in 2026", September 2026. Prices checked on vendor pages on September 24, 2026. No paid placement. Compared with the 43 AI observability tools we track.

Frequently Asked Questions

What is the best LLM monitoring tool in 2026?

Langfuse, if you want cost and error alerts without a per-seat fee. Core includes 100,000 units, 90 days of access, unlimited users, and 20 alerts, at the monthly price in the comparison table. Hobby is free and stops at 2 alerts and 50,000 units. Choose Datadog if the page has to land in an existing on-call tool, LangSmith if the app is LangChain, and Helicone if you can only change a base URL.

How much does an LLM monitoring tool cost in 2026?

As of September 24, 2026, Langfuse Core is $29/mo and Pro is $199/mo, then $8 per extra 100,000 units at the first overage step. Datadog's first 100,000 LLM spans are $160/mo on an annual bill, $200 month-to-month, or $240 on demand. LangSmith Plus is $39 per seat with 10,000 base traces included. Helicone Pro is $79/mo, Arize AX Pro is $50/mo, and Portkey Production is $49/mo. Opik Pro is $19/mo, Galileo Pro is $100/mo when billed yearly, Braintrust Pro is $249/mo, and Fiddler Developer is $0.002 per trace. Enterprise tiers on these products publish no list price.

Is there a free LLM monitoring tool in 2026?

Yes, and every free option below stops on a hard cap. Langfuse Hobby includes 50,000 units, 30 days, and 2 alerts. Datadog's price list marks 40,000 LLM spans free. LangSmith Developer includes 5,000 base traces on one seat until you add a card. Helicone Hobby includes 10,000 requests and 7 days. Arize AX Free includes 25,000 spans. Portkey Developer includes 10,000 logs and then stops recording. Opik's open-source build is free, and Free Cloud includes 25,000 spans. Galileo Free includes 5,000 traces. Braintrust Starter includes 1 GB and 10,000 scores. Fiddler's free plan is the guardrail, not the per-trace monitor.

How does Langfuse compare with Datadog for LLM monitoring?

Langfuse is the seat-free LLM chart: Core is the first paid row in the table, alerts are a printed cap, and a unit is a trace plus observations plus scores. You can self-host the MIT build and pay no unit fee, which suits a team that will run the upgrades itself. Datadog is the span line on an existing on-call stack: the annual block in the table covers the first 100,000 spans, and token and cost metrics stay on full traffic when you sample traces. Choose Langfuse when the LLM app is the thing you monitor. Choose Datadog when the page has to sit next to the rest of the infrastructure alerts, and keep host pricing off that comparison.

What is the difference between LLM monitoring and LLM observability?

Monitoring is the chart and the page: cost, latency, errors, and a threshold that notifies someone. Observability is the replay: the prompt, the tool calls, and the score on one request. LangSmith's monitoring tab still works after a 14-day base trace is deleted, which is the split in one product. If you need the replay, use the LLM observability guide. If you need the page, stay on this list.

Does the monitoring chart survive after the prompt is deleted?

On LangSmith, yes for the monitoring tab: base traces last 14 days, and the tab reads metadata kept for more than 30 days. On Portkey Production, metrics last 90 days and logs last 30. On Langfuse, the data-access clock is the chart clock: 30 days on Hobby, 90 on Core, and 3 years on Pro. If the review needs the prompt text itself, buy the longer retention. The chart is not a copy of the text.

Which LLM monitoring tool fits a LangChain app?

LangSmith Plus, at the per-seat price in the table, includes 10,000 base traces. The monitoring chart remains after the 14-day base trace expires, which is the production question. Langfuse is the better fit when the same app should not pay per seat and you will add the SDK yourself. Helicone fits when you will only change the base URL and can accept that in-process tool calls may never hit the proxy.

Cite this page: Toolradar, "Best LLM Monitoring Tools in 2026", updated September 2026, https://toolradar.com/guides/best-llm-monitoring-tools

Sources

Prices and plan details on this page come from each vendor's own pricing page, re-checked by the Toolradar pricing tracker:

Related Guides