answer·4 min read

Agent observability pricing compared

Free tiers, entry prices and billing meters for eight agent observability platforms. Compare what you pay to trace sessions, evaluate behavior, find recurring failures, alert your team and act on what the system finds.

the short answer

Published entry prices run from $29 a month for Langfuse Core to $299 a month for Raindrop Pro. Failproof AI Team and Latitude Pro cost $99, Galileo Pro is listed at $100 billed yearly through a demo, Braintrust Pro costs $249, and Future AGI is usage-based with add-ons from $250. Judgment Labs does not publish pricing. Compare the products on the work they remove as well as the usage meter: Failproof AI is built to find failure modes automatically with code-based and LLM-based evaluations, turn them into findings and recommended fixes, and then steer or block the agent when needed.

Lowest paid entry
Langfuse Core, $29 a month.
No published price
Judgment Labs - demo only.
SSO below Enterprise
Failproof AI Scale ($599), Future AGI Boost ($250, OAuth) or Scale ($750, SAML), Langfuse Pro plus a $300 add-on.
Failure intelligence
Failproof AI finds recurring failures with code and LLM evaluations, recommends fixes, and can turn findings into agent policies.

The Table

VendorFree tierEntry paid tierWhat is meteredSSO fromSelf-hostCan it stop an action?
Failproof AI5,000 runs and 100 evals a month, hard cap; 1 user; 30-day retentionTeam: $99/month - 50,000 runs and 50,000 evals, 5 usersRuns ($1.00 per 1k over) and evals ($0.05 each over)Scale, $599/monthMIT CLI free; self-hosted Cloud on EnterpriseYes - hook-level policies, including the free CLI
Judgment LabsNot publishedNot publishedNot publishedNot publishedSDK Apache-2.0; platform "coming soon"No
RaindropHobby: 1,000 events a month, then ingestion stopsPro: $299/month, 14-day trialEvents: $0.003 each to 1M, then $0.002EnterpriseBeta for select partnersNo
LatitudeStarter: 20K credits, 30-day retention, unlimited seatsPro: $99/month - 100K credits, 90-day retentionCredits ($20 per extra 10K); unit not definedEnterprise (SAML)Free under MIT; on-prem on EnterpriseNo
Future AGI50 GB storage, 2K AI credits, 100K gateway requests, unlimited seatsPay-as-you-go; Boost add-on $250/monthStorage, AI credits, gateway requests, simulationBoost $250 (OAuth); Scale $750 (SAML)Free under Apache-2.0; on-prem gateway needs an enterprise licenceYes - Protect and tool permissions, for traffic through its gateway
Galileo5,000 traces, unlimited users; the only self-serve tierPro: $100/month billed yearly, 50,000 traces; "Book a Demo"Traces; rate above Pro not publishedEnterpriseVPC or on-prem on EnterpriseReal-time guardrails on Enterprise; Agent Control checks LLM and tool inputs and outputs
LangfuseHobby: 50k units, 30 days of data, 2 usersCore: $29/month - 100k units, 90 daysUnits (traces, observations, scores): $8 per 100k overPro + $300 Teams add-on, or Enterprise ($2,499)Free under MIT; Enterprise self-host at custom pricingNo
BraintrustStarter: 1 GB processed data, 10k scores, 14-day retentionPro: $249/month - 5 GB, 50k scoresProcessed data ($3/GB on Pro), scores ($1.50 per 1k on Pro), model creditsEnterpriseHybrid on EnterpriseNone documented
Prices were checked against each vendor's pricing page or plan documentation. Raindrop's /pricing page returns a 404, so its figures come from its plan documentation; Judgment Labs has no pricing page. Galileo is now Splunk Agent Observability and is listed under the name buyers still search for.

Read the meter column before the price column. A run, an event, a trace, a unit, a credit and a gigabyte are not the same thing, and the cheapest plan for your team depends on how many of each your agent produces. Failproof AI's own prices are on /pricing.

What Each Meter Actually Counts

  • Runs and evals (Failproof AI): agent runs ingested into Cloud, and evaluations counted separately - each session-and-evaluation pair is one billable evaluation, so two evaluations on one session count twice. The free tier is a hard cap with no overage.
  • Events (Raindrop): the plans doc prices events without defining one, so count what your SDK actually sends.
  • Traces (Galileo): priced by trace count; how many your agent produces depends on how you instrument it.
  • Units (Langfuse): "any tracing data point" - traces, observations (spans, events and generations) and scores. An agent session with twenty tool calls is twenty-odd units, not one.
  • Credits (Latitude): the pricing page does not say what a credit measures. Ask before you estimate.
  • Processed data and scores (Braintrust): gigabytes of data processed plus each score, and model credits at token rates. Trace size drives the bill.
  • Storage, AI credits and gateway requests (Future AGI): storage at $2 per GB above 50 GB, AI credits for evaluations, guardrails and synthetic data at $10 per 1,000, gateway requests at $5 per 100,000, plus simulation. The page does not say how many credits one evaluation uses.

Two Worked Examples

Every number below rests on assumptions you should replace with your own. One agent session is one Failproof AI run, one Galileo trace and one Raindrop event, and each evaluated session runs one Failproof AI evaluation. In Langfuse a session produces 1 trace, 20 observations and 2 scores, so 23 units. In Braintrust a session is 40 KB of processed data, and each evaluated session gets 2 scores. Overage is billed pro rata. Judge-model tokens are excluded, because on most of these platforms the judge is a model you pay for separately.

A: 20,000 Sessions a Month, 5 People, 10% of Sessions Evaluated

VendorPlanArithmeticPer month
Failproof AITeam20,000 runs and 2,000 evals, both inside the allowance$99
LangfuseCore20,000 × 23 = 460,000 units; $29 + 3.6 × $8$57.80
BraintrustStarter0.8 GB and 4,000 scores, inside the allowance; 14-day retention$0
BraintrustProInside the allowance; 30-day retention$249
GalileoPro20,000 traces, inside 50,000; Pro is sold through a demo$100, billed as $1,200 a year
RaindropPro$299 + 20,000 × $0.003$359
LatitudeProCredits per session not published$99 + unknown
Future AGIFree + pay-as-you-go0.8 GB storage, inside 50 GB; credits per evaluation not publishedNot computable
Judgment Labs-Not publishedNot published

B: 500,000 Sessions a Month, 25 People, SSO Required, 4% of Sessions Evaluated

VendorPlanArithmeticPer month
Failproof AIScale500,000 runs and 20,000 evals, inside the allowance; SSO/SAML included$599
LangfusePro + Teams add-on11.5M units; $199 + $300 + 114 × $8, before volume discounts$1,411
Future AGIScale add-on + usage$750 for SAML SSO; 20 GB storage inside 50 GB; eval credits not computable$750 + credits
BraintrustEnterpriseSSO is Enterprise-onlyNot published
GalileoEnterpriseSSO, and volume above 50,000 traces, are EnterpriseNot published
RaindropEnterpriseSSO is Enterprise-onlyNot published
LatitudeEnterpriseSAML SSO is Enterprise-onlyNot published
Judgment Labs-Not publishedNot published

Drop the SSO requirement and the picture changes. Braintrust Pro would be $249 + 15 GB × $3 = $294 (20 GB processed, 40,000 scores inside the allowance), and Raindrop Pro $299 + 500,000 × $0.003 = $1,799. Failproof AI Team would be $99 + 450 × $1.00 = $549: all 20,000 evaluations fit inside its 50,000-evaluation allowance. That makes Team $50 cheaper than Scale at this usage, though Scale adds SSO/SAML, more users and substantially more included capacity. Trace size does the most work in the Braintrust figure: at 200 KB a session it becomes 100 GB and $534.

What the Numbers Mean

Three patterns hold across the table. The free tiers are for trying, not running: they cap at 1,000 to 50,000 of something, and two of them, Failproof AI and Raindrop, stop rather than bill past the cap. SSO is the real price line for a mid-sized company: on most platforms it only comes with Enterprise, which means a sales call and an unpublished price. A published price is not always a self-serve one, either: Galileo lists Pro at $100 a month but routes it through "Book a Demo", so its Free tier is the only one you can start alone. And unpublished is not the same as expensive; it is slower to find out.

An evaluation allowance and the model cost behind an LLM-based evaluation are separate parts of the estimate unless a vendor explicitly says otherwise. Failproof AI counts each session-and-evaluation pair against the plan's evaluation allowance. Braintrust bills model credits at token rates, Future AGI charges AI credits for evaluations, and Galileo prices Luna-2 by token on Enterprise. For every platform, confirm whether the published allowance includes model usage or whether you also pay the model provider, then budget sessions evaluated × tokens per judgment × model price where applicable.

What Happens After the Platform Finds a Failure?

Detection is only useful when it leads to a decision. Failproof AI uses code-based and LLM-based evaluations to find failures across production sessions, groups related failures into findings, and recommends a fix. Alerts bring the right findings to the team. When the fix is a change to agent behavior, the finding can become a policy that steers or blocks the agent.

  • Failproof AI connects the finding to the fix. A generated policy can be tested against past agent activity, deployed in observe mode and then used to steer or block a tool action.
  • Future AGI checks model and tool traffic sent through its gateway or SDK. Protect and Tool Permissions can block, warn, mask or log.
  • Galileo offers hosted guardrails on Enterprise. Agent Control checks model and tool inputs and outputs, and can deny or steer a request before allowing it through.

Judgment Labs, Raindrop, Latitude, Langfuse and Braintrust evaluate completed activity and alert the team. They do not provide the same path from a discovered failure to a tested policy. If the agent can run commands, move money, change files or send messages, check whether the product helps your team fix the behavior as well as detect it.

When You Do Not Need to Buy Anything

If one person looks after one agent that handles a few thousand sessions a month, the free tiers above cover it, and so does a pytest file with a handful of checks. If you need everything on your own infrastructure and have someone to run it, Langfuse and Latitude self-host under MIT. Future AGI has an Apache-2.0 core with separately licensed enterprise code. In each case, the practical price includes your servers and the time required to operate the stack.

If you only need inexpensive trace storage and scoring, Failproof AI is not the cheapest option at small volume. In example A, Braintrust Starter comes to $0 and Langfuse Core to $57.80, compared with $99 for Failproof AI Team. Failproof AI becomes useful when nobody can read every session and you need the system to find recurring failure modes, explain what is going wrong and recommend a fix. Steering or blocking the agent is an additional step for failures that can be prevented at runtime.

FAQ

Which agent observability platform is cheapest?

On published entry prices, Langfuse Core at $29 a month is the lowest paid tier, and Braintrust Starter is $0 with metered overage. The meters differ - runs, events, traces, units, credits and gigabytes - so the cheapest platform depends on your volume. At 20,000 sessions a month, Braintrust Starter came to $0 and Langfuse Core to $57.80 in the worked example here.

Does Judgment Labs publish pricing?

No. Judgment Labs has no pricing page, its sitemap lists none, and the homepage call to action is "Book demo". Plans, a free tier and feature gating are not published. The judgeval SDK is Apache-2.0 and free to install; the platform it reports to is sold through sales.

Which platforms include SSO without an enterprise contract?

Failproof AI includes RBAC and SSO/SAML on Scale at $599 a month. Future AGI adds OAuth SSO with its $250 Boost add-on and SAML with SCIM on its $750 Scale add-on. Langfuse offers enterprise SSO on Pro with a $300-a-month Teams add-on. Raindrop, Latitude, Galileo and Braintrust put SSO on Enterprise.

Are judge model costs included in these prices?

Do not assume that an evaluation allowance includes the model tokens used by an LLM-based evaluation. Failproof AI counts each session-and-evaluation pair against the allowance shown on /pricing. Braintrust bills model credits at token rates, Future AGI charges AI credits for evaluations, and Galileo prices Luna-2 per token on Enterprise. Confirm the model-cost treatment with each vendor and estimate any separate judge spend from sessions evaluated × tokens per judgment × model price.

What does a Langfuse billable unit count?

Langfuse defines a billable unit as any tracing data point sent to the platform: traces, observations (spans, events and generations) and scores. One agent session with 20 LLM calls and tool spans plus 2 scores is about 23 units. Core includes 100,000 units for $29 a month and charges $8 per 100,000 beyond that, lower with volume.

Get Started

Failproof AI is free to start. It finds recurring failure modes across agent sessions using code-based and LLM-based evaluations, groups the evidence into findings, and recommends fixes. Bring the eval suite you already have, alert the right owner when behavior drifts, and turn a tested fix into a policy that prevents the failure from recurring. See pricing for the tiers.

Sources

Checked against each vendor's own site and docs on 2026-09-14. Products change; if a detail here is out of date, tell us at support@befailproof.ai.

  1. Judgment Labs homepage
  2. Judgment Labs sitemap
  3. Judgment Labs docs: Self-hosting
  4. Raindrop docs: Plans
  5. Latitude pricing
  6. Future AGI pricing
  7. Future AGI docs: Protect
  8. Galileo pricing
  9. Galileo docs: Luna-2
  10. Galileo docs: Agent Control
  11. Galileo: Announcing Agent Control
  12. Galileo release notes
  13. Langfuse pricing
  14. Langfuse self-hosted pricing
  15. Braintrust pricing
  16. Braintrust docs: Self-hosting
  17. Failproof AI docs: Policy editor
  18. Failproof AI docs: Supported harnesses
  19. Failproof AI docs: Evaluations