Cubbie Conference December 10, 2026 in San Francisco Get tickets

Know when your AI is wrong

Evals, tracing, and monitoring so model regressions show up in a dashboard before they show up in churn.

Trace-to-quality AI loop

LangSmith + Braintrust + Arize

Company size
Purchasing
Compliance (as stated by vendors)

Individual products and alternatives

Sort

Braintrust

AI Evaluation and Testing

Braintrust is an AI evaluation platform for testing prompts, models, and application behavior with production-like datasets and scoring workflows.

Cubbie's pickFree planFree trialSOC 2GDPRHIPAA
Tiered plans

Price$249 / month

Price position: above the category median.
View profileBuy / SubscribeCompare

LangSmith

AI Observability

LangSmith is LangChain's platform for debugging, testing, evaluating, and monitoring LLM applications and agent workflows.

Cubbie's pickFree planFree trialTiered plansLLM OrchestrationAI Observability

Price$39+ / seat / month

Starting price: below the category median.
View profileBuy / SubscribeCompare

Langfuse

AI Observability

Langfuse is an open-source platform for tracing, prompt management, evaluations, and observability across LLM applications and agent systems.

Cubbie's pickFree planFree trialSOC 2ISO 27001GDPR
Tiered plans

Price$29+ / month

Price position: below this category's displayed range.
View profileBuy / SubscribeCompare

Arize

Observability Platforms

Arize helps teams monitor, evaluate, and improve LLM applications, retrieval systems, and machine learning products once they are live in production.

Cubbie's pickFree planFree trialSOC 2ISO 27001PCI DSS
Tiered plans

Price$50+ / month

Starting price: above the category median.
View profileBuy / SubscribeCompare

PromptLayer

Prompt Engineering Tools

PromptLayer is a platform for prompt management, AI observability, evaluation workflows, and release control across LLM apps and agent systems.

Free planSOC 2GDPRHIPAACCPATiered plans

Price$50+ / seat / month

View profileBuy / SubscribeCompare

LangWatch

AI Observability

LangWatch helps teams monitor prompts, traces, and quality signals across production LLM applications and agent workflows.

Free planFree trialTiered plansPrompt Engineering ToolsAI Observability

Price$69+ / month

Starting price: below the category median.
View profileBuy / SubscribeCompare

Maxim AI

AI Evaluation and Testing

End-to-end AI evaluation and monitoring platform for LLM product teams with prompt management and datasets.

Free planFree trialSOC 2ISO 27001HIPAAGDPR
Per seat

Price$39+ / user / month

Starting price: below the category median.
View profileBuy / SubscribeCompare

Prompt Security

AI Security Platforms

Enterprise AI security platform protecting against prompt injection, data leakage, and harmful AI outputs.

Free trialTiered plansAI GovernanceAI Security Platforms

Price$834+ / month

View profileBuy / SubscribeCompare

Opik

AI Observability

Open-source LLM evaluation and observability platform from Comet.

Free planFree trialTiered plansAI ObservabilityMLOps Platforms

Price$19 / month

Price position: below this category's displayed range.
View profileBuy / SubscribeCompare

Evidently AI

AI Evaluation and Testing

Open-source ML monitoring tool for evaluating, testing, and monitoring data and model quality in production.

Free planFree trialTiered plansAI Evaluation and TestingMLOps Platforms

Price$80 / month

Price position: below the category median.
View profileBuy / SubscribeCompare

Orq.ai

LLM Ops

LLM operations platform for managing prompts, models, and deployments in production with collaboration tools for AI product teams.

Free planFree trialSOC 2GDPRHIPAATiered plans

Price$5+ / month

View profileBuy / SubscribeCompare

Aimon

LLM Orchestration

LLM observability and monitoring platform detecting hallucinations, toxicity, and quality issues in production.

LLM Orchestration

Price$49 / month

View profileBuy / SubscribeCompare

Vellum

Prompt Engineering Tools

Vellum helps teams design, test, and orchestrate prompt-driven workflows and LLM-powered products with a more controlled production pipeline.

Free planFree trialTiered plansPrompt Engineering ToolsLLM Orchestration

Price$50+ / month

View profileBuy / SubscribeCompare

Mosaic AI

AI Infrastructure

Databricks' Mosaic AI suite for building, tuning, and serving LLMs with RAG, evaluation, and monitoring.

Free trialUsage-basedAI Infrastructure

PriceContact sales

View profileBuy / SubscribeCompare

Traceloop

Observability Platforms

Open-source observability platform for LLM applications with tracing, monitoring, and prompt versioning via OpenLLMetry.

Free planFree trialSOC 2HIPAATiered plansAI Observability
View profileBuy / SubscribeCompare

Promptfoo

AI Evaluation and Testing

Open-source tool for testing and evaluating LLM prompts, models, and RAG pipelines with side-by-side comparisons.

Free planFree trialAI Evaluation and TestingAI Translation
View profileBuy / SubscribeCompare

Parea

AI Observability

LLM observability and evaluation platform for AI teams with prompt testing and production monitoring.

Free planFree trialTiered plansLLM OrchestrationAI Observability
View profileBuy / SubscribeCompare