
Braintrust
AI Evaluation and Testing
Braintrust is an AI evaluation platform for testing prompts, models, and application behavior with production-like datasets and scoring workflows.
Price$249 / month
Price position: above the category median.Evals, tracing, and monitoring so model regressions show up in a dashboard before they show up in churn.

AI Evaluation and Testing
Braintrust is an AI evaluation platform for testing prompts, models, and application behavior with production-like datasets and scoring workflows.
Price$249 / month
Price position: above the category median.
AI Observability
LangSmith is LangChain's platform for debugging, testing, evaluating, and monitoring LLM applications and agent workflows.
Price$39+ / seat / month
Starting price: below the category median.
AI Observability
Langfuse is an open-source platform for tracing, prompt management, evaluations, and observability across LLM applications and agent systems.
Price$29+ / month
Price position: below this category's displayed range.
Observability Platforms
Arize helps teams monitor, evaluate, and improve LLM applications, retrieval systems, and machine learning products once they are live in production.
Price$50+ / month
Starting price: above the category median.
Prompt Engineering Tools
PromptLayer is a platform for prompt management, AI observability, evaluation workflows, and release control across LLM apps and agent systems.
Price$50+ / seat / month
AI Observability
LangWatch helps teams monitor prompts, traces, and quality signals across production LLM applications and agent workflows.
Price$69+ / month
Starting price: below the category median.
AI Evaluation and Testing
End-to-end AI evaluation and monitoring platform for LLM product teams with prompt management and datasets.
Price$39+ / user / month
Starting price: below the category median.
AI Security Platforms
Enterprise AI security platform protecting against prompt injection, data leakage, and harmful AI outputs.
Price$834+ / month

AI Observability
Open-source LLM evaluation and observability platform from Comet.
Price$19 / month
Price position: below this category's displayed range.
AI Evaluation and Testing
Open-source ML monitoring tool for evaluating, testing, and monitoring data and model quality in production.
Price$80 / month
Price position: below the category median.
LLM Ops
LLM operations platform for managing prompts, models, and deployments in production with collaboration tools for AI product teams.
Price$5+ / month
LLM Orchestration
LLM observability and monitoring platform detecting hallucinations, toxicity, and quality issues in production.
Price$49 / month

Prompt Engineering Tools
Vellum helps teams design, test, and orchestrate prompt-driven workflows and LLM-powered products with a more controlled production pipeline.
Price$50+ / month
AI Infrastructure
Databricks' Mosaic AI suite for building, tuning, and serving LLMs with RAG, evaluation, and monitoring.
PriceContact sales

Observability Platforms
Open-source observability platform for LLM applications with tracing, monitoring, and prompt versioning via OpenLLMetry.

AI Evaluation and Testing
Open-source tool for testing and evaluating LLM prompts, models, and RAG pipelines with side-by-side comparisons.

AI Observability
LLM observability and evaluation platform for AI teams with prompt testing and production monitoring.
AIOps Platforms
Open source observability platform for AI apps with prompt tracking