Braintrust is an AI evaluation platform for testing prompts, models, and application behavior with production-like datasets and scoring workflows.
Know when your AI is wrong
Evals, tracing, and monitoring so model regressions show up in a dashboard before they show up in churn.
LangSmith is LangChain's platform for debugging, testing, evaluating, and monitoring LLM applications and agent workflows.
Langfuse is an open-source platform for tracing, prompt management, evaluations, and observability across LLM applications and agent systems.
Arize helps teams monitor, evaluate, and improve LLM applications, retrieval systems, and machine learning products once they are live in production.
Databricks' Mosaic AI suite for building, tuning, and serving LLMs with RAG, evaluation, and monitoring.
PromptLayer is a platform for prompt management, AI observability, evaluation workflows, and release control across LLM apps and agent systems.
End-to-end AI evaluation and monitoring platform for LLM product teams with prompt management and datasets.
Open-source tool for testing and evaluating LLM prompts, models, and RAG pipelines with side-by-side comparisons.
Open-source observability platform for LLM applications with tracing, monitoring, and prompt versioning via OpenLLMetry.
LangWatch helps teams monitor prompts, traces, and quality signals across production LLM applications and agent workflows.
Open-source LLM evaluation and observability platform from Comet.
Open source observability platform for AI apps with prompt tracking
LLM observability and evaluation platform for AI teams with prompt testing and production monitoring.
LLM evaluation platform providing automated testing and benchmarking for AI applications with custom metrics.
LLM observability and evaluation platform for building production-grade AI applications with tracing, monitoring, and annotation tools.
LLM observability and evaluation platform for monitoring, debugging, and improving AI application performance.
AI observability platform for monitoring, explaining, and analyzing ML model predictions in production.
Helicone is an open-source LLM observability platform providing request logging, cost tracking, caching, and evaluation tools for AI applications in production.
Open-source ML monitoring tool for evaluating, testing, and monitoring data and model quality in production.
LLM observability platform for collecting user feedback, evaluating model performance, and analytics for GenAI apps.
LLM operations platform for managing prompts, models, and deployments in production with collaboration tools for AI product teams.
AI security platform defending LLM applications from prompt injection and exploits.
LLM testing and deployment platform with prompt playground and observability for AI developers.
Vellum helps teams design, test, and orchestrate prompt-driven workflows and LLM-powered products with a more controlled production pipeline.
Open-source LLM observability and analytics platform with logging, user analytics, and prompt templates.
Open-source LLM evaluation and observability tool for production AI pipelines.
LLM app platform with prompt engineering, evaluation, and observability for teams
ML observability platform for monitoring models in production with drift and performance tracking.
Open source observability and evaluation platform for LLM applications with tracing.
AI observability platform for monitoring errors and debugging issues in LLM applications.