
Braintrust
AI Evaluation and Testing
Braintrust is an AI evaluation platform for testing prompts, models, and application behavior with production-like datasets and scoring workflows.
Price$249 / month
Curated software listings.

AI Evaluation and Testing
Braintrust is an AI evaluation platform for testing prompts, models, and application behavior with production-like datasets and scoring workflows.
Price$249 / month

AI Observability
LangSmith is LangChain's platform for debugging, testing, evaluating, and monitoring LLM applications and agent workflows.
Price$39+ / seat / month

Prompt Engineering Tools
PromptLayer is a platform for prompt management, AI observability, evaluation workflows, and release control across LLM apps and agent systems.
Price$50+ / seat / month

AI Evaluation and Testing
LLM testing and evaluation platform with continuous testing pipelines.
Price$299+ / month
AI Evaluation and Testing
Open-source evaluation tool for testing and benchmarking LLM applications in CI/CD pipelines.

AI Evaluation and Testing
LLM monitoring and evaluation platform detecting hallucinations, toxicity, and quality issues in production AI.

AI Evaluation and Testing
Open-source platform for prompt engineering, evaluation, and deployment of LLM applications with team collaboration.

Talent Marketplace
AI recruiting platform that uses video interviews to match candidates with tech companies.

AI Evaluation and Testing
AI platform delivering annotation, evaluation, and red teaming for enterprise LLMs including adversarial testing for safety and security.
PriceContact sales

AI Observability
LLM observability and evaluation platform for AI teams with prompt testing and production monitoring.

AI Evaluation and Testing
Open-source tool for testing and evaluating LLM prompts, models, and RAG pipelines with side-by-side comparisons.

AI Developer Tools
Developer platform for building, evaluating, and deploying generative AI applications to production.

LLM Ops
Open-source library for evaluating and tracking LLM and RAG application quality with feedback functions.

AI Evaluation and Testing
Open-source AI governance testing framework by Singapore's IMDA for assessing AI system trustworthiness.

AI Evaluation and Testing
AI testing platform for debugging and evaluating ML and generative AI systems.

AI Observability
LLM observability and evaluation platform for monitoring, debugging, and improving AI application performance.
AI Evaluation and Testing
LLM evaluation platform providing automated testing and benchmarking for AI applications with custom metrics.

AI Evaluation and Testing
AI application development platform for testing, evaluation, and deployment of LLM systems.