Braintrust is an AI evaluation platform for testing prompts, models, and application behavior with production-like datasets and scoring workflows.
Platforms for systematically evaluating, benchmarking, and testing AI model outputs for accuracy, safety, and reliability.
Braintrust is an AI evaluation platform for testing prompts, models, and application behavior with production-like datasets and scoring workflows.
LangWatch helps teams monitor prompts, traces, and quality signals across production LLM applications and agent workflows.
Browserbase provides a reliable, managed browser infrastructure for AI agents, enabling web scraping, automated testing, and agent-based web interactions at scale.
LLM evaluation platform providing automated testing and benchmarking for AI applications with custom metrics.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Structured interview platform with question libraries, scorecards, and collaborative evaluation for reducing hiring bias.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Open-source evaluation tool for testing and benchmarking LLM applications in CI/CD pipelines.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Open-source AI quality and testing platform for detecting vulnerabilities, biases, and hallucinations in ML models.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Open-source ML monitoring tool for evaluating, testing, and monitoring data and model quality in production.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Open-source LLM evaluation and observability tool for production AI pipelines.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.LLM monitoring and evaluation platform detecting hallucinations, toxicity, and quality issues in production AI.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Angular is a TypeScript-based web application framework by Google for building enterprise-grade single-page applications with comprehensive tooling, testing, and state management.
Checkly is a monitoring-as-code platform for API monitoring and end-to-end testing that integrates with CI/CD pipelines to detect issues before they affect users.
Pricing$0 to $360
$0 to $360 per seat per year, against a category median of 1,188 dollars.LLM testing and evaluation platform with continuous testing pipelines.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.End-to-end AI evaluation and monitoring platform for LLM product teams with prompt management and datasets.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Open-source platform for prompt engineering, evaluation, and deployment of LLM applications with team collaboration.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Open-source ML testing and monitoring platform for validating data and models throughout the development lifecycle.
Pricing$0
$0 per seat per year, against a category median of 1,188 dollars.Chromatic is a visual testing and review platform for Storybook components that catches UI bugs by comparing component snapshots across changes.
Pricing$0 to $1,788
$0 to $1,788 per seat per year, against a category median of 1,188 dollars.ArchiCAD by Graphisoft provides BIM authoring software for architects with integrated design and documentation in a single model.
Pricing$2,940
$2,940 per seat per year, against a category median of 1,188 dollars.Out of 62 products here. A product can suit more than one size, and one that states none is in no bar.