1 dépôt
Utilities for quantifying the quality and reliability of multi-step agent workflows using automated metrics.
Distinct from Performance Metrics: Distinct from Performance Metrics: focuses on the evaluation of complex agent workflow performance rather than basic model precision/recall statistics.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Workflow Performance Scorers. Refine with filters or upvote what's useful.
Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs
Quantifies the quality of AI outputs and agent workflows using automated scoring to ensure consistent performance.