19 个仓库
Systems that automate the assessment of artificial intelligence outputs and reasoning quality through comparative analysis or secondary model verification.
Explore 19 awesome GitHub repositories matching artificial intelligence & ml · AI Evaluation Frameworks. Refine with filters or upvote what's useful.
MetaGPT is an agentic workflow engine and multi-agent orchestration framework designed to automate complex software engineering and data analysis tasks. It functions as an automated software factory that transforms high-level natural language requirements into functional web applications, technical documentation, and production-ready code. By utilizing a runtime environment that manages the lifecycle of specialized agents, the platform bridges the gap between user intent and finished software components. The system distinguishes itself through role-based agent orchestration and dynamic task d
Validates results by running multiple AI teams in parallel to compare outputs and test workflows for correctness and robustness.
Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model
Provides a system for running systematic experiments, benchmarking model outputs, and automating quality scoring.
This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap
Ships a systematic approach for assessing generative AI outputs and reasoning quality through comparative analysis.
Evals is a framework designed for automating, managing, and executing repeatable benchmarking suites to analyze the quality and performance of language models. It provides a platform for running standardized tests to measure model accuracy and track behavioral changes over time. The system distinguishes itself through a modular architecture that uses a standardized adapter layer to normalize inputs and outputs, allowing different models to be swapped and tested interchangeably. It supports the creation of custom benchmarks using proprietary data, enabling quality assurance on sensitive tasks
Enables the definition of bespoke evaluation logic and datasets to assess unique model behaviors.
Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn
Provides a comprehensive framework for automated testing, dataset management, and model-as-a-judge scoring.
This project is a comprehensive framework for building and managing autonomous agent systems. It provides a unified architecture for orchestrating multi-agent societies, where specialized agents collaborate through roleplay to decompose and solve complex tasks. The system integrates language models with external environments, enabling agents to perform real-world actions through a standardized tool-calling abstraction layer. The framework distinguishes itself through its focus on iterative reasoning and data reliability. It employs automated feedback loops to refine agent outputs and self-eva
Captures and structures the information returned by an agent to facilitate further reasoning or integration into a multi-agent workflow.
ChatALL is a desktop application that functions as a multi-model chat client and aggregator for artificial intelligence services. It enables users to send a single prompt to multiple AI models simultaneously, allowing for the side-by-side comparison of generated responses within a unified interface. The application distinguishes itself through a local-first approach to data management, ensuring that all conversation logs and user configurations are stored directly on the user's device. This architecture supports privacy and offline access while providing a centralized system for managing and
Facilitates comparative analysis by sending prompts to multiple AI services and displaying their responses side by side.
This project serves as a comprehensive educational resource and technical handbook for engineers building applications powered by large language models. It provides a structured framework for mastering the principles of artificial intelligence engineering, covering the full lifecycle of model development from initial design to production deployment. The repository distinguishes itself by offering a deep dive into the practical implementation of advanced design patterns, including retrieval-augmented generation, agentic tool orchestration, and parameter-efficient model adaptation. It emphasize
Automates the assessment of artificial intelligence outputs and reasoning quality through comparative analysis.
This project is a development platform for managing the lifecycle of generative artificial intelligence models. It provides a unified environment for accessing, fine-tuning, and deploying large language models, serving as an orchestrator that handles the integration of diverse models into custom applications. The platform distinguishes itself by offering a managed infrastructure for hosting and scaling models, which removes the requirement for manual server maintenance or configuration. It includes integrated tools for supervised fine-tuning and vector embedding optimization, allowing for the
Assessing the performance and reliability of generated content using automated testing services and rubrics to ensure alignment with project requirements.
Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin
Compares outputs from multi-step AI workflows to verify correctness.
Hive is an artificial intelligence workflow automation engine and development platform designed for building and deploying autonomous agents. It provides a framework for orchestrating complex, multi-step business processes by coordinating tasks across multiple specialized agents using directed graph structures. The platform distinguishes itself through a focus on production-grade reliability and state management. It maintains persistent execution context and conversation history on disk, enabling crash recovery and continuity for long-running automated sessions. Furthermore, it incorporates a
Validates agent outputs using a multi-level pipeline of deterministic rules, semantic assessment, and human oversight.
This project is an AI-powered IDE extension and LLM coding assistant that provides a conversational interface for generating, refactoring, and debugging code. It functions as an AI agent framework and a Model Context Protocol client, connecting AI models to external data sources and tools to automate complex development tasks. The system is distinguished by its use of autonomous AI agents capable of multi-step task execution, including the ability to read files, modify code, and run terminal commands iteratively. It supports recursive agent orchestration through subagent delegation and employ
Produces the necessary code and assets to implement evaluation frameworks for testing AI agents.
Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo
Uses an AI agent to suggest and define evaluator patterns based on a target task description.
Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of
Provides a comprehensive framework for automating the assessment of AI outputs and reasoning quality.
adk-go is an agent orchestration engine and multi-agent framework for building, coordinating, and scaling systems of large language model agents. It provides a tool integration kit to connect agents with external APIs, custom functions, and diverse data sources. The project utilizes graph-based workflow orchestration to blend deterministic logic with adaptive reasoning. It supports modular multi-agent composition, allowing specialized agents to be organized into hierarchical structures to manage complex tasks through coordinated workflows. The framework includes tools for performance evaluat
Measures agent reliability by comparing model responses against predefined schemas and validation tools.
本项目是一套标准化的抽象与推理问题集,旨在评估人工智能模型学习新规则的能力。它作为流体智力测试和推理基准,利用网格谜题和程序合成数据集来评估智能体如何从示例中生成算法。 该项目专注于衡量通用流体智力和零样本泛化能力,测试系统在不依赖特定任务训练的情况下,能否将学到的逻辑应用于未见过的任务。它为抽象推理研究提供了一个框架,专门评估程序合成以及模型发现新颖网格变换中潜在模式的能力。 系统结合了基于网格的领域表示和几何与拓扑操作的组合搜索空间,并包含一个人工参与界面,允许手动构建输出网格,从而为验证和基准测试定义事实标准。
Provides a standardized set of abstraction and reasoning problems to assess the quality of AI reasoning.
该项目是一系列生成式 AI 实现,专注于 AI 代理、检索增强生成(RAG)流水线和向量搜索集成的开发。它提供了一个将托管云数据库连接到语言模型的框架,以创建上下文感知应用。 该项目涵盖了使用多步推理和外部工具完成任务的自主代理的编排。它包括使用高维嵌入进行语义检索的实现,以及使用模型无关的提示词(prompting)以确保不同大语言模型之间输出的一致性。 其他功能包括使用地面真值(ground-truth)评估框架来衡量 AI 性能的准确性和可靠性。该项目还演示了用于存储和管理应用数据及向量嵌入的云数据库账号设置。
Uses AI evaluation frameworks to measure the accuracy and reliability of model outputs.
这是一个自托管的 AI 监控技术栈,集 LLM 可观测性平台、AI 评估框架和 OpenTelemetry 链路分析器于一体。它旨在捕获并分析 LLM 链路、会话和遥测数据,以监控 AI Agent 的性能。 该平台作为 Model Context Protocol 服务器,将工作区功能暴露为 AI 编码 Agent 的工具。它支持将生产环境中的失败链路转换为回归测试数据集,并利用基于语义的会话聚类来发现新兴的用户行为模式。 系统涵盖了广泛的功能领域,包括 Agent 执行路径的遥测采集、实时流量的自动化评估评分,以及用于隔离交互模式的语义搜索。此外,它还提供信号回归告警、工具故障行为分析以及遥测数据的 PII 脱敏功能。 该软件可通过 Docker Compose、Kubernetes 或 Helm charts 部署在私有基础设施上,支持单机安装或可扩展集群部署。
Implements a system for scoring live traffic and executing regression tests using automated semantic judgments.
PhiCookBook is a technical guide and implementation framework for integrating small language models into applications. It provides instructions for deploying these lightweight models to perform reasoning, coding, and math tasks across various hardware environments and serving platforms. The project functions as a tutorial for developing intelligent AI applications by chaining prompts and code into executable sequences. It includes a framework for evaluating model behavior and calculating quality metrics to verify the accuracy and reliability of these workflows. The repository covers a broad
Implements a framework to automate the assessment of AI outputs and reasoning quality through quality metrics and behavior testing.