19 مستودعات
Systems that automate the assessment of artificial intelligence outputs and reasoning quality through comparative analysis or secondary model verification.
Explore 19 awesome GitHub repositories matching artificial intelligence & ml · AI Evaluation Frameworks. Refine with filters or upvote what's useful.
MetaGPT is an agentic workflow engine and multi-agent orchestration framework designed to automate complex software engineering and data analysis tasks. It functions as an automated software factory that transforms high-level natural language requirements into functional web applications, technical documentation, and production-ready code. By utilizing a runtime environment that manages the lifecycle of specialized agents, the platform bridges the gap between user intent and finished software components. The system distinguishes itself through role-based agent orchestration and dynamic task d
Validates results by running multiple AI teams in parallel to compare outputs and test workflows for correctness and robustness.
Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model
Provides a system for running systematic experiments, benchmarking model outputs, and automating quality scoring.
This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap
Ships a systematic approach for assessing generative AI outputs and reasoning quality through comparative analysis.
Evals is a framework designed for automating, managing, and executing repeatable benchmarking suites to analyze the quality and performance of language models. It provides a platform for running standardized tests to measure model accuracy and track behavioral changes over time. The system distinguishes itself through a modular architecture that uses a standardized adapter layer to normalize inputs and outputs, allowing different models to be swapped and tested interchangeably. It supports the creation of custom benchmarks using proprietary data, enabling quality assurance on sensitive tasks
Enables the definition of bespoke evaluation logic and datasets to assess unique model behaviors.
Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn
Provides a comprehensive framework for automated testing, dataset management, and model-as-a-judge scoring.
This project is a comprehensive framework for building and managing autonomous agent systems. It provides a unified architecture for orchestrating multi-agent societies, where specialized agents collaborate through roleplay to decompose and solve complex tasks. The system integrates language models with external environments, enabling agents to perform real-world actions through a standardized tool-calling abstraction layer. The framework distinguishes itself through its focus on iterative reasoning and data reliability. It employs automated feedback loops to refine agent outputs and self-eva
Captures and structures the information returned by an agent to facilitate further reasoning or integration into a multi-agent workflow.
ChatALL is a desktop application that functions as a multi-model chat client and aggregator for artificial intelligence services. It enables users to send a single prompt to multiple AI models simultaneously, allowing for the side-by-side comparison of generated responses within a unified interface. The application distinguishes itself through a local-first approach to data management, ensuring that all conversation logs and user configurations are stored directly on the user's device. This architecture supports privacy and offline access while providing a centralized system for managing and
Facilitates comparative analysis by sending prompts to multiple AI services and displaying their responses side by side.
This project serves as a comprehensive educational resource and technical handbook for engineers building applications powered by large language models. It provides a structured framework for mastering the principles of artificial intelligence engineering, covering the full lifecycle of model development from initial design to production deployment. The repository distinguishes itself by offering a deep dive into the practical implementation of advanced design patterns, including retrieval-augmented generation, agentic tool orchestration, and parameter-efficient model adaptation. It emphasize
Automates the assessment of artificial intelligence outputs and reasoning quality through comparative analysis.
This project is a development platform for managing the lifecycle of generative artificial intelligence models. It provides a unified environment for accessing, fine-tuning, and deploying large language models, serving as an orchestrator that handles the integration of diverse models into custom applications. The platform distinguishes itself by offering a managed infrastructure for hosting and scaling models, which removes the requirement for manual server maintenance or configuration. It includes integrated tools for supervised fine-tuning and vector embedding optimization, allowing for the
Assessing the performance and reliability of generated content using automated testing services and rubrics to ensure alignment with project requirements.
Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin
Compares outputs from multi-step AI workflows to verify correctness.
Hive is an artificial intelligence workflow automation engine and development platform designed for building and deploying autonomous agents. It provides a framework for orchestrating complex, multi-step business processes by coordinating tasks across multiple specialized agents using directed graph structures. The platform distinguishes itself through a focus on production-grade reliability and state management. It maintains persistent execution context and conversation history on disk, enabling crash recovery and continuity for long-running automated sessions. Furthermore, it incorporates a
Validates agent outputs using a multi-level pipeline of deterministic rules, semantic assessment, and human oversight.
This project is an AI-powered IDE extension and LLM coding assistant that provides a conversational interface for generating, refactoring, and debugging code. It functions as an AI agent framework and a Model Context Protocol client, connecting AI models to external data sources and tools to automate complex development tasks. The system is distinguished by its use of autonomous AI agents capable of multi-step task execution, including the ability to read files, modify code, and run terminal commands iteratively. It supports recursive agent orchestration through subagent delegation and employ
Produces the necessary code and assets to implement evaluation frameworks for testing AI agents.
Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo
Uses an AI agent to suggest and define evaluator patterns based on a target task description.
Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of
Provides a comprehensive framework for automating the assessment of AI outputs and reasoning quality.
adk-go is an agent orchestration engine and multi-agent framework for building, coordinating, and scaling systems of large language model agents. It provides a tool integration kit to connect agents with external APIs, custom functions, and diverse data sources. The project utilizes graph-based workflow orchestration to blend deterministic logic with adaptive reasoning. It supports modular multi-agent composition, allowing specialized agents to be organized into hierarchical structures to manage complex tasks through coordinated workflows. The framework includes tools for performance evaluat
Measures agent reliability by comparing model responses against predefined schemas and validation tools.
هذا المشروع هو مجموعة قياسية من مشاكل التجريد والاستدلال المصممة لقياس قدرة نماذج الذكاء الاصطناعي على تعلم قواعد جديدة. يعمل كاختبار للذكاء السائل ومعيار للاستدلال، باستخدام مجموعة من الألغاز القائمة على الشبكة ومجموعة بيانات لتوليد البرامج لتقييم كيفية توليد الوكلاء للخوارزميات من الأمثلة. يركز المشروع على قياس الذكاء السائل العام والقدرة على التعميم بدون تدريب مسبق (zero-shot)، واختبار ما إذا كان النظام يمكنه تطبيق المنطق المتعلم على مشاكل غير مرئية دون الاعتماد على تدريب خاص بالمهمة. يوفر إطار عمل لبحوث الاستدلال التجريدي، وتقييم توليد البرامج بشكل خاص وقدرة النماذج على اكتشاف الأنماط الكامنة لتحويلات الشبكة الجديدة. يدمج النظام تمثيلاً للمجال قائماً على الشبكة ومساحة بحث توافقية للعمليات الهندسية والطوبولوجية. يتضمن واجهة إنسان في الحلقة تسمح بالبناء اليدوي لشبكات المخرجات لتحديد الحقيقة الأرضية للتحقق والمقارنة.
Provides a standardized set of abstraction and reasoning problems to assess the quality of AI reasoning.
هذا المشروع عبارة عن مجموعة من تنفيذات الذكاء الاصطناعي التوليدي التي تركز على تطوير وكلاء الذكاء الاصطناعي، وخطوط أنابيب التوليد المعزز بالاسترجاع (RAG)، وتكامل البحث المتجهي. يوفر إطار عمل لربط قواعد بيانات سحابية مدارة بنماذج اللغة لإنشاء تطبيقات واعية بالسياق. يغطي المشروع تنسيق الوكلاء المستقلين الذين يستخدمون التفكير متعدد الخطوات وأدوات خارجية لإكمال المهام. يتضمن تنفيذات للاسترجاع الدلالي باستخدام تضمينات عالية الأبعاد واستخدام توجيهات (prompting) مستقلة عن النموذج لضمان مخرجات متسقة عبر نماذج لغة كبيرة مختلفة. تشمل القدرات الإضافية استخدام أطر تقييم الحقيقة الأساسية (ground-truth) لقياس دقة وموثوقية أداء الذكاء الاصطناعي. يوضح المشروع أيضاً إعداد حسابات قواعد البيانات السحابية لتخزين وإدارة بيانات التطبيقات والتضمينات المتجهية.
Uses AI evaluation frameworks to measure the accuracy and reliability of model outputs.
هذا المشروع عبارة عن حزمة مراقبة ذكاء اصطناعي مستضافة ذاتياً تعمل كمنصة لمراقبة LLM، وإطار عمل لتقييم الذكاء الاصطناعي، ومحلل تتبع OpenTelemetry. صُمم لالتقاط وتحليل تتبعات LLM والجلسات والقياسات عن بُعد لمراقبة أداء وكيل الذكاء الاصطناعي. تتميز المنصة كخادم لبروتوكول سياق النموذج (Model Context Protocol)، حيث تعرض وظائف مساحة العمل كأدوات لوكلاء برمجة الذكاء الاصطناعي. تمكن من تحويل تتبعات الإنتاج الفاشلة إلى مجموعات بيانات اختبار لاختبار الانحدار وتستخدم تجميع الجلسات القائم على الدلالات لاكتشاف أنماط سلوك المستخدم الناشئة. يغطي النظام مجالات قدرات واسعة بما في ذلك جمع القياسات عن بُعد لمسارات تنفيذ الوكيل، وتسجيل التقييم التلقائي لحركة المرور المباشرة، والبحث الدلالي لعزل أنماط التفاعل. كما يوفر تنبيهات لانحدارات الإشارة، وتحليل السلوك لفشل الأدوات، وتصحيح PII لبيانات القياس عن بُعد. يمكن نشر البرنامج على بنية تحتية خاصة كتثبيت لمضيف واحد أو كمجموعة قابلة للتوسع باستخدام Docker Compose أو Kubernetes أو Helm charts.
Implements a system for scoring live traffic and executing regression tests using automated semantic judgments.
PhiCookBook is a technical guide and implementation framework for integrating small language models into applications. It provides instructions for deploying these lightweight models to perform reasoning, coding, and math tasks across various hardware environments and serving platforms. The project functions as a tutorial for developing intelligent AI applications by chaining prompts and code into executable sequences. It includes a framework for evaluating model behavior and calculating quality metrics to verify the accuracy and reliability of these workflows. The repository covers a broad
Implements a framework to automate the assessment of AI outputs and reasoning quality through quality metrics and behavior testing.