awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

19 रिपॉजिटरी

Awesome GitHub RepositoriesAI Evaluation Frameworks

Systems that automate the assessment of artificial intelligence outputs and reasoning quality through comparative analysis or secondary model verification.

Explore 19 awesome GitHub repositories matching artificial intelligence & ml · AI Evaluation Frameworks. Refine with filters or upvote what's useful.

Awesome AI Evaluation Frameworks GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • foundationagents/metagptFoundationAgents का अवतार

    FoundationAgents/MetaGPT

    68,844GitHub पर देखें↗

    MetaGPT is an agentic workflow engine and multi-agent orchestration framework designed to automate complex software engineering and data analysis tasks. It functions as an automated software factory that transforms high-level natural language requirements into functional web applications, technical documentation, and production-ready code. By utilizing a runtime environment that manages the lifecycle of specialized agents, the platform bridges the gap between user intent and finished software components. The system distinguishes itself through role-based agent orchestration and dynamic task d

    Validates results by running multiple AI teams in parallel to compare outputs and test workflows for correctness and robustness.

    Pythonagentgptllm
    GitHub पर देखें↗68,844
  • langfuse/langfuselangfuse का अवतार

    langfuse/langfuse

    29,190GitHub पर देखें↗

    Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model

    Provides a system for running systematic experiments, benchmarking model outputs, and automating quality scoring.

    TypeScriptanalyticsautogenevaluation
    GitHub पर देखें↗29,190
  • datawhalechina/prompt-engineering-for-developersdatawhalechina का अवतार

    datawhalechina/prompt-engineering-for-developers

    24,267GitHub पर देखें↗

    This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap

    Ships a systematic approach for assessing generative AI outputs and reasoning quality through comparative analysis.

    Jupyter Notebook
    GitHub पर देखें↗24,267
  • openai/evalsopenai का अवतार

    openai/evals

    18,702GitHub पर देखें↗

    Evals is a framework designed for automating, managing, and executing repeatable benchmarking suites to analyze the quality and performance of language models. It provides a platform for running standardized tests to measure model accuracy and track behavioral changes over time. The system distinguishes itself through a modular architecture that uses a standardized adapter layer to normalize inputs and outputs, allowing different models to be swapped and tested interchangeably. It supports the creation of custom benchmarks using proprietary data, enabling quality assurance on sensitive tasks

    Enables the definition of bespoke evaluation logic and datasets to assess unique model behaviors.

    Python
    GitHub पर देखें↗18,702
  • comet-ml/opikcomet-ml का अवतार

    comet-ml/opik

    17,787GitHub पर देखें↗

    Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn

    Provides a comprehensive framework for automated testing, dataset management, and model-as-a-judge scoring.

    Pythonevaluationhacktoberfesthacktoberfest2025
    GitHub पर देखें↗17,787
  • camel-ai/camelcamel-ai का अवतार

    camel-ai/camel

    17,253GitHub पर देखें↗

    This project is a comprehensive framework for building and managing autonomous agent systems. It provides a unified architecture for orchestrating multi-agent societies, where specialized agents collaborate through roleplay to decompose and solve complex tasks. The system integrates language models with external environments, enabling agents to perform real-world actions through a standardized tool-calling abstraction layer. The framework distinguishes itself through its focus on iterative reasoning and data reliability. It employs automated feedback loops to refine agent outputs and self-eva

    Captures and structures the information returned by an agent to facilitate further reasoning or integration into a multi-agent workflow.

    Pythonagentai-societiesartificial-intelligence
    GitHub पर देखें↗17,253
  • ai-shifu/chatallai-shifu का अवतार

    ai-shifu/ChatALL

    16,283GitHub पर देखें↗

    ChatALL is a desktop application that functions as a multi-model chat client and aggregator for artificial intelligence services. It enables users to send a single prompt to multiple AI models simultaneously, allowing for the side-by-side comparison of generated responses within a unified interface. The application distinguishes itself through a local-first approach to data management, ensuring that all conversation logs and user configurations are stored directly on the user's device. This architecture supports privacy and offline access while providing a centralized system for managing and

    Facilitates comparative analysis by sending prompts to multiple AI services and displaying their responses side by side.

    JavaScriptbingchatchatbotchatgpt
    GitHub पर देखें↗16,283
  • chiphuyen/aie-bookchiphuyen का अवतार

    chiphuyen/aie-book

    13,779GitHub पर देखें↗

    This project serves as a comprehensive educational resource and technical handbook for engineers building applications powered by large language models. It provides a structured framework for mastering the principles of artificial intelligence engineering, covering the full lifecycle of model development from initial design to production deployment. The repository distinguishes itself by offering a deep dive into the practical implementation of advanced design patterns, including retrieval-augmented generation, agentic tool orchestration, and parameter-efficient model adaptation. It emphasize

    Automates the assessment of artificial intelligence outputs and reasoning quality through comparative analysis.

    Jupyter Notebook
    GitHub पर देखें↗13,779
  • googlecloudplatform/generative-aiGoogleCloudPlatform का अवतार

    GoogleCloudPlatform/generative-ai

    12,700GitHub पर देखें↗

    This project is a development platform for managing the lifecycle of generative artificial intelligence models. It provides a unified environment for accessing, fine-tuning, and deploying large language models, serving as an orchestrator that handles the integration of diverse models into custom applications. The platform distinguishes itself by offering a managed infrastructure for hosting and scaling models, which removes the requirement for manual server maintenance or configuration. It includes integrated tools for supervised fine-tuning and vector embedding optimization, allowing for the

    Assessing the performance and reliability of generated content using automated testing services and rubrics to ensure alignment with project requirements.

    Jupyter Notebookagentsgcpgemini
    GitHub पर देखें↗12,700
  • vibrantlabsai/ragasvibrantlabsai का अवतार

    vibrantlabsai/ragas

    12,659GitHub पर देखें↗

    Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin

    Compares outputs from multi-step AI workflows to verify correctness.

    Pythonevaluationllmllmops
    GitHub पर देखें↗12,659
  • aden-hive/hiveaden-hive का अवतार

    aden-hive/hive

    10,578GitHub पर देखें↗

    Hive is an artificial intelligence workflow automation engine and development platform designed for building and deploying autonomous agents. It provides a framework for orchestrating complex, multi-step business processes by coordinating tasks across multiple specialized agents using directed graph structures. The platform distinguishes itself through a focus on production-grade reliability and state management. It maintains persistent execution context and conversation history on disk, enabling crash recovery and continuity for long-running automated sessions. Furthermore, it incorporates a

    Validates agent outputs using a multi-level pipeline of deterministic rules, semantic assessment, and human oversight.

    Pythonagentagent-frameworkagent-skills
    GitHub पर देखें↗10,578
  • microsoft/vscode-copilot-chatmicrosoft का अवतार

    microsoft/vscode-copilot-chat

    9,493GitHub पर देखें↗

    This project is an AI-powered IDE extension and LLM coding assistant that provides a conversational interface for generating, refactoring, and debugging code. It functions as an AI agent framework and a Model Context Protocol client, connecting AI models to external data sources and tools to automate complex development tasks. The system is distinguished by its use of autonomous AI agents capable of multi-step task execution, including the ability to read files, modify code, and run terminal commands iteratively. It supports recursive agent orchestration through subagent delegation and employ

    Produces the necessary code and assets to implement evaluation frameworks for testing AI agents.

    TypeScript
    GitHub पर देखें↗9,493
  • oumi-ai/oumioumi-ai का अवतार

    oumi-ai/oumi

    8,858GitHub पर देखें↗

    Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo

    Uses an AI agent to suggest and define evaluator patterns based on a target task description.

    Pythondpoevaluationfine-tuning
    GitHub पर देखें↗8,858
  • evidentlyai/evidentlyevidentlyai का अवतार

    evidentlyai/evidently

    7,137GitHub पर देखें↗

    Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of

    Provides a comprehensive framework for automating the assessment of AI outputs and reasoning quality.

    Jupyter Notebookdata-driftdata-qualitydata-science
    GitHub पर देखें↗7,137
  • google/adk-gogoogle का अवतार

    google/adk-go

    6,958GitHub पर देखें↗

    adk-go is an agent orchestration engine and multi-agent framework for building, coordinating, and scaling systems of large language model agents. It provides a tool integration kit to connect agents with external APIs, custom functions, and diverse data sources. The project utilizes graph-based workflow orchestration to blend deterministic logic with adaptive reasoning. It supports modular multi-agent composition, allowing specialized agents to be organized into hierarchical structures to manage complex tasks through coordinated workflows. The framework includes tools for performance evaluat

    Measures agent reliability by comparing model responses against predefined schemas and validation tools.

    Goa2aagentsagents-sdk
    GitHub पर देखें↗6,958
  • fchollet/arc-agifchollet का अवतार

    fchollet/ARC-AGI

    4,787GitHub पर देखें↗

    यह प्रोजेक्ट कृत्रिम बुद्धिमत्ता मॉडल्स की नए नियम सीखने की क्षमता को बेंचमार्क करने के लिए डिज़ाइन की गई एब्स्ट्रैक्शन और रीजनिंग समस्याओं का एक स्टैंडर्ड सेट है। यह एक फ्लूइड इंटेलिजेंस टेस्ट और एक रीजनिंग बेंचमार्क के रूप में कार्य करता है, जो ग्रिड-आधारित पहेलियों के संग्रह और एक प्रोग्राम सिंथेसिस डेटासेट का उपयोग करके यह मूल्यांकन करता है कि एजेंट्स उदाहरणों से एल्गोरिदम कैसे उत्पन्न करते हैं। यह प्रोजेक्ट सामान्य फ्लूइड इंटेलिजेंस और ज़ीरो-शॉट सामान्यीकरण की क्षमता को मापने पर केंद्रित है, यह परीक्षण करता है कि क्या कोई सिस्टम कार्य-विशिष्ट प्रशिक्षण पर भरोसा किए बिना अनदेखी समस्याओं पर सीखे गए लॉजिक को लागू कर सकता है। यह एब्स्ट्रैक्ट रीजनिंग रिसर्च के लिए एक फ्रेमवर्क प्रदान करता है, जो विशेष रूप से प्रोग्राम सिंथेसिस और उपन्यास ग्रिड ट्रांसफॉर्मेशन के लिए लेटेंट पैटर्न्स खोजने की मॉडल्स की क्षमता का मूल्यांकन करता है। यह सिस्टम एक ग्रिड-आधारित डोमेन रिप्रेजेंटेशन और ज्यामितीय और टोपोलॉजिकल ऑपरेशन्स के एक कॉम्बिनेटरियल सर्च स्पेस को शामिल करता है। इसमें एक ह्यूमन-इन-द-लूप इंटरफेस शामिल है जो वैलिडेशन और बेंचमार्किंग के लिए ग्राउंड ट्रुथ को परिभाषित करने के लिए आउटपुट ग्रिड्स के मैन्युअल निर्माण की अनुमति देता है।

    Provides a standardized set of abstraction and reasoning problems to assess the quality of AI reasoning.

    JavaScriptartificial-intelligenceintelligence-testingprogram-synthesis
    GitHub पर देखें↗4,787
  • mongodb-developer/genai-showcasemongodb-developer का अवतार

    mongodb-developer/GenAI-Showcase

    4,236GitHub पर देखें↗

    यह प्रोजेक्ट जेनरेटिव AI कार्यान्वयनों का एक संग्रह है जो AI एजेंटों, रिट्रीवल-ऑगमेंटेड जनरेशन पाइपलाइन्स, और वेक्टर सर्च एकीकरण के विकास पर केंद्रित है। यह संदर्भ-जागरूक एप्लिकेशन बनाने के लिए प्रबंधित क्लाउड डेटाबेस को भाषा मॉडल्स से जोड़ने के लिए एक फ्रेमवर्क प्रदान करता है। यह प्रोजेक्ट स्वायत्त एजेंटों के ऑर्केस्ट्रेशन को कवर करता है जो कार्यों को पूरा करने के लिए मल्टी-स्टेप रीजनिंग और बाहरी टूल्स का उपयोग करते हैं। इसमें उच्च-आयामी एम्बेडिंग का उपयोग करके सिमेंटिक रिट्रीवल और विभिन्न लार्ज लैंग्वेज मॉडल्स में सुसंगत आउटपुट सुनिश्चित करने के लिए मॉडल-अज्ञेयवादी प्रॉम्प्टिंग का उपयोग शामिल है। अतिरिक्त क्षमताओं में AI प्रदर्शन की सटीकता और विश्वसनीयता को मापने के लिए ग्राउंड-ट्रुथ मूल्यांकन फ्रेमवर्क्स का उपयोग शामिल है। प्रोजेक्ट एप्लिकेशन डेटा और वेक्टर एम्बेडिंग्स को संग्रहीत और प्रबंधित करने के लिए क्लाउड डेटाबेस खातों के सेटअप को भी प्रदर्शित करता है।

    Uses AI evaluation frameworks to measure the accuracy and reliability of model outputs.

    Jupyter Notebookagentsartificial-intelligencegenerative-ai
    GitHub पर देखें↗4,236
  • latitude-dev/latitude-llmlatitude-dev का अवतार

    latitude-dev/latitude-llm

    4,145GitHub पर देखें↗

    This project is a self-hosted AI monitoring stack that functions as an LLM observability platform, AI evaluation framework, and OpenTelemetry trace analyzer. It is designed to capture and analyze LLM traces, sessions, and telemetry to monitor AI agent performance. The platform distinguishes itself as a Model Context Protocol server, exposing workspace functions as tools for AI coding agents. It enables the conversion of failing production traces into test datasets for regression testing and utilizes semantic-based session clustering to discover emerging user behavior patterns. The system cov

    Implements a system for scoring live traffic and executing regression tests using automated semantic judgments.

    TypeScript
    GitHub पर देखें↗4,145
  • microsoft/phicookbookmicrosoft का अवतार

    microsoft/PhiCookBook

    3,755GitHub पर देखें↗

    PhiCookBook is a technical guide and implementation framework for integrating small language models into applications. It provides instructions for deploying these lightweight models to perform reasoning, coding, and math tasks across various hardware environments and serving platforms. The project functions as a tutorial for developing intelligent AI applications by chaining prompts and code into executable sequences. It includes a framework for evaluating model behavior and calculating quality metrics to verify the accuracy and reliability of these workflows. The repository covers a broad

    Implements a framework to automate the assessment of AI outputs and reasoning quality through quality metrics and behavior testing.

    Jupyter Notebookcookbooklanguage-modelphi-4
    GitHub पर देखें↗3,755
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Model Evaluation and Analysis
  6. AI Evaluation Frameworks

सब-टैग एक्सप्लोर करें

  • Evaluation Asset Generation1 सब-टैगAutomatic generation of code and datasets required to implement AI evaluation frameworks. **Distinct from AI Evaluation Frameworks:** Focuses on producing the code that runs the evaluation, rather than the framework that performs the assessment.
  • Multi-Agent Output EvaluationSystems that compare outputs from parallel AI agents to verify correctness and workflow robustness.