awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

33 dépôts

Awesome GitHub RepositoriesAI Observability and Evaluation

Diagnostic and benchmarking tools designed to inspect model reasoning, trace execution flows, and validate performance metrics.

Explore 33 awesome GitHub repositories matching artificial intelligence & ml · AI Observability and Evaluation. Refine with filters or upvote what's useful.

Awesome AI Observability and Evaluation GitHub Repositories

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • ripienaar/free-for-devAvatar de ripienaar

    ripienaar/free-for-dev

    123,154Voir sur GitHub↗

    This project is a community-maintained directory of technical resources, tools, and services that offer free tiers for developers. It serves as a centralized reference point for discovering infrastructure, software, and educational materials, helping individuals and teams minimize operational costs while building and scaling applications. The directory distinguishes itself through a collaborative, community-driven curation model that aggregates metadata about third-party services. By utilizing a hierarchical taxonomy and storing all content in version-controlled, plain-text files, the project

    Track telemetry data to inspect the performance, latency, and output of artificial intelligence models and agents in production environments.

    HTMLawesome-listfree-for-developers
    Voir sur GitHub↗123,154
  • shubhamsaboo/awesome-llm-appsAvatar de Shubhamsaboo

    Shubhamsaboo/awesome-llm-apps

    114,725Voir sur GitHub↗

    This repository serves as a comprehensive collection of resources, templates, and starter code for building artificial intelligence applications. It provides a centralized hub for developers to access practical implementations of common workflows, including retrieval-augmented generation pipelines and autonomous agent loops, alongside educational materials designed to support rapid prototyping and experimentation. The project distinguishes itself by offering a dual focus on technical implementation and critical analysis. It provides a library of lightweight, single-file agents and tutorials f

    Assesses stylistic consistency and linguistic traits to help verify machine-generated output.

    Pythonagentsllmspython
    Voir sur GitHub↗114,725
  • garrytan/gstackAvatar de garrytan

    garrytan/gstack

    110,596Voir sur GitHub↗

    gstack is an AI agent framework and development workflow system designed to automate the software development lifecycle. It coordinates specialized AI personas to manage tasks across product design, engineering management, and quality assurance, transforming product intent into technical specifications and final releases. The project is distinguished by its deep integration of headless browser automation and semantic code memory. It utilizes a persistent Chromium daemon for web scraping and visual auditing, and implements a searchable knowledge base that logs architectural decisions and repos

    Compares latency, token usage, and quality scores by running identical prompts across different AI providers.

    TypeScript
    Voir sur GitHub↗110,596
  • jaywcjlove/awesome-macAvatar de jaywcjlove

    jaywcjlove/awesome-mac

    105,841Voir sur GitHub↗

    This project is a comprehensive, curated collection of software resources designed for the macOS ecosystem. It serves as a centralized directory for discovering applications across a wide range of functional domains, including professional development, system management, and personal productivity. The directory distinguishes itself by offering a highly granular classification of tools that cater to specific technical and creative workflows. It highlights specialized software for software engineering, such as terminal emulators, version control clients, and API development tools, alongside a b

    Monitors performance metrics and output latency for artificial intelligence agents.

    Swiftappappleapplication
    Voir sur GitHub↗105,841
  • fighting41love/funnlpAvatar de fighting41love

    fighting41love/funNLP

    81,299Voir sur GitHub↗

    This project is a community-driven knowledge base and curated repository focused on natural language processing and large language model development. It serves as a centralized index for high-quality tools, libraries, and research materials, organizing technical resources into structured, version-controlled documentation to assist developers in navigating the evolving artificial intelligence ecosystem. The repository distinguishes itself by acting as an aggregator for AI model evaluation and benchmarking. It provides access to tools that enable the simultaneous comparison of multiple conversa

    Curates methodologies and frameworks for assessing the performance and output quality of various conversational artificial intelligence models.

    Python
    Voir sur GitHub↗81,299
  • openhands/openhandsAvatar de OpenHands

    OpenHands/OpenHands

    77,330Voir sur GitHub↗

    OpenHands is an autonomous agent framework designed for software engineering workflows. It provides a modular platform for orchestrating AI agents that reason, plan, and execute tasks within isolated, containerized development environments. By integrating with standard version control and development tools, the system enables agents to autonomously navigate codebases, implement features, and resolve issues through iterative reasoning and tool execution. The platform distinguishes itself through a model-agnostic orchestrator that connects diverse language models to a unified tool registry. It

    Intercepts internal thinking blocks through event callbacks to audit and visualize the step-by-step reasoning chains used by models.

    Pythonagentartificial-intelligencechatgpt
    Voir sur GitHub↗77,330
  • microsoft/ai-agents-for-beginnersAvatar de microsoft

    microsoft/ai-agents-for-beginners

    67,369Voir sur GitHub↗

    This project is a structured educational resource and technical guide for designing and implementing autonomous systems using large language models. It provides a comprehensive curriculum and code samples focused on agentic design patterns, autonomous development, and the creation of systems capable of planning and executing multi-step tasks. The resource details the implementation of agentic retrieval-augmented generation, where models autonomously plan and refine data searches. It covers a wide array of orchestrators and design patterns, including metacognitive reflection for self-correctin

    Ships tools to visualize and audit step-by-step reasoning chains, enabling the evaluation and adjustment of internal decision processes.

    Jupyter Notebookagentic-aiagentic-frameworkagentic-rag
    Voir sur GitHub↗67,369
  • patchy631/ai-engineering-hubAvatar de patchy631

    patchy631/ai-engineering-hub

    35,826Voir sur GitHub↗

    This project serves as an educational resource and technical guide for building production-ready intelligent systems. It provides a collection of hands-on tutorials, blueprints, and documentation focused on the development of applications powered by large language models, autonomous agentic workflows, and retrieval-augmented generation. The repository distinguishes itself by offering structured implementations for multi-agent orchestration and standardized communication protocols. It enables developers to integrate external tools and data sources into their systems, ensuring interoperability

    Instruments system outputs with monitoring hooks to benchmark performance, track reliability, and validate model behavior.

    Jupyter Notebookagentsaillms
    Voir sur GitHub↗35,826
  • sgl-project/sglangAvatar de sgl-project

    sgl-project/sglang

    29,079Voir sur GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Surfaces internal reasoning steps within API responses using unified configuration parameters.

    Pythonattentionblackwellcuda
    Voir sur GitHub↗29,079
  • virattt/dexterAvatar de virattt

    virattt/dexter

    27,085Voir sur GitHub↗

    Dexter is an autonomous research platform designed to decompose complex inquiries into structured, multi-step workflows. It functions as an agent orchestration system that utilizes iterative tool-calling loops and language models to gather data, perform analysis, and validate findings against internal criteria to ensure accuracy. The platform distinguishes itself through its specialized focus on financial research and messaging integration. It autonomously interprets real-time market data, including income statements and regulatory filings, to generate evidence-based insights. By connecting d

    Logs tool calls, raw data, and internal thought processes to provide transparency into how conclusions are reached.

    TypeScript
    Voir sur GitHub↗27,085
  • datawhalechina/llm-cookbookAvatar de datawhalechina

    datawhalechina/llm-cookbook

    24,263Voir sur GitHub↗

    This repository is a comprehensive set of tutorials and examples for building software powered by large language models. It serves as an application development guide and a prompt engineering framework, providing instructional content for integrating model logic with user interfaces and external data sources. The project provides technical walkthroughs for specialized workflows, including the implementation of retrieval augmented generation using vector databases and semantic search. It includes guidance on adapting pre-trained model weights through fine-tuning with private datasets and the o

    Provides systematic tools for tracking and debugging AI outputs to ensure quality and consistency.

    Jupyter Notebookcookbookllm
    Voir sur GitHub↗24,263
  • liguodongiot/llm-actionAvatar de liguodongiot

    liguodongiot/llm-action

    23,169Voir sur GitHub↗

    This project is a comprehensive framework for the training, fine-tuning, and deployment of large language models. It functions as a distributed deep learning platform that enables users to scale model workflows across multiple hardware nodes while providing tools for model evaluation and performance benchmarking. The platform distinguishes itself by offering specialized utilities for model compression and weight transformation, allowing users to reduce memory footprints and latency through quantization and pruning. It supports the adaptation of large models for consumer-grade hardware, facili

    Evaluates language model capabilities through standardized reasoning and instruction-following benchmarks.

    HTMLllmllm-inferencellm-serving
    Voir sur GitHub↗23,169
  • typpo/promptfooAvatar de typpo

    typpo/promptfoo

    22,295Voir sur GitHub↗

    promptfoo is an evaluation framework for measuring the performance of large language model prompts, agents, and retrieval augmented generation pipelines. It provides a suite of tools for conducting comparative benchmarking and executing automated quality and security regressions. The system features a benchmarking suite for running identical prompts across different model providers to compare output quality side-by-side. It also includes a dedicated red teaming tool for identifying security vulnerabilities and prompt injection risks through automated penetration testing. The framework suppor

    Provides frameworks for running standardized tests to compare the performance and reliability of different LLM providers.

    TypeScript
    Voir sur GitHub↗22,295
  • huggingface/trlAvatar de huggingface

    huggingface/trl

    18,653Voir sur GitHub↗

    This library provides a comprehensive framework for fine-tuning, aligning, and distilling transformer-based language models. It serves as a toolkit for adapting models to specialized domains through supervised learning, while offering advanced methodologies to improve output quality and reasoning capabilities. The project distinguishes itself through specialized alignment and optimization techniques, including direct preference optimization and reinforcement learning, which allow models to be tuned against human preferences without complex reward modeling. It further supports training efficie

    Tracks evaluation traces and model predictions to validate performance and reasoning quality.

    Python
    Voir sur GitHub↗18,653
  • pydantic/pydantic-aiAvatar de pydantic

    pydantic/pydantic-ai

    17,791Voir sur GitHub↗

    PydanticAI is a Python framework designed for building production-grade autonomous agents. It provides a unified interface for interacting with diverse language models, enabling developers to construct agents that perform complex tasks through structured data validation, tool execution, and multi-turn conversation management. The library centers on type-safe schema enforcement, ensuring that model inputs and outputs remain consistent and reliable throughout the agent's lifecycle. The framework distinguishes itself through a robust architecture that emphasizes modularity and testability. It ut

    The framework runs tasks against defined datasets to measure performance and generate comprehensive reports on model behavior and accuracy.

    Pythonagent-frameworkgenaillm
    Voir sur GitHub↗17,791
  • anthropics/claude-quickstartsAvatar de anthropics

    anthropics/claude-quickstarts

    17,085Voir sur GitHub↗

    Claude Quickstarts is a development framework and collection of reference implementations designed for building autonomous agents. It provides the foundational patterns necessary to orchestrate multi-agent workflows, enabling models to perform complex, multi-step tasks across software engineering, customer support, and computer-use domains. The platform distinguishes itself through specialized capabilities for desktop and browser automation, allowing agents to interact with graphical interfaces by capturing visual context and executing precise mouse and keyboard inputs. It includes robust inf

    Visualizes real-time thinking steps and source citations to provide transparency into agent decision-making.

    Python
    Voir sur GitHub↗17,085
  • tencent/weknoraAvatar de Tencent

    Tencent/WeKnora

    16,974Voir sur GitHub↗

    WeKnora is a multi-tenant retrieval-augmented generation (RAG) knowledge platform and autonomous AI agent framework. It transforms raw documents into queryable knowledge bases and integrates large language models with vector databases to provide grounded AI responses. The system also functions as a Model Context Protocol (MCP) tool server, exposing knowledge search and agentic capabilities to external AI clients. The platform distinguishes itself through an autonomous agent framework that utilizes iterative reasoning, tool calling, and web search to solve multi-step tasks. It implements a sta

    Provides visualization and auditing of step-by-step reasoning chains and tool execution for token observability.

    Goagentagenticai
    Voir sur GitHub↗16,974
  • langbot-app/langbotAvatar de langbot-app

    langbot-app/LangBot

    15,311Voir sur GitHub↗

    LangBot is an orchestration platform designed for building, managing, and deploying AI agents. It functions as a comprehensive framework for integrating large language models with custom workflows, enabling developers to connect intelligent agents to various messaging platforms and external tools. The platform distinguishes itself through a modular, plugin-based architecture that allows for the extension of agent capabilities via custom tools and file parsers. It features a secure, sandbox-isolated runtime environment that executes untrusted code and plugin logic within resource-constrained c

    Formats and displays agent reasoning steps, tool execution logs, and knowledge base citations.

    Pythonagentcozedeepseek
    Voir sur GitHub↗15,311
  • google-ai-edge/galleryAvatar de google-ai-edge

    google-ai-edge/gallery

    15,162Voir sur GitHub↗

    This project is a development framework for building edge-based AI agents that perform multimodal inference and system-level automation directly on mobile devices. By prioritizing local-first execution, the platform ensures data privacy and offline functionality, allowing developers to run large language models on hardware without requiring external server connectivity. The framework distinguishes itself through an integrated orchestration layer that connects language models to custom tools, scripts, and native device intents. It provides a structured registry for mapping natural language ins

    Visualizes and audits the step-by-step reasoning chains used by models to provide transparency into complex problem solving.

    Kotlin
    Voir sur GitHub↗15,162
  • deepmind/deepmind-researchAvatar de deepmind

    deepmind/deepmind-research

    15,024Voir sur GitHub↗

    This project is an AI research implementation library and machine learning research repository. It provides a collection of reference code, illustrative implementations, and open-source research datasets used to verify hypotheses and build upon existing models in artificial intelligence. The repository focuses on scientific research reproduction by translating theoretical findings from published papers into executable code. It includes specialized scientific simulation environments designed to test the behavior of autonomous agents and models within controlled settings. The project covers AI

    Utilizes standardized datasets and environments to benchmark the performance of new machine learning models.

    Jupyter Notebook
    Voir sur GitHub↗15,024
Préc.12Suivant
  1. Home
  2. Artificial Intelligence & ML
  3. Artificial Intelligence Tooling
  4. AI Observability and Evaluation

Explorer les sous-tags

  • AI Content Analysis ToolsTools that evaluate the quality, safety, and sentiment of content generated or processed by artificial intelligence.
  • AI Model BenchmarkingFrameworks and services for running standardized tests to assess the performance and reliability of machine learning models.
  • AI Monitoring ToolsUtilities for monitoring, inspecting, and analyzing the performance, latency, and output of artificial intelligence applications.
  • Listing Evaluation1 sous-tagAutomated assessment of web listings using AI to determine if products meet specific criteria. **Distinct from AI Observability and Evaluation:** Focuses on evaluating external product content rather than benchmarking the model's internal performance or reasoning.
  • Reasoning Process MonitorsTools that visualize and audit the step-by-step reasoning chains used by models to reach conclusions.