awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

94 مستودعات

Awesome GitHub RepositoriesEvaluation Datasets

Structured collections of inputs and expected outputs used for benchmarking language model performance.

Distinct from Dataset Management: Distinct from general dataset management: focuses on evaluation-specific benchmarking inputs rather than training data samples.

Explore 94 awesome GitHub repositories matching artificial intelligence & ml · Evaluation Datasets. Refine with filters or upvote what's useful.

Awesome Evaluation Datasets GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • datawhalechina/hello-agentsالصورة الرمزية لـ datawhalechina

    datawhalechina/hello-agents

    59,685عرض على GitHub↗

    This project provides a comprehensive framework for building, training, and managing autonomous agents. It enables the construction of systems that utilize language models to plan, manage memory, and execute multi-step tasks through iterative reasoning loops and tool-based actions. The framework distinguishes itself by offering specialized capabilities for interacting with graphical user interfaces and legacy software, allowing agents to perceive visual elements and perform actions like a human user. It supports complex, cross-application workflows through graph-based orchestration and provid

    Maintains fixed request signatures and evaluation contexts to ensure model performance improvements are measurable and comparable across different training iterations.

    Pythonagentllmrag
    عرض على GitHub↗59,685
  • roboflow/supervisionالصورة الرمزية لـ roboflow

    roboflow/supervision

    44,437عرض على GitHub↗

    Supervision is a computer vision toolset for normalizing model outputs, managing datasets, and visualizing annotations. It provides a framework to convert predictions from various classification and detection models into a standardized data format to ensure interoperability across different computer vision pipelines. The library features a post-processor for filtering, counting, and tracking detected objects across image frames and video streams. It includes capabilities for large image tiling to improve the detection of small objects and tools for assigning persistent identities to objects t

    Converts computer vision datasets between different formats to ensure compatibility between training and evaluation frameworks.

    Pythonclassificationcococomputer-vision
    عرض على GitHub↗44,437
  • langchain-ai/deepagentsالصورة الرمزية لـ langchain-ai

    langchain-ai/deepagents

    25,006عرض على GitHub↗

    Deepagents is an LLM agent orchestration platform and stateful application server designed for deploying and managing AI agents built with computational graphs. It provides a containerized runtime environment that handles agent execution, state persistence, and the versioning of AI assistants. The platform distinguishes itself through deep integration with the Model Context Protocol, allowing agents to function as servers that expose tools and capabilities to external clients. It features a sophisticated observability suite for capturing execution traces, performing LLM-based evaluations agai

    Runs defined evaluators against dataset examples to measure agent performance automatically.

    Pythonagentsdeepagentslangchain
    عرض على GitHub↗25,006
  • recommenders-team/recommendersالصورة الرمزية لـ recommenders-team

    recommenders-team/recommenders

    21,769عرض على GitHub↗

    This project is a recommendation system framework designed for building, evaluating, and operationalizing personalized item suggestion engines. It provides a comprehensive toolkit for implementing collaborative filtering and content-based algorithms, supported by an end-to-end machine learning pipeline for preparing datasets and deploying predictive models. The framework distinguishes itself through the integration of knowledge graphs to provide richer context for recommendations and the use of industry-specific patterns to accelerate system deployment. It also includes a specialized model ev

    Provides stratified dataset splitting to ensure representative model evaluation.

    Pythonaiartificial-intelligencedata-science
    عرض على GitHub↗21,769
  • mastra-ai/mastraالصورة الرمزية لـ mastra-ai

    mastra-ai/mastra

    21,221عرض على GitHub↗

    Mastra is an orchestration framework designed for building, deploying, and managing autonomous AI agents and multi-agent systems. It provides a comprehensive suite of primitives for creating resilient AI applications, including durable workflow orchestration, event-driven agent loops, and semantic memory management. By integrating these core components, the platform enables developers to build complex, multi-step processes that can reason about goals and execute tasks without manual intervention. The framework distinguishes itself through its focus on observability and secure, isolated execut

    Organizes structured collections of inputs and expected outputs for benchmarking agent and workflow performance.

    TypeScriptagentsaichatbots
    عرض على GitHub↗21,221
  • letta-ai/lettaالصورة الرمزية لـ letta-ai

    letta-ai/letta

    21,168عرض على GitHub↗

    Letta is a framework for building, deploying, and managing autonomous AI agents that maintain persistent state across long-term interactions. It provides a comprehensive suite of primitives for defining agents with configurable personas, modular memory blocks, and tool-use capabilities, enabling them to retain user preferences and conversation history over extended sessions. The platform distinguishes itself through its advanced memory management and orchestration capabilities. It allows agents to autonomously update their own memory, perform retrieval-augmented generation, and coordinate com

    Organizes collections of test cases in JSONL or CSV formats to systematically measure agent performance.

    Pythonaiai-agentsllm
    عرض على GitHub↗21,168
  • pydantic/pydantic-aiالصورة الرمزية لـ pydantic

    pydantic/pydantic-ai

    17,791عرض على GitHub↗

    PydanticAI is a Python framework designed for building production-grade autonomous agents. It provides a unified interface for interacting with diverse language models, enabling developers to construct agents that perform complex tasks through structured data validation, tool execution, and multi-turn conversation management. The library centers on type-safe schema enforcement, ensuring that model inputs and outputs remain consistent and reliable throughout the agent's lifecycle. The framework distinguishes itself through a robust architecture that emphasizes modularity and testability. It ut

    The framework organizes collections of test scenarios and expected outcomes into structured datasets to systematically validate AI tasks and functions.

    Pythonagent-frameworkgenaillm
    عرض على GitHub↗17,791
  • wkentaro/labelmeالصورة الرمزية لـ wkentaro

    wkentaro/labelme

    15,984عرض على GitHub↗

    Labelme هي أداة تعليق صور تعتمد على Python تُستخدم لإنشاء مجموعات بيانات الرؤية الحاسوبية. تعمل كمحرر مرئي للتقسيم الدلالي، مما يسمح للمستخدمين بتحديد حدود الكائنات باستخدام المضلعات والمستطيلات والنقاط والدوائر. يعمل التطبيق أيضاً كمعلق صور متعدد الأطياف، ويدعم ملفات TIFF ذات العمق العالي المستخدمة في صور الأقمار الصناعية والصور العلمية. تتضمن الأداة قدرات تصنيف بمساعدة الذكاء الاصطناعي لأتمتة إنشاء الأقنعة والمضلعات. تسمح هذه الميزات بإنشاء الأشكال المدفوع بمطالبات نصية أو اختيارات نقاط تفاعلية، والتي تقترح حدوداً بناءً على نقاط إيجابية وسلبية يضعها المستخدم. يغطي البرنامج مجموعة واسعة من مهام إدارة البيانات والتعليق، بما في ذلك إنشاء أقنعة بكسل كثيفة، وصناديق التحديد الدوارة، وتسلسل إطارات الفيديو. يتضمن خط أنابيب لترجمة استمرارية حالة JSON الداخلية إلى تنسيقات مجموعات بيانات قياسية مثل COCO و Pascal VOC. تشمل القدرات الإضافية علامات تصنيف على مستوى الصورة، وأدوات تحسين الهندسة، واستيراد الصور المجمعة.

    Converts custom annotations into standardized vision formats like VOC or COCO for consistent model training.

    Python
    عرض على GitHub↗15,984
  • confident-ai/deepevalالصورة الرمزية لـ confident-ai

    confident-ai/deepeval

    13,733عرض على GitHub↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Manages structured evaluation datasets to ensure consistent benchmarking across model versions and prompt iterations.

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    عرض على GitHub↗13,733
  • rasbt/python-machine-learning-bookالصورة الرمزية لـ rasbt

    rasbt/python-machine-learning-book

    12,614عرض على GitHub↗

    This project is an educational resource providing practical code examples and implementations of machine learning algorithms using the Python language. It serves as a guide for constructing predictive pipelines, clustering models, and dimensionality reduction within the Scikit-Learn ecosystem. The repository includes comprehensive demonstrations for supervised and unsupervised learning, as well as detailed examples for implementing neural networks and deep architectures. It also provides practical guidance on exporting model parameters to JSON and wrapping trained models in web APIs for produ

    Provides utilities for partitioning datasets into separate training and testing sets to evaluate generalization.

    Jupyter Notebook
    عرض على GitHub↗12,614
  • vibrantlabsai/ragasالصورة الرمزية لـ vibrantlabsai

    vibrantlabsai/ragas

    12,659عرض على GitHub↗

    Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin

    Provides structured collections of inputs and expected outputs used for benchmarking language model performance.

    Pythonevaluationllmllmops
    عرض على GitHub↗12,659
  • salesforce/lavisالصورة الرمزية لـ salesforce

    salesforce/LAVIS

    11,236عرض على GitHub↗

    LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and

    Imports and prepares multimodal datasets for use in training and evaluating language-vision models.

    Jupyter Notebook
    عرض على GitHub↗11,236
  • facebookresearch/parlaiالصورة الرمزية لـ facebookresearch

    facebookresearch/ParlAI

    10,625عرض على GitHub↗

    ParlAI is a conversational AI research framework designed for training, evaluating, and sharing dialogue models using a unified interface for datasets and agents. It functions as a PyTorch-based training platform and a dialogue data collection system, providing a centralized model zoo for the distribution of versioned pretrained agents. The project distinguishes itself through a knowledge-grounded retrieval system that combines dense and sparse indexing to ground responses in external information. It also provides a comprehensive infrastructure for gathering human-AI interaction data via inte

    Computes standard performance metrics on a held-out dataset to measure model quality after training.

    Python
    عرض على GitHub↗10,625
  • raulmur/orb_slam2الصورة الرمزية لـ raulmur

    raulmur/ORB_SLAM2

    10,105عرض على GitHub↗

    ORB_SLAM2 is a visual simultaneous localization and mapping system that tracks camera movement and builds 3D environments from image data. It functions as a real-time visual odometry tool and sparse 3D reconstructor, computing the position and orientation of a camera while generating a point cloud map of a physical space. The system utilizes a camera relocalization engine to identify a camera's position within a known map after tracking failure or system restarts. It incorporates a spatial tracker to enable the precise insertion and composition of virtual 3D objects into real-world planar reg

    Provides utilities to analyze image sequence datasets for measuring the accuracy and performance of spatial mapping algorithms.

    C++
    عرض على GitHub↗10,105
  • joelgrus/data-science-from-scratchالصورة الرمزية لـ joelgrus

    joelgrus/data-science-from-scratch

    9,636عرض على GitHub↗

    This project is a collection of foundational machine learning algorithms and data science tools implemented in Python. It focuses on building the logic of these tools using basic programming primitives rather than relying on specialized libraries. The implementation covers several core domains, including a linear algebra library for matrix and vector operations, a statistical analysis toolkit for probability and hypothesis testing, and a framework for map-reduce distributed processing. It also includes implementations for natural language processing, graph theory for network analysis, and var

    Provides utilities for dividing datasets into training and testing sets to evaluate model performance.

    Python
    عرض على GitHub↗9,636
  • sjwhitworth/golearnالصورة الرمزية لـ sjwhitworth

    sjwhitworth/golearn

    9,438عرض على GitHub↗

    GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and a toolkit for building, training, and evaluating predictive models through a standardized interface. The project implements a data frame system that loads CSV files into structured grids for matrix operations. It includes a preprocessing library for discretizing continuous variables and a model evaluation toolkit that utilizes confusion matrices and cross-validation to measure precision and recall. The library covers data engineering and management, including the ability to

    Ships a toolkit for measuring model accuracy and precision via confusion matrices and cross-validation.

    Go
    عرض على GitHub↗9,438
  • facebookresearch/maskrcnn-benchmarkالصورة الرمزية لـ facebookresearch

    facebookresearch/maskrcnn-benchmark

    9,370عرض على GitHub↗

    This project is a modular PyTorch framework for training and evaluating object detection and instance segmentation models. It serves as a computer vision research tool and a deep learning inference engine designed to identify object locations, classes, and pixel-level masks within images. The framework implements a two-stage inference pipeline that utilizes region proposal networks and a symmetric mask-head architecture. It provides specialized capabilities for instance segmentation, object bounding box detection, and human pose estimation via anatomical keypoint detection. The system includ

    Includes utilities to standardize diverse vision dataset annotations into a uniform format for batch processing.

    Python
    عرض على GitHub↗9,370
  • redditsota/state-of-the-art-result-for-machine-learning-problemsالصورة الرمزية لـ RedditSota

    RedditSota/state-of-the-art-result-for-machine-learning-problems

    8,900عرض على GitHub↗

    This project is a machine learning benchmark directory and performance registry. It serves as a centralized catalog of state-of-the-art results, research papers, and associated source code across multiple artificial intelligence domains. The repository functions as a reference index that maps high-performing models to their specific datasets and evaluation metrics. It organizes these results through domain-specific categorization to help track progress and monitor performance trends in machine learning research. The platform supports algorithmic baseline research, model performance tracking,

    Provides a centralized registry mapping high-performance models to their specific evaluation datasets and metrics.

    عرض على GitHub↗8,900
  • facebookresearch/ditالصورة الرمزية لـ facebookresearch

    facebookresearch/DiT

    8,642عرض على GitHub↗

    DiT is a latent diffusion model and transformer-based generative AI framework implemented in PyTorch. It functions as a class-conditional image generator that replaces traditional convolutional backbones with a transformer architecture to synthesize high-fidelity images. The project utilizes patch-based latent processing and latent space compression to operate on low-dimensional image representations. It incorporates class-conditional guidance and adjustable guidance scales to control the visual content of generated images during the sampling process. The framework covers distributed model t

    Enables parallel generation of large image batches to evaluate model performance across diverse datasets.

    Python
    عرض على GitHub↗8,642
  • arize-ai/phoenixالصورة الرمزية لـ Arize-ai

    Arize-ai/phoenix

    8,605عرض على GitHub↗

    Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and

    Creates structured collections of inputs and outputs specifically for benchmarking and evaluating model performance.

    Jupyter Notebookagentsai-monitoringai-observability
    عرض على GitHub↗8,605
السابق1234…5التالي
  1. Home
  2. Artificial Intelligence & ML
  3. Dataset Management
  4. Evaluation Datasets

استكشف الوسوم الفرعية

  • Automated Dataset EvaluationSystems that automatically execute evaluators against structured benchmark datasets. **Distinct from Evaluation Datasets:** Focuses on the execution of the evaluation process on a dataset, not the dataset storage itself.
  • Benchmark RegistriesDirectories that map specific models to evaluation datasets and performance metrics for benchmarking. **Distinct from Evaluation Datasets:** Distinct from Evaluation Datasets: focuses on the registry mapping results to datasets, not the datasets themselves.
  • Cross-Dataset ValidationTesting a model trained on one dataset against a different target dataset to measure generalization. **Distinct from Evaluation Datasets:** Focuses on performance validation across different datasets rather than just using a standard evaluation set
  • Dataset Curation4 وسوم فرعيةProcesses and labels raw interaction data to create high-quality reference datasets for evaluation. **Distinct from Evaluation Datasets:** Focuses on the process of collecting and refining data for a baseline, rather than the storage of the final dataset
  • Dataset Splitting Utilities3 وسوم فرعيةTools for dividing interaction data into training and testing sets using various sampling methods. **Distinct from Evaluation Datasets:** Focuses on the process of splitting data rather than the curated collections of evaluation datasets
  • Evaluation Dataset Standardizers5 وسوم فرعيةTools for converting raw interaction data into structured formats for consistent performance benchmarking. **Distinct from Evaluation Datasets:** Distinct from Evaluation Datasets: focuses on the standardization and conversion process rather than the datasets themselves.
  • Evaluation Dataset Structurers1 وسم فرعيTools for organizing inputs and metadata into formats that support performance slicing and systematic analysis. **Distinct from Evaluation Datasets:** Distinct from Evaluation Datasets: focuses on the structural organization and slicing capabilities rather than the dataset content.
  • Evaluation Output Collection1 وسم فرعيUtilities for generating and aggregating model responses to create datasets for benchmarking. **Distinct from Evaluation Datasets:** Focuses on collecting the actual model outputs for testing rather than the structured inputs of the dataset
  • Evaluation Sample BatchersUtilities for grouping evaluation data into chunks to ensure proportional representation and efficient processing. **Distinct from Evaluation Datasets:** Distinct from Evaluation Datasets: focuses on the batching and chunking logic for processing rather than the dataset structure itself.
  • Experiment Management Interfaces1 وسم فرعيInteractive tools for registering scorers, running batch evaluations, and curating datasets for model testing. **Distinct from Evaluation Datasets:** Focuses on the management interface for evaluation experiments, distinct from the datasets themselves.
  • GenerativeDatasets created specifically to benchmark the performance of generative models. **Distinct from Evaluation Datasets:** Focuses on the generation of visual datasets for benchmarking, not just static collections of inputs for LLMs.
  • Judge Performance BenchmarkingProcesses that compare automated judge scores against a human-annotated golden dataset to measure evaluator accuracy. **Distinct from Automated Dataset Evaluation:** Evaluates the performance of the judge itself, while Automated Dataset Evaluation evaluates the model using a judge.
  • Model Experiment ExecutionRunning a set of tasks against a dataset and applying evaluators to compare results across versions. **Distinct from Automated Dataset Evaluation:** Focuses on comparative experimentation rather than just the execution of a single automated evaluation.
  • Stratified Splitting ToolsUtilities for dividing datasets while maintaining class proportions to ensure representative evaluation. **Distinct from Evaluation Datasets:** Focuses on the specific method of stratified splitting rather than general evaluation dataset management
  • Test Dataset AnalyzersUtilities for exporting and reviewing generated test data before evaluation. **Distinct from Evaluation Datasets:** Distinct from evaluation datasets: focuses on the analysis and manual review of test data rather than the dataset storage itself.