awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

10 repositorios

Awesome GitHub RepositoriesModel Evaluation Suites

Frameworks for benchmarking and comparing machine learning model performance through automated and human-in-the-loop testing.

Distinguishing note: None of the candidates are relevant; they focus on UI layout or security, whereas this is a domain-specific AI benchmarking capability.

Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Model Evaluation Suites. Refine with filters or upvote what's useful.

Awesome Model Evaluation Suites GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • vision-cair/minigpt-4Avatar de Vision-CAIR

    Vision-CAIR/MiniGPT-4

    25,679Ver en GitHub↗

    MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar

    Ships a suite of assessment scripts for benchmarking accuracy in vision and language understanding.

    Python
    Ver en GitHub↗25,679
  • letta-ai/lettaAvatar de letta-ai

    letta-ai/letta

    21,168Ver en GitHub↗

    Letta is a framework for building, deploying, and managing autonomous AI agents that maintain persistent state across long-term interactions. It provides a comprehensive suite of primitives for defining agents with configurable personas, modular memory blocks, and tool-use capabilities, enabling them to retain user preferences and conversation history over extended sessions. The platform distinguishes itself through its advanced memory management and orchestration capabilities. It allows agents to autonomously update their own memory, perform retrieval-augmented generation, and coordinate com

    Runs defined evaluation tasks against components and generates output results to verify performance and behavior.

    Pythonaiai-agentsllm
    Ver en GitHub↗21,168
  • conardli/easy-datasetAvatar de ConardLi

    ConardLi/easy-dataset

    13,394Ver en GitHub↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    Facilitates side-by-side model testing by anonymizing outputs to capture unbiased human preferences and objective performance metrics.

    JavaScriptdatasetfine-tuningjavascript
    Ver en GitHub↗13,394
  • owainlewis/awesome-artificial-intelligenceAvatar de owainlewis

    owainlewis/awesome-artificial-intelligence

    12,960Ver en GitHub↗

    This project is a comprehensive repository and curated index of resources, research papers, and development frameworks designed to support the construction and deployment of intelligent systems. It serves as a centralized knowledge base for developers seeking to navigate the technical landscape of artificial intelligence, ranging from foundational educational materials to specialized implementation guides. The repository distinguishes itself by providing structured directories for comparing generative artificial intelligence providers, including aggregated performance metrics, pricing data, a

    Executes standardized evaluation suites to ensure consistent quality and reliability in model outputs.

    aiartificial-intelligencedeep-learning
    Ver en GitHub↗12,960
  • pyannote/pyannote-audioAvatar de pyannote

    pyannote/pyannote-audio

    9,203Ver en GitHub↗

    Pyannote.audio is a PyTorch toolkit for speaker diarization, speaker identification, and speech activity detection. Its primary purpose is to partition audio recordings into segments and assign each segment to a specific speaker identity to determine who spoke when. The project includes a framework for classifying speaker identities and a pipeline for distinguishing human speech from background noise. It provides specialized tools for handling symmetric-overlap speech, where multiple speakers talk simultaneously, and employs learnable band-pass filters for raw waveform feature extraction. Th

    Provides a comprehensive suite of metrics for computing diarization error rates and speaker boundary precision.

    Jupyter Notebookoverlapped-speech-detectionpretrained-modelspytorch
    Ver en GitHub↗9,203
  • lianjiatech/belleAvatar de LianjiaTech

    LianjiaTech/BELLE

    8,273Ver en GitHub↗

    BELLE is a specialized implementation of Chinese conversational large language models, encompassing a full instruction tuning framework. It provides a pipeline for training, evaluating, and deploying models optimized for natural language understanding and dialogue tasks in the Chinese language. The project is distinguished by its integrated approach to model refinement, combining the curation of multi-million entry instruction datasets with a distributed training pipeline. This pipeline supports both full fine-tuning and low-rank adaptation to optimize conversational performance. The system

    Provides a comprehensive suite of categorized test benchmarks and scoring prompts to assess model outputs.

    HTMLbloomchinese-nlpgpt-evaluation
    Ver en GitHub↗8,273
  • pytorch/igniteAvatar de pytorch

    pytorch/ignite

    4,770Ver en GitHub↗

    Ignite es un framework de entrenamiento de alto nivel para redes neuronales en PyTorch que sirve como motor de entrenamiento y gestor del ciclo de vida del aprendizaje profundo. Proporciona un sistema estructurado para organizar y automatizar bucles de entrenamiento y evaluación, gestionando iteradores de datos y activando manejadores de eventos en hitos específicos durante el proceso de entrenamiento del modelo. El proyecto se distingue por una suite integral de herramientas para el entrenamiento distribuido y la evaluación de modelos. Incluye utilidades para sincronizar gradientes y coordinar la comunicación colectiva a través de múltiples GPUs o nodos, así como una suite de evaluación para calcular métricas de rendimiento y realizar validación cruzada k-fold. Sus capacidades más amplias cubren la automatización del flujo de trabajo de entrenamiento, incluyendo la programación de la tasa de aprendizaje, parada temprana y optimización de hiperparámetros. El framework también proporciona herramientas de observabilidad para el seguimiento de experimentos, perfilado de tiempo de ejecución y entrenamiento de precisión mixta para optimizar el uso de memoria. Se incluyen mecanismos de persistencia de estado para gestionar checkpoints del modelo y recuperar sesiones de entrenamiento. Hay entornos contenedorizados disponibles para simplificar el despliegue y la configuración del entorno.

    Ships a comprehensive suite for computing performance metrics and performing cross-validation on PyTorch models.

    Python
    Ver en GitHub↗4,770
  • districtdatalabs/yellowbrickAvatar de DistrictDataLabs

    DistrictDataLabs/yellowbrick

    4,398Ver en GitHub↗

    Yellowbrick es una librería de visualización de machine learning y herramienta de diagnóstico de modelos diseñada para analizar la importancia de las características, distribuciones objetivo y métricas de error del modelo. Sirve como un kit de herramientas visual para diagnosticar el subajuste (underfitting) y sobreajuste (overfitting) mediante el uso de curvas de validación y aprendizaje. El proyecto proporciona suites especializadas para evaluar modelos predictivos y aprendizaje no supervisado. Permite la determinación de conteos de clústeres óptimos mediante métodos de codo y coeficientes de silueta, y evalúa la calidad de clasificadores y regresores a través de curvas ROC, matrices de confusión y gráficos de residuos. La librería cubre varias áreas de capacidad de alto nivel, incluyendo análisis de ingeniería de características para identificar variables predictivas, ajuste de hiperparámetros para ajustar la complejidad del modelo y diagnóstico de errores de regresión para identificar puntos de datos influyentes. También incluye herramientas para proyección de aprendizaje de variedades (manifold learning) para visualizar datos de alta dimensión y corpus de texto. La herramienta se integra con la API de Scikit-Learn para consumir métodos estándar de fit y predict.

    Offers a suite for generating ROC curves, confusion matrices, and residual plots to assess model quality.

    Python
    Ver en GitHub↗4,398
  • starsfieldai/r1-vAvatar de StarsfieldAI

    StarsfieldAI/R1-V

    4,060Ver en GitHub↗

    R1-V es un conjunto de herramientas para el desarrollo de modelos multimodales, que proporciona un entorno de entrenamiento de bajo coste diseñado para optimizar los bucles de razonamiento y retroalimentación de modelos grandes de visión-lenguaje. Integra un framework de entrenamiento, pipelines de ajuste fino (fine-tuning) y herramientas de evaluación de rendimiento. El proyecto cuenta con un framework de aprendizaje por refuerzo que mejora el razonamiento visual y la generalización al recompensar las salidas correctas basadas en verificación visual. También incluye un pipeline de ajuste fino supervisado para personalizar modelos de visión-lenguaje en tareas específicas mediante datasets etiquetados y archivos de configuración. La suite abarca herramientas de evaluación de razonamiento visual y datasets específicos para evaluar el rendimiento del modelo en tareas de conteo y geometría.

    Ships dedicated benchmark datasets and scripts to measure accuracy on counting and geometry problems.

    Python
    Ver en GitHub↗4,060
  • evolvinglmms-lab/lmms-evalAvatar de EvolvingLMMs-Lab

    EvolvingLMMs-Lab/lmms-eval

    3,701Ver en GitHub↗

    lmms-eval is a benchmarking system and performance analysis suite designed to measure the capabilities of large multimodal models. It provides a framework for evaluating models across text, image, audio, and video datasets, serving as a multimodal dataset orchestrator and benchmarking tool to quantify accuracy and efficiency. The project distinguishes itself through a unified multimodal message protocol that structures diverse media inputs for consistent model consumption. It features specialized benchmarking for audio, video, visual, document, and spatial reasoning, alongside tools for model

    Ships a suite for computing statistical significance, throughput, and token usage for multimodal evaluations.

    Pythonagiaudio-evaluationbenchmark
    Ver en GitHub↗3,701
  1. Home
  2. Artificial Intelligence & ML
  3. Model Evaluation Suites

Explorar subetiquetas

  • Diarization Evaluation Suites1 sub-etiquetaSpecialized benchmarking frameworks for measuring diarization error rates and boundary precision. **Distinct from Model Evaluation Suites:** Specific to speaker diarization rather than general ML model evaluation