awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 repositorios

Awesome GitHub RepositoriesModel Quality Metrics

Calculations for accuracy, perplexity, and F1 scores to quantify the performance of language models.

Distinct from Accuracy Calculators: Distinct from Accuracy Calculators: encompasses a broader set of quality metrics including perplexity and F1 scores.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Model Quality Metrics. Refine with filters or upvote what's useful.

Awesome Model Quality Metrics GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • evolvinglmms-lab/lmms-evalAvatar de EvolvingLMMs-Lab

    EvolvingLMMs-Lab/lmms-eval

    3,701Ver en GitHub↗

    lmms-eval is a benchmarking system and performance analysis suite designed to measure the capabilities of large multimodal models. It provides a framework for evaluating models across text, image, audio, and video datasets, serving as a multimodal dataset orchestrator and benchmarking tool to quantify accuracy and efficiency. The project distinguishes itself through a unified multimodal message protocol that structures diverse media inputs for consistent model consumption. It features specialized benchmarking for audio, video, visual, document, and spatial reasoning, alongside tools for model

    Calculates accuracy, perplexity, and F1 scores using configurable aggregation methods to quantify model quality.

    Pythonagiaudio-evaluationbenchmark
    Ver en GitHub↗3,701
  • huggingface/transfer-learning-conv-aiAvatar de huggingface

    huggingface/transfer-learning-conv-ai

    1,757Ver en GitHub↗

    This framework is a research-oriented toolkit designed for training, fine-tuning, and evaluating conversational agents using transformer-based language architectures. It provides an integrated environment for adapting large pre-trained models to specific dialogue datasets, enabling the development of systems capable of generating coherent, human-like responses. The project distinguishes itself through its support for multi-GPU distributed training, which accelerates the optimization of large-scale models. It also features configurable probabilistic decoding strategies, such as nucleus and gre

    Calculates standard metrics like perplexity and F1 scores to quantify the quality of conversational models.

    Pythonchatbotsdeep-learningdialog
    Ver en GitHub↗1,757
  1. Home
  2. Artificial Intelligence & ML
  3. Prediction Visualization
  4. Model Quality Metrics