awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

10 रिपॉजिटरी

Awesome GitHub RepositoriesModel Evaluation Suites

Frameworks for benchmarking and comparing machine learning model performance through automated and human-in-the-loop testing.

Distinguishing note: None of the candidates are relevant; they focus on UI layout or security, whereas this is a domain-specific AI benchmarking capability.

Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Model Evaluation Suites. Refine with filters or upvote what's useful.

Awesome Model Evaluation Suites GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • vision-cair/minigpt-4Vision-CAIR का अवतार

    Vision-CAIR/MiniGPT-4

    25,679GitHub पर देखें↗

    MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar

    Ships a suite of assessment scripts for benchmarking accuracy in vision and language understanding.

    Python
    GitHub पर देखें↗25,679
  • letta-ai/lettaletta-ai का अवतार

    letta-ai/letta

    21,168GitHub पर देखें↗

    Letta is a framework for building, deploying, and managing autonomous AI agents that maintain persistent state across long-term interactions. It provides a comprehensive suite of primitives for defining agents with configurable personas, modular memory blocks, and tool-use capabilities, enabling them to retain user preferences and conversation history over extended sessions. The platform distinguishes itself through its advanced memory management and orchestration capabilities. It allows agents to autonomously update their own memory, perform retrieval-augmented generation, and coordinate com

    Runs defined evaluation tasks against components and generates output results to verify performance and behavior.

    Pythonaiai-agentsllm
    GitHub पर देखें↗21,168
  • conardli/easy-datasetConardLi का अवतार

    ConardLi/easy-dataset

    13,394GitHub पर देखें↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    Facilitates side-by-side model testing by anonymizing outputs to capture unbiased human preferences and objective performance metrics.

    JavaScriptdatasetfine-tuningjavascript
    GitHub पर देखें↗13,394
  • owainlewis/awesome-artificial-intelligenceowainlewis का अवतार

    owainlewis/awesome-artificial-intelligence

    12,960GitHub पर देखें↗

    This project is a comprehensive repository and curated index of resources, research papers, and development frameworks designed to support the construction and deployment of intelligent systems. It serves as a centralized knowledge base for developers seeking to navigate the technical landscape of artificial intelligence, ranging from foundational educational materials to specialized implementation guides. The repository distinguishes itself by providing structured directories for comparing generative artificial intelligence providers, including aggregated performance metrics, pricing data, a

    Executes standardized evaluation suites to ensure consistent quality and reliability in model outputs.

    aiartificial-intelligencedeep-learning
    GitHub पर देखें↗12,960
  • pyannote/pyannote-audiopyannote का अवतार

    pyannote/pyannote-audio

    9,203GitHub पर देखें↗

    Pyannote.audio is a PyTorch toolkit for speaker diarization, speaker identification, and speech activity detection. Its primary purpose is to partition audio recordings into segments and assign each segment to a specific speaker identity to determine who spoke when. The project includes a framework for classifying speaker identities and a pipeline for distinguishing human speech from background noise. It provides specialized tools for handling symmetric-overlap speech, where multiple speakers talk simultaneously, and employs learnable band-pass filters for raw waveform feature extraction. Th

    Provides a comprehensive suite of metrics for computing diarization error rates and speaker boundary precision.

    Jupyter Notebookoverlapped-speech-detectionpretrained-modelspytorch
    GitHub पर देखें↗9,203
  • lianjiatech/belleLianjiaTech का अवतार

    LianjiaTech/BELLE

    8,273GitHub पर देखें↗

    BELLE is a specialized implementation of Chinese conversational large language models, encompassing a full instruction tuning framework. It provides a pipeline for training, evaluating, and deploying models optimized for natural language understanding and dialogue tasks in the Chinese language. The project is distinguished by its integrated approach to model refinement, combining the curation of multi-million entry instruction datasets with a distributed training pipeline. This pipeline supports both full fine-tuning and low-rank adaptation to optimize conversational performance. The system

    Provides a comprehensive suite of categorized test benchmarks and scoring prompts to assess model outputs.

    HTMLbloomchinese-nlpgpt-evaluation
    GitHub पर देखें↗8,273
  • pytorch/ignitepytorch का अवतार

    pytorch/ignite

    4,770GitHub पर देखें↗

    Ignite PyTorch न्यूरल नेटवर्क के लिए एक उच्च-स्तरीय ट्रेनिंग फ्रेमवर्क है जो एक ट्रेनिंग इंजन और डीप लर्निंग लाइफसाइकिल मैनेजर के रूप में कार्य करता है। यह ट्रेनिंग और इवैल्यूएशन लूप को व्यवस्थित और स्वचालित करने, डेटा इटरेटर्स को प्रबंधित करने और मॉडल ट्रेनिंग प्रक्रिया के दौरान विशिष्ट मील के पत्थर पर इवेंट हैंडलर्स को ट्रिगर करने के लिए एक संरचित सिस्टम प्रदान करता है। यह प्रोजेक्ट डिस्ट्रीब्यूटेड ट्रेनिंग और मॉडल इवैल्यूएशन के लिए टूल्स के एक व्यापक सूट के माध्यम से खुद को अलग करता है। इसमें ग्रेडिएंट्स को सिंक्रोनाइज़ करने और कई GPUs या नोड्स के बीच सामूहिक संचार का समन्वय करने के लिए यूटिलिटीज, साथ ही परफॉरमेंस मेट्रिक्स की गणना और k-fold क्रॉस-वैलिडेशन करने के लिए एक इवैल्यूएशन सूट शामिल है। इसकी व्यापक क्षमताएं ट्रेनिंग वर्कफ़्लो ऑटोमेशन को कवर करती हैं, जिसमें लर्निंग रेट शेड्यूलिंग, अर्ली स्टॉपिंग और हाइपरपैरामीटर ऑप्टिमाइज़ेशन शामिल हैं। फ्रेमवर्क एक्सपेरिमेंट ट्रैकिंग, निष्पादन समय प्रोफाइलिंग और मेमोरी उपयोग को ऑप्टिमाइज़ करने के लिए मिक्स्ड प्रिसिजन ट्रेनिंग के लिए ऑब्जर्वेबिलिटी टूल्स भी प्रदान करता है। मॉडल चेकपॉइंट्स को प्रबंधित करने और ट्रेनिंग सत्रों को रिकवर करने के लिए स्टेट पर्सिस्टेंस मैकेनिज्म शामिल हैं। डिप्लॉयमेंट और एनवायरनमेंट सेटअप को सरल बनाने के लिए कंटेनराइज़्ड एनवायरनमेंट उपलब्ध हैं।

    Ships a comprehensive suite for computing performance metrics and performing cross-validation on PyTorch models.

    Python
    GitHub पर देखें↗4,770
  • districtdatalabs/yellowbrickDistrictDataLabs का अवतार

    DistrictDataLabs/yellowbrick

    4,398GitHub पर देखें↗

    Yellowbrick is a machine learning visualization library and model diagnostic tool designed to analyze feature importance, target distributions, and model error metrics. It serves as a visual toolkit for diagnosing underfitting and overfitting through the use of validation and learning curves. The project provides specialized suites for evaluating predictive models and unsupervised learning. It enables the determination of optimal cluster counts via elbow methods and silhouette coefficients, and assesses classifier and regressor quality through ROC curves, confusion matrices, and residual plot

    Offers a suite for generating ROC curves, confusion matrices, and residual plots to assess model quality.

    Python
    GitHub पर देखें↗4,398
  • starsfieldai/r1-vStarsfieldAI का अवतार

    StarsfieldAI/R1-V

    4,060GitHub पर देखें↗

    R1-V मल्टीमॉडल मॉडल के विकास के लिए एक टूलसेट है, जो बड़े विज़न-लैंग्वेज मॉडल के तर्क (reasoning) और फीडबैक लूप को अनुकूलित करने के लिए डिज़ाइन किया गया एक कम लागत वाला प्रशिक्षण वातावरण प्रदान करता है। यह एक प्रशिक्षण फ्रेमवर्क, फाइन-ट्यूनिंग पाइपलाइन्स और प्रदर्शन मूल्यांकन उपकरणों को एकीकृत करता है। इस प्रोजेक्ट में एक रीइन्फोर्समेंट लर्निंग फ्रेमवर्क है जो दृश्य सत्यापन (visual verification) के आधार पर सही आउटपुट को पुरस्कृत करके दृश्य तर्क और सामान्यीकरण में सुधार करता है। इसमें लेबल किए गए डेटासेट और कॉन्फ़िगरेशन फ़ाइलों का उपयोग करके विशिष्ट कार्यों के लिए विज़न-लैंग्वेज मॉडल को कस्टमाइज़ करने के लिए एक सुपरवाइज्ड फाइन-ट्यूनिंग पाइपलाइन भी शामिल है। यह सूट दृश्य तर्क मूल्यांकन उपकरणों और विशेष रूप से गिनती और ज्यामिति कार्यों पर मॉडल के प्रदर्शन का आकलन करने के लिए डेटासेट को शामिल करता है।

    Ships dedicated benchmark datasets and scripts to measure accuracy on counting and geometry problems.

    Python
    GitHub पर देखें↗4,060
  • evolvinglmms-lab/lmms-evalEvolvingLMMs-Lab का अवतार

    EvolvingLMMs-Lab/lmms-eval

    3,701GitHub पर देखें↗

    lmms-eval is a benchmarking system and performance analysis suite designed to measure the capabilities of large multimodal models. It provides a framework for evaluating models across text, image, audio, and video datasets, serving as a multimodal dataset orchestrator and benchmarking tool to quantify accuracy and efficiency. The project distinguishes itself through a unified multimodal message protocol that structures diverse media inputs for consistent model consumption. It features specialized benchmarking for audio, video, visual, document, and spatial reasoning, alongside tools for model

    Ships a suite for computing statistical significance, throughput, and token usage for multimodal evaluations.

    Pythonagiaudio-evaluationbenchmark
    GitHub पर देखें↗3,701
  1. Home
  2. Artificial Intelligence & ML
  3. Model Evaluation Suites

सब-टैग एक्सप्लोर करें

  • Diarization Evaluation Suites1 सब-टैगSpecialized benchmarking frameworks for measuring diarization error rates and boundary precision. **Distinct from Model Evaluation Suites:** Specific to speaker diarization rather than general ML model evaluation