10 रिपॉजिटरी
Frameworks for benchmarking and comparing machine learning model performance through automated and human-in-the-loop testing.
Distinguishing note: None of the candidates are relevant; they focus on UI layout or security, whereas this is a domain-specific AI benchmarking capability.
Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Model Evaluation Suites. Refine with filters or upvote what's useful.
MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar
Ships a suite of assessment scripts for benchmarking accuracy in vision and language understanding.
Letta is a framework for building, deploying, and managing autonomous AI agents that maintain persistent state across long-term interactions. It provides a comprehensive suite of primitives for defining agents with configurable personas, modular memory blocks, and tool-use capabilities, enabling them to retain user preferences and conversation history over extended sessions. The platform distinguishes itself through its advanced memory management and orchestration capabilities. It allows agents to autonomously update their own memory, perform retrieval-augmented generation, and coordinate com
Runs defined evaluation tasks against components and generates output results to verify performance and behavior.
Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side
Facilitates side-by-side model testing by anonymizing outputs to capture unbiased human preferences and objective performance metrics.
This project is a comprehensive repository and curated index of resources, research papers, and development frameworks designed to support the construction and deployment of intelligent systems. It serves as a centralized knowledge base for developers seeking to navigate the technical landscape of artificial intelligence, ranging from foundational educational materials to specialized implementation guides. The repository distinguishes itself by providing structured directories for comparing generative artificial intelligence providers, including aggregated performance metrics, pricing data, a
Executes standardized evaluation suites to ensure consistent quality and reliability in model outputs.
Pyannote.audio is a PyTorch toolkit for speaker diarization, speaker identification, and speech activity detection. Its primary purpose is to partition audio recordings into segments and assign each segment to a specific speaker identity to determine who spoke when. The project includes a framework for classifying speaker identities and a pipeline for distinguishing human speech from background noise. It provides specialized tools for handling symmetric-overlap speech, where multiple speakers talk simultaneously, and employs learnable band-pass filters for raw waveform feature extraction. Th
Provides a comprehensive suite of metrics for computing diarization error rates and speaker boundary precision.
BELLE is a specialized implementation of Chinese conversational large language models, encompassing a full instruction tuning framework. It provides a pipeline for training, evaluating, and deploying models optimized for natural language understanding and dialogue tasks in the Chinese language. The project is distinguished by its integrated approach to model refinement, combining the curation of multi-million entry instruction datasets with a distributed training pipeline. This pipeline supports both full fine-tuning and low-rank adaptation to optimize conversational performance. The system
Provides a comprehensive suite of categorized test benchmarks and scoring prompts to assess model outputs.
Ignite PyTorch न्यूरल नेटवर्क के लिए एक उच्च-स्तरीय ट्रेनिंग फ्रेमवर्क है जो एक ट्रेनिंग इंजन और डीप लर्निंग लाइफसाइकिल मैनेजर के रूप में कार्य करता है। यह ट्रेनिंग और इवैल्यूएशन लूप को व्यवस्थित और स्वचालित करने, डेटा इटरेटर्स को प्रबंधित करने और मॉडल ट्रेनिंग प्रक्रिया के दौरान विशिष्ट मील के पत्थर पर इवेंट हैंडलर्स को ट्रिगर करने के लिए एक संरचित सिस्टम प्रदान करता है। यह प्रोजेक्ट डिस्ट्रीब्यूटेड ट्रेनिंग और मॉडल इवैल्यूएशन के लिए टूल्स के एक व्यापक सूट के माध्यम से खुद को अलग करता है। इसमें ग्रेडिएंट्स को सिंक्रोनाइज़ करने और कई GPUs या नोड्स के बीच सामूहिक संचार का समन्वय करने के लिए यूटिलिटीज, साथ ही परफॉरमेंस मेट्रिक्स की गणना और k-fold क्रॉस-वैलिडेशन करने के लिए एक इवैल्यूएशन सूट शामिल है। इसकी व्यापक क्षमताएं ट्रेनिंग वर्कफ़्लो ऑटोमेशन को कवर करती हैं, जिसमें लर्निंग रेट शेड्यूलिंग, अर्ली स्टॉपिंग और हाइपरपैरामीटर ऑप्टिमाइज़ेशन शामिल हैं। फ्रेमवर्क एक्सपेरिमेंट ट्रैकिंग, निष्पादन समय प्रोफाइलिंग और मेमोरी उपयोग को ऑप्टिमाइज़ करने के लिए मिक्स्ड प्रिसिजन ट्रेनिंग के लिए ऑब्जर्वेबिलिटी टूल्स भी प्रदान करता है। मॉडल चेकपॉइंट्स को प्रबंधित करने और ट्रेनिंग सत्रों को रिकवर करने के लिए स्टेट पर्सिस्टेंस मैकेनिज्म शामिल हैं। डिप्लॉयमेंट और एनवायरनमेंट सेटअप को सरल बनाने के लिए कंटेनराइज़्ड एनवायरनमेंट उपलब्ध हैं।
Ships a comprehensive suite for computing performance metrics and performing cross-validation on PyTorch models.
Yellowbrick is a machine learning visualization library and model diagnostic tool designed to analyze feature importance, target distributions, and model error metrics. It serves as a visual toolkit for diagnosing underfitting and overfitting through the use of validation and learning curves. The project provides specialized suites for evaluating predictive models and unsupervised learning. It enables the determination of optimal cluster counts via elbow methods and silhouette coefficients, and assesses classifier and regressor quality through ROC curves, confusion matrices, and residual plot
Offers a suite for generating ROC curves, confusion matrices, and residual plots to assess model quality.
R1-V मल्टीमॉडल मॉडल के विकास के लिए एक टूलसेट है, जो बड़े विज़न-लैंग्वेज मॉडल के तर्क (reasoning) और फीडबैक लूप को अनुकूलित करने के लिए डिज़ाइन किया गया एक कम लागत वाला प्रशिक्षण वातावरण प्रदान करता है। यह एक प्रशिक्षण फ्रेमवर्क, फाइन-ट्यूनिंग पाइपलाइन्स और प्रदर्शन मूल्यांकन उपकरणों को एकीकृत करता है। इस प्रोजेक्ट में एक रीइन्फोर्समेंट लर्निंग फ्रेमवर्क है जो दृश्य सत्यापन (visual verification) के आधार पर सही आउटपुट को पुरस्कृत करके दृश्य तर्क और सामान्यीकरण में सुधार करता है। इसमें लेबल किए गए डेटासेट और कॉन्फ़िगरेशन फ़ाइलों का उपयोग करके विशिष्ट कार्यों के लिए विज़न-लैंग्वेज मॉडल को कस्टमाइज़ करने के लिए एक सुपरवाइज्ड फाइन-ट्यूनिंग पाइपलाइन भी शामिल है। यह सूट दृश्य तर्क मूल्यांकन उपकरणों और विशेष रूप से गिनती और ज्यामिति कार्यों पर मॉडल के प्रदर्शन का आकलन करने के लिए डेटासेट को शामिल करता है।
Ships dedicated benchmark datasets and scripts to measure accuracy on counting and geometry problems.
lmms-eval is a benchmarking system and performance analysis suite designed to measure the capabilities of large multimodal models. It provides a framework for evaluating models across text, image, audio, and video datasets, serving as a multimodal dataset orchestrator and benchmarking tool to quantify accuracy and efficiency. The project distinguishes itself through a unified multimodal message protocol that structures diverse media inputs for consistent model consumption. It features specialized benchmarking for audio, video, visual, document, and spatial reasoning, alongside tools for model
Ships a suite for computing statistical significance, throughput, and token usage for multimodal evaluations.