2 रिपॉजिटरी
Standardized tests and metrics for evaluating the quality and performance of speech and audio models.
Distinct from Model Performance Benchmarking: Specifically targets audio and speech fidelity and accuracy, whereas the parent is general model performance benchmarking.
Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Audio Performance Benchmarks. Refine with filters or upvote what's useful.
Kimi-Audio is a large language model audio foundation model designed to understand audio input and generate high-fidelity speech responses in real time. It functions as a unified system encompassing a text-to-speech synthesis engine and a speech-to-text transcription tool. The project enables real-time audio conversations through a multi-modal conversation loop and chunk-wise streaming detokenization to reduce playback latency. It provides controls over speech speed, accent, and emotional tone during conversational audio generation. The system covers audio intelligence capabilities, includin
Ships a benchmarking harness with standardized metrics and side-by-side inference recipes for audio models.
lmms-eval is a benchmarking system and performance analysis suite designed to measure the capabilities of large multimodal models. It provides a framework for evaluating models across text, image, audio, and video datasets, serving as a multimodal dataset orchestrator and benchmarking tool to quantify accuracy and efficiency. The project distinguishes itself through a unified multimodal message protocol that structures diverse media inputs for consistent model consumption. It features specialized benchmarking for audio, video, visual, document, and spatial reasoning, alongside tools for model
Assesses model capabilities in speech recognition, speech translation, and audio-based question answering.