34 रिपॉजिटरी
Numerical representations used to condition models on specific voice characteristics.
Distinguishing note: Focuses on input conditioning rather than general audio processing.
Explore 34 awesome GitHub repositories matching artificial intelligence & ml · Speaker Embeddings. Refine with filters or upvote what's useful.
This project is a neural text-to-speech engine and voice cloning toolkit designed to generate synthetic speech that mimics the vocal characteristics of a target speaker. It functions as a real-time audio synthesizer, utilizing a deep learning pipeline to convert written text into high-fidelity speech output with minimal latency. The system employs a transfer learning framework that leverages pre-trained speaker verification models to adapt synthesis to new, unseen vocal identities. By using an encoder-based speaker embedding process, the toolkit maps variable-length audio samples into a laten
Encodes variable-length audio inputs into fixed-dimensional latent vectors that capture unique speaker characteristics.
LocalAI is a local generative AI platform and inference engine designed to host large language, vision, and audio models on private hardware. It functions as an API compatible gateway that mimics proprietary service endpoints, allowing existing third-party software to integrate with a self-hosted backend. The platform distinguishes itself as a distributed AI model orchestrator, capable of scaling inference across machine clusters using VRAM-aware routing and hardware coordination. It provides a unified interface for diverse open-source backends and supports self-hosted RAG infrastructure thro
Identifies specific speakers and analyzes voice characteristics including age, gender, and emotion.
यह प्रोजेक्ट एक डीप लर्निंग टेक्स्ट-टू-स्पीच टूलकिट है जिसका उपयोग न्यूरल स्पीच सिंथेसिस मॉडल को प्रशिक्षित और तैनात करने के लिए किया जाता है। यह लिखित टेक्स्ट को बोले गए ऑडियो में बदलने के लिए एक व्यापक फ्रेमवर्क प्रदान करता है। टूलकिट में एक वॉयस क्लोनिंग सिस्टम शामिल है जो छोटे ऑडियो नमूनों से स्पीकर एम्बेडिंग निकालकर विशिष्ट मानव आवाजों की नकल करता है। यह मल्टी-स्पीकर ऑडियो सिंथेसिस का भी समर्थन करता है। सिस्टम में स्पीच डेटासेट क्यूरेशन, प्रदर्शन ट्रैकिंग के साथ कस्टम मॉडल प्रशिक्षण और ऑडियो जनरेशन के लिए कमांड-लाइन इंटरफेस सहित पूर्ण स्पीच सिंथेसिस पाइपलाइन शामिल है।
Extracts speaker embeddings from audio samples to condition the synthesis model on specific vocal characteristics.
ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles. The system is distinguished by its ability to generate interactive dialogue with realistic vocal nuances. It utilizes a speech nuance controller to insert specific tokens that trigger non-verbal elements, such as laughter, pauses, and interjections, during the synthesis process. The project includes a streaming audio generato
Uses learned vector representations to maintain consistent vocal characteristics across different speakers.
This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a two-stage autoregressive transformer architecture that separates semantic token prediction from acoustic detail reconstruction to balance linguistic accuracy with audio quality. The system is designed to support multilingual output and conversational AI development, enabling the generation of context-aware speech that maintains flow across multiple dialogue turns. The platform distinguishes itself through a production-ready inference server that employs continuous batching to
Uses dedicated identifiers to manage and switch between distinct voice characteristics.
TensorFlow.js is a JavaScript machine learning library used for training and deploying models in web browsers and server-side environments. It functions as a browser-based model trainer, a WebAssembly inference engine, and a WebGPU accelerated tensor library for low-level linear algebra. The project also includes a model converter to transform Python-based models into optimized formats for JavaScript execution. The library distinguishes itself through a pluggable backend architecture that allows mathematical operations to be executed via CPU, WebGL, or WebGPU. It supports the conversion of Py
Provides specialized capabilities to group sentences by comparing word embeddings to determine textual similarity.
Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By utilizing deep learning models and phoneme-based sequence modeling, the system transforms text into natural-sounding audio waveforms suitable for a variety of accessibility and media applications. The platform functions as a server-side inference pipeline that provides a programmatic interface for integrating voice generation into external applications. It distinguishes itself through asynchronous audio streaming, which buffers and delivers generated speech chunks in real time to
Injects specific speaker identity parameters into the synthesis model to allow for distinct vocal characteristics.
Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web
Converts audio waveforms into mathematical embedding vectors representing unique vocal characteristics.
यह प्रोजेक्ट PyTorch के लिए एक कंप्यूटर विजन एक्सप्लेनबल AI लाइब्रेरी और फ्रेमवर्क है, जो गहरे न्यूरल नेटवर्क की आंतरिक निर्णय लेने की प्रक्रियाओं को देखने और ऑडिट करने के लिए टूल का एक सूट प्रदान करता है। यह न्यूरल नेटवर्क एट्रिब्यूशन टूल और डीबगिंग यूटिलिटी के रूप में कार्य करता है ताकि यह पहचान सके कि कौन से छवि क्षेत्र मॉडल भविष्यवाणियों को प्रेरित करते हैं। लाइब्रेरी ग्रेडिएंट-आधारित और ग्रेडिएंट-मुक्त एट्रिब्यूशन विधियों दोनों के लिए अपने समर्थन द्वारा प्रतिष्ठित है, जो मूल मॉडल स्रोत कोड में संशोधनों की आवश्यकता के बिना विजुअल हीटमैप और एट्रिब्यूशन मैप के निर्माण की अनुमति देती है। यह आगे विजुअल कॉन्सेप्ट डिस्कवरी के माध्यम से खुद को अलग करती है, आंतरिक सक्रियणों को व्याख्या योग्य पैटर्न में विघटित करने के लिए मैट्रिक्स फैक्टराइजेशन का उपयोग करती है और लेटेंट एम्बेडिंग को पिक्सेल महत्व पर मैप करती है। फ्रेमवर्क हीटमैप निर्माण और शोधन, विजन ट्रांसफार्मर जैसे आर्किटेक्चर के लिए स्थानिक परिवर्तन, और ऑब्जेक्ट डिटेक्शन और सिमेंटिक सेगमेंटेशन जैसे मल्टी-टास्क विजन लक्ष्यों के लिए अनुकूलन सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। इसमें एक मॉडल फिडेलिटी मूल्यांकन सूट भी शामिल है जो उत्पन्न स्पष्टीकरणों की निष्ठा को मापने के लिए पर्टरबेशन विश्लेषण, एब्लेशन अध्ययन और स्थानीयकरण माप का उपयोग करता है। प्रोजेक्ट विभिन्न मॉडल आउटपुट से एक्सप्लेनबिलिटी टूल को जोड़ने के लिए डायनामिक एक्टिवेशन हुकिंग, कस्टम आर्किटेक्चर अनुकूलन, और लक्ष्य-संचालित उद्देश्य कॉन्फ़िगरेशन के लिए तंत्र प्रदान करता है।
Visualizes the image regions that contribute most to the similarity between an output feature vector and a reference vector.
Omi is an open-source wearable AI platform that captures audio and screen data to provide real-time conversational assistance and memory. It integrates a wearable hardware development kit with a vector memory database and large language model capabilities to create a persistent digital record of user interactions. The platform is distinguished by its BLE audio streaming pipeline, which transmits raw audio from wearable hardware for real-time transcription and speaker identification. It utilizes a plugin-based agent tool framework that allows AI assistants to autonomously invoke custom functio
Matches live voice embeddings against stored profiles to authenticate and distinguish between different speakers.
PaddleSpeech is a comprehensive toolkit of neural models for speech recognition, synthesis, and translation built on the PaddlePaddle deep learning framework. It provides a collection of frameworks and tools for converting spoken audio into written text, synthesizing natural audio from text, and performing direct speech translation. The toolkit includes specialized capabilities for keyword spotting to detect trigger words and speaker verification systems that extract unique voiceprints to identify and distinguish between individuals. It also features end-to-end translation tools that map audi
Generates fixed-dimensional numerical representations of voices to identify and verify individual speaker identities.
Spark-TTS is a deep learning text-to-speech synthesis engine designed to convert written text into high-fidelity audio. It utilizes a transformer-based architecture and autoregressive sequence modeling to generate coherent speech, transforming linguistic input into natural-sounding waveforms through neural speech codec synthesis. The platform distinguishes itself through zero-shot voice cloning, which allows users to mimic a target speaker’s unique vocal identity using only a short reference audio sample without requiring additional model training. It also features cross-lingual phonetic mapp
Extracts acoustic features from short audio samples to condition synthesis models without additional training.
Piper is a local neural text-to-speech engine designed to convert written text into natural human speech entirely on your own hardware. By utilizing a neural synthesis framework, it operates without the need for internet connectivity, ensuring that all audio generation remains private and secure. The system distinguishes itself through a modular architecture that allows for the dynamic loading of speaker embeddings and voice configurations. This enables users to switch between various vocal personas and styles without requiring a full reload of the core synthesis model. By processing input th
Supports dynamic loading of speaker embeddings to adjust vocal characteristics without reloading the core model.
This project is a comprehensive suite for neural speech synthesis, featuring a deep learning text-to-speech engine, a neural speech synthesis trainer, and a voice cloning toolkit. It provides a system for synthesizing human-like speech from text using neural network models and high-fidelity vocoders. The suite includes a speech model conversion utility to transform deep learning models between different formats for deployment across various hardware runtimes. It also provides a self-contained HTTP server to expose pre-trained text-to-speech models as a remote audio API. Capabilities include
Generates numerical representations of vocal characteristics to enable voice cloning and multi-speaker synthesis.
Monolith is a distributed recommendation model framework and asynchronous training engine designed to build and train large-scale deep learning architectures. It functions as a distributed model trainer that processes massive datasets across multiple compute nodes using asynchronous update mechanisms. The system features a dedicated embedding table manager that creates unique, feature-isolated tables to prevent representation collisions. It also includes a real-time weight updater to capture immediate changes in user interest and data hotspots through continuous parameter synchronization. Th
Manages unique embedding tables for different identity features to prevent representation collisions in large models.
This project is an AI voice assistant backend and gateway server designed to connect ESP32 hardware to large language models. It enables real-time conversational AI by processing streaming speech-to-text and text-to-speech interactions, allowing hardware devices to engage in natural language dialogue. The system is distinguished by a modular plugin framework that loads custom feature extensions at runtime and a retrieval-augmented generation engine that queries external knowledge bases for factual accuracy. It further personalizes interactions by using voiceprint mapping to identify individua
Matches incoming audio signatures against stored voiceprints to verify and identify the speaker.
Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo
Identifies and ranks keywords from text using a configurable algorithm.
BERTopic is a topic modeling library used to extract interpretable themes from collections of text documents and images. It functions as a document clustering framework that transforms unstructured data into numerical vectors to group semantically similar content. The project distinguishes itself through a multimodal embedding tool that allows for joint clustering of text and images in a shared vector space. It also features a class-based TF-IDF representation engine to identify representative words for clusters and an integrated system for using large language models to generate natural lang
Matches documents to user-defined labels using cosine similarity while clustering remaining documents.
PyTorch Metric Learning is an open-source library for training neural networks to produce similarity-preserving embedding spaces. It provides a modular framework where interchangeable loss functions, mining strategies, and evaluation tools can be composed to learn representations that map similar items to nearby points and dissimilar items to distant points in the embedding space. The library distinguishes itself through a highly configurable architecture that separates concerns across several interchangeable components. Users can assemble custom loss functions from pluggable distance metrics
Implements loss functions that use configurable similarity measures to separate classes in embedding space.
StyleTTS2 is an adversarial text-to-speech model that uses style diffusion and large speech language models to generate natural-sounding speech from text input. It combines adversarial training with large pre-trained speech models to improve speech quality and reduce artifacts, while employing a style diffusion process that extracts prosodic and timbral features from reference audio to guide speech generation. The model supports multi-speaker voice synthesis by conditioning the diffusion process on speaker-specific embeddings derived from reference utterances, enabling voice cloning and adapt
Controls voice identity by conditioning the diffusion process on speaker-specific embeddings derived from reference utterances.