25 रिपॉजिटरी
Tools for converting written text into spoken audio using AI language models.
Distinct from Text-to-Speech Integrations: Distinct from Text-to-Speech Integrations: focuses on audio generation from text, not on providing interface methods for integration.
Explore 25 awesome GitHub repositories matching artificial intelligence & ml · Text-to-Speech Conversions. Refine with filters or upvote what's useful.
Duix-Avatar is an AI digital human toolkit used to create, clone, and animate realistic virtual personas. It functions as a digital persona cloning tool and a text-to-speech animation API that converts written text or audio into synthetic voice and facial motion markers. The framework provides an offline video generation engine that renders digital human animations and lip-synced videos on local hardware. It includes a specialized lip sync engine to synchronize mouth movements with audio waveforms and a pipeline for extracting facial and vocal features from source media to create synthetic re
Converts written text into synthetic voice and corresponding facial motion markers for virtual character rendering.
Home Assistant is a local home automation platform and server that acts as an IoT device orchestrator. It integrates diverse smart home hardware by wrapping third-party APIs into a standardized logic layer and stores all system state and historical statistics on local hardware to eliminate cloud dependencies. The system functions as a Matter IoT controller and an MQTT home automation bridge, allowing for local interoperability between different manufacturers. It features a state-based entity model and an internal event bus that decouple physical device logic from system automation. The platf
Home Assistant generates spoken audio from text to provide audible feedback or alerts.
LiveTalking is an interactive talking head engine and AI avatar management platform designed to synchronize synthetic speech with facial movements. It functions as a real-time orchestrator that connects large language models and text-to-speech services to neural-rendered digital humans. The project distinguishes itself through low-latency streaming capabilities and the ability to handle real-time conversational interruptions. It supports advanced audio-visual customization, including human voice cloning and the ability to drive avatar expressions using real-time webcam data. The platform cov
Convert written text into spoken audio using a model optimized for fast inference and short-form audio.
ChatTTS-ui is a web-based interface and API wrapper for the ChatTTS model, designed to convert written text and mixed language input into spoken audio. It functions as an AI speech synthesis dashboard and a programmatic generator for creating naturalistic voice output. The project focuses on custom voice profiling and speech nuance control. It allows for the maintenance of consistent speaker characteristics using seed values and data files, while providing controls for tone, laughter, and pauses through behavioral prompts and sampling parameters. The system includes a client-server architect
Converts written text into spoken audio using AI language models for various projects.
MeloTTS is an open-source text-to-speech library that generates natural-sounding speech across six languages, with the ability to mix two languages within a single utterance. Its architecture combines a token-based text frontend with a language-agnostic acoustic model, enabling it to handle bilingual code-switching and produce streaming audio output in real time. The system is designed to run efficiently on standard CPU hardware without requiring a dedicated GPU, using a lightweight neural network for real-time inference. It supports English, Spanish, French, Chinese, Japanese, and Korean, an
Converts written text into natural-sounding speech across multiple languages including English, Spanish, French, Chinese, Japanese, and Korean.
Zonos is a controllable audio synthesis engine and large language model for text-to-speech. It serves as a multilingual speech generator capable of producing audio in English, Japanese, Chinese, French, and German. The system provides zero-shot voice cloning, allowing the replication of specific human voices using short audio samples. It supports the capture of nuanced behaviors, such as whispering, and provides parametric control over speaking rate, pitch, frequency, and emotional tone. The project covers a broad range of expressive speech synthesis and custom audio generation capabilities,
Converts written text into high-quality spoken audio across English, Japanese, Chinese, French, and German.
espeak-ng एक बहुभाषी टेक्स्ट-टू-स्पीच इंजन और C-आधारित लाइब्रेरी है जो लिखित टेक्स्ट को विभिन्न भाषाओं, लहजों और क्षेत्रीय बोलियों में बोले गए ऑडियो में परिवर्तित करती है। यह बाहरी एप्लिकेशन में संश्लेषण क्षमताओं को एम्बेड करने के लिए एक प्रोग्रामेटिक इंटरफेस और एक फोनेटिक टेक्स्ट कनवर्टर दोनों के रूप में कार्य करता है जो लिखित टेक्स्ट को फोनेम कोड में अनुवादित करता है। यह सिस्टम कई संश्लेषण विधियों का उपयोग करता है, जिसमें गणितीय रूप से मुखर ध्वनियाँ उत्पन्न करने के लिए फॉर्मेंट संश्लेषण और पूर्व-रिकॉर्ड किए गए फोनेटिक सेगमेंट को जोड़कर ऑडियो बनाने के लिए डिफोन संश्लेषण शामिल है। इसमें ऑडियो पिच और टाइमिंग को नियंत्रित करने के लिए SSML और HTML टैग को पार्स करने में सक्षम एक स्पीच प्रोसेसर शामिल है। इंजन फोनेटिक अनुवाद मैप्स और परिभाषा फ़ाइलों के माध्यम से कस्टम वॉयस डिज़ाइन और भाषा उच्चारण अनुकूलन के लिए टूल्स प्रदान करता है। यह WAV प्रारूप में ऑडियो फ़ाइल एक्सपोर्ट, प्लेबैक गति समायोजन और भाषाई विश्लेषण के लिए फोनेटिक डेटा के निर्माण का समर्थन करता है। स्पीच जनरेशन को ट्रिगर करने और ऑडियो आउटपुट सेटिंग्स को प्रबंधित करने के लिए एक कमांड लाइन इंटरफेस उपलब्ध है।
Acts as a comprehensive engine for converting written text into spoken audio across various languages and dialects.
🎤 微软语音合成工具,使用 Electron Vue ElementPlus Vite 构建。
Converts plain text into spoken audio using Microsoft's synthesis engine, with automatic text slicing for long passages.
This is a collection of pre-trained neural models for speech recognition, synthesis, and voice activity detection. It provides a library of assets designed for speech-to-text, text-to-speech, and the identification of human speech segments within audio. The project features text-to-speech synthesis with support for multiple languages and the use of Speech Synthesis Markup Language to control prosody, pitch, and timing. For speech recognition, the system includes capabilities for transcribing audio to text with word-level timestamp extraction and an automated punctuation restorer to insert cap
Supports Speech Synthesis Markup Language (SSML) to provide precise control over pitch, timing, and prosody.
GLaDOS एक मल्टीमॉडल AI एजेंट फ्रेमवर्क है जिसे ऐसे ऑटोनॉमस सिस्टम बनाने के लिए डिज़ाइन किया गया है जो यूज़र्स और उनके वातावरण के साथ इंटरैक्ट करने के लिए टेक्स्ट, स्पीच और विज़ुअल डेटा को प्रोसेस करते हैं। यह एक AI पर्सनैलिटी फ्रेमवर्क पर केंद्रित है जो मल्टी-एजेंट आर्किटेक्चर और कॉन्फ़िगर करने योग्य व्यवहार प्रोफाइल का उपयोग करके जटिल कैरेक्टर पर्सना को एम्यूलेट करता है। यह प्रोजेक्ट एक इंटीग्रेटेड टूल लेयर के माध्यम से खुद को अलग करता है जो लैंग्वेज मॉडल्स को स्टैंडर्ड प्रोटोकॉल के ज़रिए बाहरी हार्डवेयर, स्मार्ट होम डिवाइसेस और सिस्टम APIs से जोड़ता है। इसमें लो-लेटेंसी प्लेबैक और इंटरप्शन हैंडलिंग के साथ एक कैरेक्टर टेक्स्ट-टू-स्पीच इंजन है, साथ ही एक मेमोरी और स्टेट मैनेजर है जो रिएक्टिव इमोशनल स्टेट्स को ट्रैक करता है और बातचीत में निरंतरता बनाए रखने के लिए लॉन्ग-टर्म तथ्यों को सेव रखता है। यह सिस्टम पर्यावरणीय समझ के लिए विज़न-लैंग्वेज परसेप्शन और ऑटोनॉमस एक्शन निष्पादन के लिए स्टेट-बेस्ड प्रोएक्टिव ट्रिगरिंग सहित कई क्षमताओं को कवर करता है। यह एजेंट आउटपुट को रीयल-टाइम में मॉनिटर और एडजस्ट करने के लिए एक कॉन्स्टिट्यूशनल बिहेवियर लेयर भी लागू करता है, जिससे पूर्व-निर्धारित व्यक्तित्व लक्षणों और दिशानिर्देशों का पालन सुनिश्चित होता है। सिस्टम में स्पीच रिकग्निशन, वॉयस सेटिंग्स और स्टेटस पैनल्स को मैनेज करने के लिए एक टर्मिनल कंट्रोल इंटरफ़ेस शामिल है।
Generates audible spoken responses using a variety of regional accents and gender-specific voice profiles.
PraisonAI is an autonomous AI agent platform that coordinates multiple LLM-powered agents for research, planning, and execution of complex workflows. It functions as a multi-agent orchestration framework, a workflow builder, and a Model Context Protocol server, while also providing retrieval-augmented generation through vector knowledge bases. Agents can interact via CLI, web, or standardized protocols with sandboxed code execution. The platform distinguishes itself with a rich set of agent communication protocols, including A2A, REST, WebSocket, voice and telephony integration, and MCP, allo
Converts written text into spoken audio using a language model and saves it as a file.
Bard-API Google Gemini के साथ बातचीत करने के लिए एक एसिंक्रोनस Python रैपर और क्लाइंट है। यह एक स्टेटफुल कन्वर्सेशन मैनेजर और मल्टीमॉडल इंटरफेस के रूप में कार्य करता है, जो उपयोगकर्ताओं को भाषा मॉडल में टेक्स्ट और इमेज प्रॉम्प्ट भेजने और प्रतिक्रिया प्राप्त करने की अनुमति देता है। यह लाइब्रेरी एक कुकी-आधारित प्रमाणीकरण प्रणाली का उपयोग करती है जो अनुरोधों को अधिकृत करने के लिए स्थानीय ब्राउज़र स्टोरेज से सत्र टोकन निकालती है। एक्सेस और कनेक्टिविटी का प्रबंधन करने के लिए, इसमें क्षेत्रीय प्रतिबंधों को बायपास करने और IP ब्लॉक से बचने के लिए प्रॉक्सी-आधारित अनुरोध रूटिंग शामिल है। यह प्रोजेक्ट मल्टीमॉडल AI विश्लेषण और निरंतर मल्टी-टर्न संवादों को सक्षम करने के लिए सत्र इतिहास के रखरखाव की क्षमताओं को कवर करता है। यह प्रतिक्रियाओं से इमेज लिंक निकालने, टेक्स्ट को स्पीच में बदलने और स्थानीय वातावरण के भीतर उत्पन्न कोड स्निपेट को स्वचालित रूप से निष्पादित करने के लिए उपयोगिताएँ भी प्रदान करता है।
Transforms written text strings into audio files using voice synthesis tools.
VITS-fast-fine-tuning छोटे ऑडियो डेटासेट का उपयोग करके विशिष्ट टारगेट आवाज़ों के लिए स्पीच सिंथेसिस मॉडल्स को अनुकूलित करने के लिए एक पाइपलाइन है। यह एक तेज़ स्पीकर अनुकूलन टूल और एक बहुभाषी स्पीच सिंथेसाइज़र के रूप में कार्य करता है जो विभिन्न भाषाओं में बोले गए ऑडियो को जनरेट करने में सक्षम है। यह सिस्टम मेनी-टू-मेनी वॉयस कन्वर्ज़न के लिए एक फ़्रेमवर्क प्रदान करता है, जो मूल भाषाई सामग्री को संरक्षित करते हुए एक स्पीकर की पहचान को दूसरे में बदल देता है। यह ऑडियो क्लिप्स या वीडियो स्रोतों के साथ एक प्री-ट्रेंड मॉडल को फ़ाइन-ट्यून करके टेक्स्ट-टू-स्पीच के लिए आवाज़ के अनुकूलन की अनुमति देता है। यह प्रोजेक्ट एंड-टू-एंड स्पीच सिंथेसिस और ऑडियो प्रोसेसिंग को कवर करता है, जो उच्च-निष्ठा (high-fidelity) ऑडियो उत्पन्न करने के लिए एडवरसैरियल वेवफ़ॉर्म जनरेशन और मोनोटोनिक अलाइनमेंट सर्च का उपयोग करता है। यह बोलने की लय में विविधताओं को प्रबंधित करने के लिए एक स्टोकेस्टिक ड्यूरेशन प्रेडिक्टर को शामिल करता है और प्री-ट्रेंड मॉडल ट्रांसफर का समर्थन करता है।
Supports synthesis of spoken audio across multiple languages while maintaining consistent character voices.
DiffSinger एक AI वोकल सिंथेसाइज़र और न्यूरल ऑडियो जनरेटर है जिसे हाई-फिडेलिटी सिंगिंग और स्पीच तैयार करने के लिए डिज़ाइन किया गया है। यह एक टेक्स्ट-टू-स्पीच सिस्टम और डिफ्यूजन-आधारित सिंगिंग वॉयस सिंथेसिस टूल के रूप में कार्य करता है जो टेक्स्ट और पिच को ऑडिबल ऑडियो में बदल देता है। यह सिस्टम यथार्थवादी वोकल परफॉरमेंस उत्पन्न करने के लिए शैलो डिफ्यूजन मैकेनिज्म और इटरेटिव नॉइज़ रिफाइनमेंट का उपयोग करता है। इसमें इन्फरेंस को तेज़ करने और सिंथेटिक आवाज़ें उत्पन्न करने में लगने वाले समय को कम करने के लिए विशेष सैंपलिंग प्लगइन्स और न्यूमेरिकल सॉल्वर्स शामिल हैं। यह प्रोजेक्ट टेक्स्ट को टाइम-डोमेन ऑडियो वेवफॉर्म में बदलने के लिए एकॉस्टिक मॉडलिंग, मेल-स्पेक्ट्रोग्राम सिंथेसिस और न्यूरल वोकोडर रिकंस्ट्रक्शन को कवर करता है। इसमें रिकॉर्डिंग की सोनिक क्वालिटी को बेहतर बनाने के लिए सिंथेटिक वोकल एन्हांसमेंट की क्षमताएं भी शामिल हैं।
Transforms written text into audible speech by predicting pitch and mel-spectrograms.
WhisperSpeech एक बहुभाषी स्पीच सिंथेसाइज़र और न्यूरल टेक्स्ट-टू-स्पीच सिस्टम है। यह टेक्स्ट को हाई-फिडेलिटी सिंथेटिक ऑडियो में बदलने के लिए Whisper मॉडल आर्किटेक्चर को इनवर्ट करके कार्य करता है। यह सिस्टम विशिष्ट वक्ताओं की नकल करने के लिए संदर्भ ऑडियो फ़ाइलों का उपयोग करके वॉयस क्लोनिंग को सक्षम बनाता है। यह बहुभाषी स्पीच प्रोडक्शन का समर्थन करता है, जिसमें विभिन्न भाषाओं में ऑडियो उत्पन्न करने और एक ही वाक्य के भीतर भाषा स्विचिंग को संभालने की क्षमता शामिल है। यह प्रोजेक्ट टेक्स्ट-टू-स्पीच जनरेशन और स्पीच डेटासेट तैयारी सहित स्पीच क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। इसमें स्पीच को टेक्स्ट में ट्रांसक्राइब करने, ध्वनिक टोकन निकालने, और वॉयस एक्टिविटी का पता लगाने के लिए टूल्स शामिल हैं।
Specializes in multilingual speech production with seamless mixing of multiple languages in one output.
This project is a comprehensive guide for configuring macOS for software engineering. It provides instructions for setting up a development environment, optimizing system settings, and enhancing the terminal experience to increase productivity. The guide focuses on several core areas of system customization, including the automation of software installation via Homebrew and the configuration of the Zsh shell with plugins, themes, and productivity aliases. It details a strategy for managing dotfiles using symbolic links to synchronize configuration files across multiple machines. Additional c
Enables the conversion of selected text or strings into spoken audio via the command line.
Kokoro-FastAPI is a text-to-speech API and LLM speech synthesis server that generates spoken audio from text via a REST interface. It functions as a Kubernetes-native deployment designed for orchestrated speech synthesis. The system includes a voice blending engine that creates unique vocal profiles by mixing multiple existing voices using custom weight ratios. The service provides real-time audio streaming to reduce latency and generates word-level timestamps for speech synchronization. It manages hardware efficiency through on-demand model loading to optimize VRAM usage and includes system
Provides a high-quality text-to-speech API for converting written text into spoken audio.
OpenSquilla is an LLM agent orchestration framework designed to coordinate multi-step AI workflows and tool execution using directed acyclic graphs. It functions as a centralized system for managing specialized skill packages and executing complex reasoning sequences. The project distinguishes itself through a routing gateway that directs tasks to different AI providers based on complexity, cost, and performance. It utilizes a multi-tier AI memory system that organizes working, episodic, and semantic knowledge using local embeddings and SQLite, alongside a secure execution sandbox that isolat
Transforms text responses into audio files using specialized media helper tools.
Abogen is a text-to-speech audiobook generator that transforms digital documents and subtitle files into audiobooks. It utilizes language models to perform text normalization, rewriting contractions and punctuation to produce more natural speech synthesis. The system features a voice profile mixer that blends multiple voice models using adjustable weight ratios to create personalized synthetic voices. It also includes an automated export system that sends completed audio files and metadata to a remote Audiobookshelf server via a web API. The project manages the end-to-end audiobook productio
Converts digital documents and subtitles into high-quality audio files with natural speech and timing control.
VDO.Ninja is a low-latency peer-to-peer media routing service and video streaming platform designed to integrate remote audio and video feeds into professional production workflows. It functions as a WebRTC broadcast integration tool and studio controller, allowing for the direct transmission of high-definition media between publishers and viewers with minimal delay. The platform distinguishes itself through extensive protocol bridging, converting between WebRTC, WHIP, WHEP, SRT, and RTMP to ensure compatibility across diverse network environments and professional studio software. It includes
Reads chat messages aloud using system voices or cloud APIs to provide audio feedback.