7 repositorios
Tools for adjusting voice parameters such as timbre, speed, language, and gender in text-to-speech systems.
Distinct from Speech Synthesis Caches: Existing candidates focus on caching, gateways, or specific languages rather than the parameterization of voice attributes.
Explore 7 awesome GitHub repositories matching artificial intelligence & ml · Speech Synthesis Customizations. Refine with filters or upvote what's useful.
F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s
Modifies the playback rate of generated audio by specifying a target duration for the output.
mi-gpt is a voice assistant bridge and agent orchestrator that connects smart speakers to large language models. It functions as an integration layer that routes audio requests from hardware speakers to AI providers and converts generated text back into speech via a customizable synthesis system. The project features a retrieval-augmented generation knowledge base that uses embeddings and external documents to provide context-aware responses. It includes a persona definition system for configuring behavioral rules, system prompts, and roleplay characteristics, alongside a plugin architecture
Allows replacing default system voices with high-quality models to customize the assistant's auditory profile.
This project provides an animated Live2D character widget that can be embedded on any web page as an interactive mascot. The widget renders characters using the Cubism SDK on an HTML canvas, and can be deployed either via a content delivery network for zero-setup integration or self-hosted on a personal server for full control over asset delivery. The mascot responds to visitor actions through CSS selector-based interaction binding, displaying custom speech bubbles when users hover over or click specific page elements. Visitors can click and drag the character to reposition it anywhere on the
Defines custom text messages that appear when visitors interact with specific page elements.
Audiblez is a text-to-speech audiobook generator that converts digital e-books into spoken audio files. The system processes written documents using speech synthesis and configurable voice profiles to produce audiobooks. The tool utilizes a graphical interface to manage the conversion workflow and task orchestration. It employs CUDA-accelerated processing to offload neural network computations to the GPU, increasing the speed of audio generation. The system includes capabilities for chapter-based file parsing and selective chapter conversion. Users can adjust synthesis parameters, including
Provides tools for adjusting voice parameters including speed, language, and voice identity.
Short video factory is a local AI content generator and automated video editing tool. It provides a production pipeline that uses large language models to transform text prompts into marketing scripts and rendered short-form videos. The system is designed for local-first execution, running all processing and asset management on the host machine to maintain data privacy. It distinguishes itself through a batch-processing workflow that can sequentially execute copywriting and rendering for multiple items using predefined presets. The software covers a broad range of media capabilities, includi
The software allows users to select language, gender, voice timbre, and speaking speed to customize audio narration.
TTS-WebUI is a web interface and speech synthesis manager designed to convert written text into spoken audio files. It serves as a self-hosted audio AI suite that allows users to configure speech synthesis models, manage speaker profiles, and generate audio through a graphical dashboard. The system functions as both a visual manager and a generative audio API, providing standardized endpoints and OpenAI-compatible request formats for external applications to trigger synthesis programmatically. It includes a plugin-based extension system that allows new tools and models to be added via externa
Features a dashboard for adjusting voice parameters, speaker profiles, and managing long-form text synthesis.
Esta aplicación es una plataforma para la síntesis de voz por IA y la clonación de voz neuronal. Proporciona un kit de herramientas integral para convertir texto en voz humana con sonido natural aplicando modelos de redes neuronales entrenados a medida a muestras de audio específicas. El sistema facilita todo el ciclo de vida del desarrollo de modelos de voz, incluyendo la preparación de audiolibros y transcripciones de video en conjuntos de datos de entrenamiento estructurados. Admite el entrenamiento de estos modelos en hardware local o remoto, utilizando procesamiento distribuido multi-GPU para manejar datos a gran escala y acelerar la convergencia del modelo. Más allá del entrenamiento, la plataforma incluye capacidades para gestionar y portar conjuntos de datos de voz a través de diferentes entornos de almacenamiento. Los usuarios pueden realizar inferencias ajustando variables latentes y parámetros de síntesis para modificar la prosodia, la inflexión emocional y las cualidades estilísticas de la salida de audio generada. La aplicación se basa en técnicas de deep learning para transformar representaciones acústicas en formas de onda de alta fidelidad.
Facilitates the management of voice datasets and the configuration of synthesis parameters to enable high-fidelity neural voice cloning.