How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
DaCiDian is an open-sourced chinese mandarin lexicon for automatic speech recognition(ASR)
A wrapper around speech quality metrics MOSNet, BSSEval, STOI, PESQ, SRMR, SISDR
🤖💬 Transformer TTS: Implementation of a non-autoregressive Transformer based neural network for text to speech.
ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles. The system is distinguished by its ability to generate interactive dialogue with realistic vocal nuances. It utilizes a speech nuance controller to insert specific tokens that trigger non-verbal elements, such as laughter, pauses, and interjections, during the synthesis process. The project includes a streaming audio generato
中文语音识别; Mandarin Automatic Speech Recognition;
The main features of lukhy/masr are: Speech Processing.
Projects with overlapping indexed features include: aishell-foundation/dacidian — DaCiDian is an open-sourced chinese mandarin lexicon for automatic speech recognition(ASR). aliutkus/speechmetrics — A wrapper around speech quality metrics MOSNet, BSSEval, STOI, PESQ, SRMR, SISDR. as-ideas/transformertts — 🤖💬 Transformer TTS: Implementation of a non-autoregressive Transformer based neural network for text to speech. boson-ai/higgs-audio — Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large… bytedance/megatts3 — MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English,… 2noise/chattts — ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding…