awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
snakers4 avatar

snakers4/silero-models

0
View on GitHub↗
5,977 स्टार्स·366 फोर्क्स·Jupyter Notebook·11 व्यूज़

Silero Models

This is a collection of pre-trained neural models for speech recognition, synthesis, and voice activity detection. It provides a library of assets designed for speech-to-text, text-to-speech, and the identification of human speech segments within audio.

The project features text-to-speech synthesis with support for multiple languages and the use of Speech Synthesis Markup Language to control prosody, pitch, and timing. For speech recognition, the system includes capabilities for transcribing audio to text with word-level timestamp extraction and an automated punctuation restorer to insert capitalization and punctuation into raw text.

The models are exported to the Open Neural Network Exchange format and TorchScript to enable high-performance execution across different hardware accelerators and operating systems.

Features

  • Automatic Speech Recognition - Provides pre-trained neural models for converting spoken audio into written text.
  • Text-to-Speech Synthesis - Provides pre-trained neural models that convert written text into natural sounding spoken audio.
  • ONNX Model Exporters - Exports neural models to the standardized ONNX format for high-performance execution across various hardware.
  • ONNX Model Exports - Delivers speech models exported to ONNX format for cross-engine compatibility and performance.
  • ONNX Runtime Inference - Enables the execution of audio models across different operating systems using the ONNX runtime.
  • TorchScript Exports - Packages models using TorchScript to enable high-performance inference without requiring a Python runtime.
  • Pre-trained Speech Models - Provides a comprehensive library of pre-trained neural models for speech recognition, synthesis, and activity detection.
  • Punctuation Restoration - Includes a pre-trained model that automatically inserts capitalization and punctuation into raw speech transcripts.
  • Multilingual Synthesis - Provides synthesis models capable of generating speech across a wide variety of regional and minority languages.
  • Speech to Text Transcription - Converts spoken audio into text using models optimized for performance and size.
  • Voice Activity Detection - Includes a model for detecting human speech boundaries to isolate voice from background noise.
  • Word-Level Timestamps - Provides precise start and end times for individual words during the audio transcription process.
  • Prosody Control - Allows adjustment of pitch, rate, and timing of synthesized speech to achieve natural inflection.
  • Speech Prosody Generation - Automatically applies word stress and resolves homographs to improve the natural pronunciation of spoken language.
  • SSML Conversions - Supports Speech Synthesis Markup Language (SSML) to provide precise control over pitch, timing, and prosody.
  • Phoneme-Based Speech Processors - Implements a synthesis pipeline that converts text into phonetic representations to ensure natural pronunciation.
  • Speech Synthesis Markup Controls - Integrates markup controls to refine the prosody and pronunciation of synthesized speech.

स्टार हिस्ट्री

snakers4/silero-models के लिए स्टार हिस्ट्री चार्टsnakers4/silero-models के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

Silero Models को शामिल करने वाली क्यूरेटेड खोजें

चुनिंदा कलेक्शन जहाँ Silero Models दिखाई देता है।
  • लोकल मल्टीलिंगुअल ट्रांसलेशन मॉडल्स
  • स्पीच सिंथेसिस और रिकग्निशन मॉडल्स
  • ओपन-सोर्स मशीन लर्निंग मॉडल

Silero Models के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Silero Models के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • modelscope/funasrmodelscope का अवतार

    modelscope/FunASR

    18,481GitHub पर देखें↗

    FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken audio into written text across more than fifty languages. It provides a framework for speaker diarization, an OpenAI-compatible transcription API for local server hosting, and speech models compatible with the ONNX format. The project distinguishes itself by supporting high-performance inference on edge hardware via self-contained binaries and portable model exports. It incorporates specialized capabilities for natural speech generation with adjustable timbre and emotional expre

    Pythonasraudiochinese
    GitHub पर देखें↗18,481
  • elevenlabs/elevenlabs-pythonelevenlabs का अवतार

    elevenlabs/elevenlabs-python

    2,873GitHub पर देखें↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    GitHub पर देखें↗2,873
  • livekit/agentslivekit का अवतार

    livekit/agents

    9,379GitHub पर देखें↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Pythonagentsaiopenai
    GitHub पर देखें↗9,379
  • k2-fsa/sherpa-onnxk2-fsa का अवतार

    k2-fsa/sherpa-onnx

    13,017GitHub पर देखें↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    C++aarch64androidarm32
    GitHub पर देखें↗13,017
  • Silero Models के सभी 30 विकल्प देखें→

    अक्सर पूछे जाने वाले प्रश्न

    snakers4/silero-models क्या करता है?

    This is a collection of pre-trained neural models for speech recognition, synthesis, and voice activity detection. It provides a library of assets designed for speech-to-text, text-to-speech, and the identification of human speech segments within audio.

    snakers4/silero-models की मुख्य विशेषताएं क्या हैं?

    snakers4/silero-models की मुख्य विशेषताएं हैं: Automatic Speech Recognition, Text-to-Speech Synthesis, ONNX Model Exporters, ONNX Model Exports, ONNX Runtime Inference, TorchScript Exports, Pre-trained Speech Models, Punctuation Restoration।

    snakers4/silero-models के कुछ ओपन-सोर्स विकल्प क्या हैं?

    snakers4/silero-models के ओपन-सोर्स विकल्पों में शामिल हैं: modelscope/funasr — FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken… elevenlabs/elevenlabs-python — This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of… livekit/agents — This project is a framework for developing multimodal AI agents that function as programmable participants in… k2-fsa/sherpa-onnx — Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device… fishaudio/bert-vits2 — Bert-VITS2 is a neural speech synthesis system and AI voice generator designed to convert written text into natural… ohf-voice/piper1-gpl — This project is a neural text-to-speech system and voice trainer that converts written text into spoken audio across a…