awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
snakers4 avatar

snakers4/silero-models

0
View on GitHub↗
5,977 stars·366 forks·Jupyter Notebook·21 views

Silero Models

This is a collection of pre-trained neural models for speech recognition, synthesis, and voice activity detection. It provides a library of assets designed for speech-to-text, text-to-speech, and the identification of human speech segments within audio.

The project features text-to-speech synthesis with support for multiple languages and the use of Speech Synthesis Markup Language to control prosody, pitch, and timing. For speech recognition, the system includes capabilities for transcribing audio to text with word-level timestamp extraction and an automated punctuation restorer to insert capitalization and punctuation into raw text.

The models are exported to the Open Neural Network Exchange format and TorchScript to enable high-performance execution across different hardware accelerators and operating systems.

Features

  • Automatic Speech Recognition - Provides pre-trained neural models for converting spoken audio into written text.
  • Text-to-Speech Synthesis - Provides pre-trained neural models that convert written text into natural sounding spoken audio.
  • ONNX Model Exporters - Exports neural models to the standardized ONNX format for high-performance execution across various hardware.
  • ONNX Model Exports - Delivers speech models exported to ONNX format for cross-engine compatibility and performance.
  • ONNX Runtime Inference - Enables the execution of audio models across different operating systems using the ONNX runtime.
  • TorchScript Exports - Packages models using TorchScript to enable high-performance inference without requiring a Python runtime.
  • Pre-trained Speech Models - Provides a comprehensive library of pre-trained neural models for speech recognition, synthesis, and activity detection.
  • Punctuation Restoration - Includes a pre-trained model that automatically inserts capitalization and punctuation into raw speech transcripts.
  • Multilingual Synthesis - Provides synthesis models capable of generating speech across a wide variety of regional and minority languages.
  • Speech to Text Transcription - Converts spoken audio into text using models optimized for performance and size.
  • Voice Activity Detection - Includes a model for detecting human speech boundaries to isolate voice from background noise.
  • Word-Level Timestamps - Provides precise start and end times for individual words during the audio transcription process.
  • Prosody Control - Allows adjustment of pitch, rate, and timing of synthesized speech to achieve natural inflection.
  • Speech Prosody Generation - Automatically applies word stress and resolves homographs to improve the natural pronunciation of spoken language.
  • SSML Conversions - Supports Speech Synthesis Markup Language (SSML) to provide precise control over pitch, timing, and prosody.
  • Phoneme-Based Speech Processors - Implements a synthesis pipeline that converts text into phonetic representations to ensure natural pronunciation.
  • Speech Synthesis Markup Controls - Integrates markup controls to refine the prosody and pronunciation of synthesized speech.

Star history

Star history chart for snakers4/silero-modelsStar history chart for snakers4/silero-models

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does snakers4/silero-models do?

This is a collection of pre-trained neural models for speech recognition, synthesis, and voice activity detection. It provides a library of assets designed for speech-to-text, text-to-speech, and the identification of human speech segments within audio.

What are the main features of snakers4/silero-models?

The main features of snakers4/silero-models are: Automatic Speech Recognition, Text-to-Speech Synthesis, ONNX Model Exporters, ONNX Model Exports, ONNX Runtime Inference, TorchScript Exports, Pre-trained Speech Models, Punctuation Restoration.

Which projects share features with snakers4/silero-models?

Projects with overlapping indexed features include: modelscope/funasr — FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken… elevenlabs/elevenlabs-python — This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of… livekit/agents — This project is a framework for developing multimodal AI agents that function as programmable participants in… k2-fsa/sherpa-onnx — Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device… fishaudio/bert-vits2 — Bert-VITS2 is a neural speech synthesis system and AI voice generator designed to convert written text into natural… ohf-voice/piper1-gpl — This project is a neural text-to-speech system and voice trainer that converts written text into spoken audio across a…

Projects sharing features with Silero Models

These projects share indexed features with Silero Models. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • modelscope/funasrmodelscope avatar

    modelscope/FunASR

    18,481View on GitHub↗

    FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken audio into written text across more than fifty languages. It provides a framework for speaker diarization, an OpenAI-compatible transcription API for local server hosting, and speech models compatible with the ONNX format. The project distinguishes itself by supporting high-performance inference on edge hardware via self-contained binaries and portable model exports. It incorporates specialized capabilities for natural speech generation with adjustable timbre and emotional expre

    Pythonasraudiochinese
    View on GitHub↗18,481
  • elevenlabs/elevenlabs-pythonelevenlabs avatar

    elevenlabs/elevenlabs-python

    2,873View on GitHub↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    View on GitHub↗2,873
  • livekit/agentslivekit avatar

    livekit/agents

    9,379View on GitHub↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Pythonagentsaiopenai
    View on GitHub↗9,379
  • k2-fsa/sherpa-onnxk2-fsa avatar

    k2-fsa/sherpa-onnx

    13,017View on GitHub↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    C++aarch64androidarm32
    View on GitHub↗13,017
Compare all 30 related projects→

Curated searches featuring Silero Models

Hand-picked collections where Silero Models appears.
  • Local Multilingual Translation Models
  • Speech Synthesis and Recognition Models
  • open-source machine learning model