awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
snakers4 avatar

snakers4/silero-models

0
View on GitHub↗
5,977 stars·366 forks·Jupyter Notebook·11 vues

Silero Models

Il s'agit d'une collection de modèles neuronaux pré-entraînés pour la reconnaissance vocale, la synthèse et la détection d'activité vocale. Elle fournit une bibliothèque d'actifs conçus pour la conversion parole-texte, la synthèse texte-parole et l'identification de segments de parole humaine au sein de l'audio.

Le projet propose une synthèse texte-parole avec prise en charge de plusieurs langues et l'utilisation du langage de balisage de synthèse vocale (SSML) pour contrôler la prosodie, la hauteur et le timing. Pour la reconnaissance vocale, le système inclut des capacités de transcription audio en texte avec extraction d'horodatage au niveau du mot et un restaurateur de ponctuation automatisé pour insérer les majuscules et la ponctuation dans le texte brut.

Les modèles sont exportés au format Open Neural Network Exchange et TorchScript pour permettre une exécution haute performance sur différents accélérateurs matériels et systèmes d'exploitation.

Features

  • Automatic Speech Recognition - Provides pre-trained neural models for converting spoken audio into written text.
  • Text-to-Speech Synthesis - Provides pre-trained neural models that convert written text into natural sounding spoken audio.
  • ONNX Model Exporters - Exports neural models to the standardized ONNX format for high-performance execution across various hardware.
  • ONNX Model Exports - Delivers speech models exported to ONNX format for cross-engine compatibility and performance.
  • ONNX Runtime Inference - Enables the execution of audio models across different operating systems using the ONNX runtime.
  • TorchScript Exports - Packages models using TorchScript to enable high-performance inference without requiring a Python runtime.
  • Pre-trained Speech Models - Provides a comprehensive library of pre-trained neural models for speech recognition, synthesis, and activity detection.
  • Punctuation Restoration - Includes a pre-trained model that automatically inserts capitalization and punctuation into raw speech transcripts.
  • Multilingual Synthesis - Provides synthesis models capable of generating speech across a wide variety of regional and minority languages.
  • Speech to Text Transcription - Converts spoken audio into text using models optimized for performance and size.
  • Voice Activity Detection - Includes a model for detecting human speech boundaries to isolate voice from background noise.
  • Word-Level Timestamps - Provides precise start and end times for individual words during the audio transcription process.
  • Prosody Control - Allows adjustment of pitch, rate, and timing of synthesized speech to achieve natural inflection.
  • Speech Prosody Generation - Automatically applies word stress and resolves homographs to improve the natural pronunciation of spoken language.
  • SSML Conversions - Supports Speech Synthesis Markup Language (SSML) to provide precise control over pitch, timing, and prosody.
  • Phoneme-Based Speech Processors - Implements a synthesis pipeline that converts text into phonetic representations to ensure natural pronunciation.
  • Speech Synthesis Markup Controls - Integrates markup controls to refine the prosody and pronunciation of synthesized speech.

Historique des stars

Graphique de l'historique des stars pour snakers4/silero-modelsGraphique de l'historique des stars pour snakers4/silero-models

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Collections incluant Silero Models

Sélections manuelles où Silero Models apparaît.
  • Modèles de traduction multilingue locaux
  • Modèles de synthèse et de reconnaissance vocale
  • Modèles de machine learning open source

Questions fréquentes

Que fait snakers4/silero-models ?

Il s'agit d'une collection de modèles neuronaux pré-entraînés pour la reconnaissance vocale, la synthèse et la détection d'activité vocale. Elle fournit une bibliothèque d'actifs conçus pour la conversion parole-texte, la synthèse texte-parole et l'identification de segments de parole humaine au sein de l'audio.

Quelles sont les fonctionnalités principales de snakers4/silero-models ?

Les fonctionnalités principales de snakers4/silero-models sont : Automatic Speech Recognition, Text-to-Speech Synthesis, ONNX Model Exporters, ONNX Model Exports, ONNX Runtime Inference, TorchScript Exports, Pre-trained Speech Models, Punctuation Restoration.

Quelles sont les alternatives open-source à snakers4/silero-models ?

Les alternatives open-source à snakers4/silero-models incluent : modelscope/funasr — FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken… elevenlabs/elevenlabs-python — This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of… livekit/agents — This project is a framework for developing multimodal AI agents that function as programmable participants in… k2-fsa/sherpa-onnx — Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device… fishaudio/bert-vits2 — Bert-VITS2 is a neural speech synthesis system and AI voice generator designed to convert written text into natural… ohf-voice/piper1-gpl — This project is a neural text-to-speech system and voice trainer that converts written text into spoken audio across a…

Alternatives open source à Silero Models

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Silero Models.
  • modelscope/funasrAvatar de modelscope

    modelscope/FunASR

    18,481Voir sur GitHub↗

    FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken audio into written text across more than fifty languages. It provides a framework for speaker diarization, an OpenAI-compatible transcription API for local server hosting, and speech models compatible with the ONNX format. The project distinguishes itself by supporting high-performance inference on edge hardware via self-contained binaries and portable model exports. It incorporates specialized capabilities for natural speech generation with adjustable timbre and emotional expre

    Pythonasraudiochinese
    Voir sur GitHub↗18,481
  • elevenlabs/elevenlabs-pythonAvatar de elevenlabs

    elevenlabs/elevenlabs-python

    2,873Voir sur GitHub↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    Voir sur GitHub↗2,873
  • livekit/agentsAvatar de livekit

    livekit/agents

    9,379Voir sur GitHub↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Pythonagentsaiopenai
    Voir sur GitHub↗9,379
  • k2-fsa/sherpa-onnxAvatar de k2-fsa

    k2-fsa/sherpa-onnx

    13,017Voir sur GitHub↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    C++aarch64androidarm32
    Voir sur GitHub↗13,017
Voir les 30 alternatives à Silero Models→