awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Streaming speech recognition

रैंकिंग 23 जुल॰ 2026 को अपडेट की गई

For a library for streaming speech recognition, the strongest matches are alphacep/vosk-api (Vosk is a speech recognition library designed for offline), modelscope/funasr (FunASR is a comprehensive speech recognition toolkit that provides) and ggml-org/whisper.cpp (Whisper). ggerganov/whisper.cpp and paddlepaddle/paddlespeech round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Hand-picked streaming speech recognition open-source libraries, ranked by stars and activity. Compare the top options and find the best fit.

Streaming speech recognition

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • alphacep/vosk-apialphacep का अवतार

    alphacep/vosk-api

    14,853GitHub पर देखें↗

    Vosk is an offline speech-to-text engine and API that converts spoken audio into text locally on a device. It provides a cross-platform speech toolkit with language bindings for integrating voice recognition into server environments, Android, iOS, and Raspberry Pi. The project includes a speaker identification tool to distinguish between different voices and an acoustic model trainer for building custom neural network models. These training tools enable speech feature extraction and model accuracy evaluation to improve recognition for specialized domains. The system supports real-time audio

    Vosk is a speech recognition library designed for offline, real-time transcription across multiple platforms, covering all the core capabilities and features requested.

    Jupyter NotebookCustom VocabulariesSpeech TranscriptionSpeech Recognition
    GitHub पर देखें↗14,853
  • modelscope/funasrmodelscope का अवतार

    modelscope/FunASR

    18,481GitHub पर देखें↗

    FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken audio into written text across more than fifty languages. It provides a framework for speaker diarization, an OpenAI-compatible transcription API for local server hosting, and speech models compatible with the ONNX format. The project distinguishes itself by supporting high-performance inference on edge hardware via self-contained binaries and portable model exports. It incorporates specialized capabilities for natural speech generation with adjustable timbre and emotional expre

    FunASR is a comprehensive speech recognition toolkit that provides real-time streaming transcription, multilingual support, offline execution, and speaker diarization, which directly matches your search for a real-time speech-to-text framework.

    PythonAutomatic Speech RecognitionSpeaker DiarizationSpeech Transcription
    GitHub पर देखें↗18,481
  • ggml-org/whisper.cppggml-org का अवतार

    ggml-org/whisper.cpp

    50,770GitHub पर देखें↗

    Whisper.cpp is a high-performance, local-first speech recognition engine designed to run large-scale machine learning models on consumer hardware. It functions as a portable library that converts audio into text, supporting both static file transcription and real-time stream processing. By utilizing a lightweight inference engine and weight quantization, the project minimizes memory and compute overhead, allowing for efficient execution without reliance on external cloud APIs or internet connectivity. The project distinguishes itself through a hardware-agnostic compute abstraction that offloa

    Whisper.cpp is a local-first speech recognition engine that provides high-performance speech-to-text capabilities and real-time stream processing as a portable C++ library, though it lacks built-in speaker diarization.

    C++Local Inference EnginesSpeaker DiarizationSpeech Transcription
    GitHub पर देखें↗50,770
  • ggerganov/whisper.cppggerganov का अवतार

    ggerganov/whisper.cpp

    50,791GitHub पर देखें↗

    whisper.cpp is a C++ implementation of the Whisper speech-to-text model, serving as a lightweight machine learning inference engine and quantized runtime. It provides high-performance automatic speech recognition and real-time audio transcription without requiring a Python environment. The project utilizes model quantization to reduce memory usage and increase inference speed on local hardware. It incorporates hardware acceleration to optimize processing speed across different processors. The system covers audio processing capabilities including voice activity detection, speaker diarization,

    This project is a high-performance C++ inference engine for real-time automatic speech recognition and speaker diarization, which directly matches your search for a speech recognition library supporting offline execution and streaming.

    C++Automatic Speech RecognitionSpeaker Diarization
    GitHub पर देखें↗50,791
  • paddlepaddle/paddlespeechPaddlePaddle का अवतार

    PaddlePaddle/PaddleSpeech

    12,626GitHub पर देखें↗

    PaddleSpeech is a comprehensive toolkit of neural models for speech recognition, synthesis, and translation built on the PaddlePaddle deep learning framework. It provides a collection of frameworks and tools for converting spoken audio into written text, synthesizing natural audio from text, and performing direct speech translation. The toolkit includes specialized capabilities for keyword spotting to detect trigger words and speaker verification systems that extract unique voiceprints to identify and distinguish between individuals. It also features end-to-end translation tools that map audi

    PaddleSpeech is a deep learning toolkit built on PaddlePaddle that provides automatic speech recognition, streaming transcription, and speaker diarization, serving as a comprehensive framework for speech-to-text applications.

    PythonAutomatic Speech RecognitionSpeaker DiarizationSpeech Transcription
    GitHub पर देखें↗12,626
  • ahmetoner/whisper-asr-webserviceahmetoner का अवतार

    ahmetoner/whisper-asr-webservice

    3,286GitHub पर देखें↗

    This project provides a self-hosted server for automatic speech recognition, functioning as a containerized inference engine for the Whisper model. It exposes core transcription and translation capabilities through a standardized web interface, allowing for the integration of speech-to-text services into external applications. The service distinguishes itself by incorporating advanced audio analysis tools, including speaker diarization to attribute text to specific individuals and voice activity detection to filter non-speech segments. It supports automated language detection and provides out

    This project provides a self-hosted automatic speech recognition server wrapper around the Whisper model rather than a direct programming library, but it natively delivers real-time-friendly API endpoints, multilingual support, offline execution, and speaker diarization.

    PythonAutomatic Speech RecognitionSpeaker DiarizationSpeech Transcription
    GitHub पर देखें↗3,286
  • nvidia/nemoNVIDIA का अवतार

    NVIDIA/NeMo

    17,394GitHub पर देखें↗

    NeMo is a multimodal AI framework and toolkit designed for the development, training, and scaling of large language models, generative AI systems, and speech-based models. It functions as an automatic speech recognition toolkit, a text-to-speech engine, and a framework for building models that process and generate combinations of text, image, and audio data. The project serves as a conversational AI orchestrator capable of managing real-time, interruptible voice interactions. It provides specialized workflows for speech translation, converting spoken audio from one language into text or speec

    NeMo is a multimodal AI framework and toolkit that provides robust automatic speech recognition, real-time streaming transcription, multilingual support, and offline execution capabilities for building speech-based applications.

    PythonAutomatic Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition Toolkits
    GitHub पर देखें↗17,394
  • pannous/tensorflow-speech-recognitionpannous का अवतार

    pannous/tensorflow-speech-recognition

    2,172GitHub पर देखें↗

    This library provides a deep learning framework for training neural networks to perform speech recognition and audio classification. It utilizes sequence-to-sequence architectures to map variable-length audio inputs into text or numerical outputs, enabling the development of custom speech-to-text transcription models. The project distinguishes itself through integrated audio processing capabilities that transform raw waveforms into spectrograms and high-dimensional numerical vectors. These tools allow for the extraction of unique vocal characteristics to identify speakers, as well as the clas

    This library provides a deep learning framework for training custom speech recognition models using TensorFlow, though it focuses on training and feature extraction rather than offering a ready-to-use real-time streaming transcription engine.

    PythonAutomatic Speech RecognitionSpeech Recognition LibrariesSpeech Transcription
    GitHub पर देखें↗2,172
  • m-bain/whisperxm-bain का अवतार

    m-bain/whisperX

    20,228GitHub पर देखें↗

    WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts. The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi

    WhisperX is an automated speech recognition toolkit providing accurate transcription, alignment, and speaker diarization, though it focuses more on post-processing and timestamp alignment than continuous low-latency real-time streaming.

    PythonAutomatic Speech RecognitionSpeaker DiarizationSpeech Transcription
    GitHub पर देखें↗20,228
  • kaldi-asr/kaldikaldi-asr का अवतार

    kaldi-asr/kaldi

    15,415GitHub पर देखें↗

    Kaldi is an automatic speech recognition toolkit used to train and deploy models that convert spoken audio into text. It functions as a framework for designing and evaluating acoustic and language models through a structured pipeline of processing tools. The system acts as a cross-platform speech engine, capable of compiling recognition logic for Android and WebAssembly to enable execution on mobile devices and web browsers. It also includes a dedicated converter for migrating speech recognition models from the HTK format into a compatible internal structure. The toolkit covers a broad range

    Kaldi is a powerful open-source speech recognition toolkit providing foundational frameworks for automatic speech recognition, though you may need to integrate additional components for modern real-time streaming or speaker diarization pipelines.

    ShellAutomatic Speech RecognitionSpeech Recognition Systems
    GitHub पर देखें↗15,415
  • facebookresearch/wav2letterfacebookresearch का अवतार

    facebookresearch/wav2letter

    6,444GitHub पर देखें↗

    wav2letter is an automatic speech recognition toolkit and deep learning framework designed to convert audio speech signals into written text. It functions as a distributed training system and an inference engine for building and deploying neural network architectures. The system enables the training of large-scale speech models across multiple compute nodes using custom architecture files and structured recipes. It includes an inference engine that allows these trained models to be executed within Python workflows to transform audio sequences into text. The framework covers the full speech r

    This automatic speech recognition toolkit covers model training and inference pipelines in C++, fitting the requirement for a speech recognition framework despite lacking explicit mention of features like real-time streaming or diarization.

    C++Automatic Speech RecognitionSpeech-to-Text EnginesASR Frameworks
    GitHub पर देखें↗6,444
  • mozilla/deepspeechmozilla का अवतार

    mozilla/DeepSpeech

    26,748GitHub पर देखें↗

    DeepSpeech is an open-source speech-to-text framework and machine learning engine designed to convert spoken audio into written text locally on a device. It provides on-device speech recognition that operates without requiring an internet connection to external servers. The system supports real-time speech transcription across a variety of hardware platforms, ranging from single-board computers and edge devices to GPU servers. This allows for audio analysis and processing directly on the local hardware.

    Mozilla DeepSpeech is an open-source speech-to-text framework designed for on-device automatic speech recognition and real-time streaming transcription, though it lacks built-in speaker diarization.

    C++Speech-to-Text EnginesSpeech Recognition
    GitHub पर देखें↗26,748
  • argmaxinc/whisperkitargmaxinc का अवतार

    argmaxinc/WhisperKit

    5,639GitHub पर देखें↗

    WhisperKit is a Swift framework for running Whisper speech recognition models locally on Apple platforms, supporting offline execution, automatic speech recognition, and streaming transcription features suited for mobile applications.

    SwiftCustom VocabulariesSpeaker DiarizationAudio Transcription
    GitHub पर देखें↗5,639
  • nl8590687/asrt_speechrecognitionnl8590687 का अवतार

    nl8590687/ASRT_SpeechRecognition

    8,375GitHub पर देखें↗

    This project is a Chinese automatic speech recognition framework and deep learning system designed to convert spoken Chinese audio into written text. It functions as a toolkit for training, evaluating, and deploying speech-to-text models, utilizing a specialized pinyin-to-text converter that transforms phonetic sequences into Chinese characters using a probability graph model. The system is distinguished by its deployment flexibility, offering a dockerized recognition server that provides transcription capabilities as a remote API. It supports high-performance streaming through a gRPC speech-

    This project is an automatic speech recognition framework that supports real-time streaming transcription and custom model training, making it a relevant tool for speech-to-text tasks even though it is specifically optimized for the Chinese language.

    PythonBidirectional Speech-to-Text StreamsAsynchronous Speech-to-Text StreamsSpeech Recognition Services
    GitHub पर देखें↗8,375
  • k2-fsa/sherpa-onnxk2-fsa का अवतार

    k2-fsa/sherpa-onnx

    13,017GitHub पर देखें↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    Sherpa-ONNX is a cross-platform speech processing toolkit for running automatic speech recognition, text-to-speech, and speaker diarization locally, providing the core speech-to-text functionality and offline execution the visitor needs.

    C++Speaker DiarizationAudio TranscriptionLocal Inference
    GitHub पर देखें↗13,017
  • openai/whisperopenai का अवतार

    openai/whisper

    102,828GitHub पर देखें↗

    This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer architecture to convert audio into text. It is built upon a weakly supervised learning framework, which leverages large-scale, unlabelled audio-transcript data to create generalized speech representations capable of performing simultaneous transcription, language identification, and translation. The system distinguishes itself through a unified multi-task modeling approach that shares token sequences across different objectives, allowing it to handle diverse languages and vocabularies

    Whisper provides a powerful transformer-based architecture for speech-to-text transcription and automatic speech recognition across multiple languages, though it primarily operates in batch mode rather than supporting real-time streaming out of the box.

    PythonAutomatic Speech RecognitionSpeech Recognition LibrariesAutomatic Speech Recognition Toolkits
    GitHub पर देखें↗102,828
  • davabase/whisper_real_timedavabase का अवतार

    davabase/whisper_real_time

    2,938GitHub पर देखें↗

    Whisper Real-Time is a speech-to-text engine designed to convert continuous microphone input into written transcripts. It functions as a real-time audio processor that leverages the OpenAI Whisper model to generate immediate textual output from live spoken language. The system utilizes a transformer-based architecture to map audio sequences to text tokens. It manages incoming data through a sliding-window buffering mechanism and a circular buffer, which ensures a steady stream of audio for the inference engine. To maintain accuracy during continuous processing, the software employs a stateful

    Whisper Real-Time is a Python-based speech-to-text engine that processes live microphone input for continuous transcription, matching the core search intent even though specific features like speaker diarization and custom vocabulary are not highlighted.

    PythonSpeech-to-Text Engines
    GitHub पर देखें↗2,938
  • guillaumekln/faster-whisperguillaumekln का अवतार

    guillaumekln/faster-whisper

    23,679GitHub पर देखें↗

    faster-whisper is an automatic speech recognition framework and an optimized implementation of the Whisper speech-to-text engine. It functions as a CTranslate2 inference engine designed to convert spoken audio into written text. The project serves as a model quantization tool that transforms large audio model weights into lower precision formats. This process reduces memory usage and increases execution speed on hardware by utilizing integer quantized weights. The framework covers a broad range of capabilities including batch audio transcription for parallel processing and voice activity det

    Faster-whisper is an optimized automatic speech recognition framework that delivers fast speech-to-text transcription, though it requires additional wrapper code to handle real-time streaming and speaker diarization out of the box.

    PythonAutomatic Speech RecognitionSpeech-to-Text Engines
    GitHub पर देखें↗23,679
  • huggingface/distil-whisperhuggingface का अवतार

    huggingface/distil-whisper

    4,084GitHub पर देखें↗

    Distil-Whisper is a compressed automatic speech recognition model designed to convert spoken audio into written text. It uses a transformer-based sequence-to-sequence architecture to provide speech-to-text transcription. The project utilizes knowledge distillation and teacher-student model compression to create a lightweight version of the Whisper model. This approach reduces GPU memory usage and accelerates token prediction while maintaining transcription accuracy. The system supports both short-form audio transcription and long-form processing through the use of sliding windows and chunked

    Distil-Whisper is an automatic speech recognition model that provides speech-to-text transcription with accelerated inference, though it is primarily a compressed model weight rather than a full framework with built-in speaker diarization.

    PythonAutomatic Speech RecognitionAudio TranscriptionSpeech Recognition Models
    GitHub पर देखें↗4,084
  • soniqo/speech-swiftsoniqo का अवतार

    soniqo/speech-swift

    896GitHub पर देखें↗

    This project is a comprehensive toolkit for on-device speech recognition, synthesis, and audio processing, specifically engineered for Apple Silicon. It provides a framework for building real-time, full-duplex voice agents that operate entirely offline, leveraging native hardware acceleration to maintain performance and privacy. By utilizing optimized machine learning models, the library enables local execution of complex audio tasks without reliance on external cloud services. The library distinguishes itself through its specialized focus on local, high-performance voice interaction. It incl

    This library provides on-device automatic speech recognition and real-time streaming capabilities tailored for Apple platforms, though it focuses more broadly on full-duplex voice agents rather than acting solely as a general-purpose transcription library.

    SwiftSpeaker DiarizationSpeech Recognition Libraries
    GitHub पर देखें↗896
  • funaudiollm/sensevoiceFunAudioLLM का अवतार

    FunAudioLLM/SenseVoice

    7,536GitHub पर देखें↗

    SenseVoice is a multilingual speech large language model designed for audio transcription, speaker diarization, and emotion recognition. It functions as an automatic speech recognition system that converts spoken audio into text across multiple languages. The system distinguishes itself by integrating acoustic event detection and speech emotion recognition, allowing it to identify non-speech sounds, such as laughter or applause, and discrete emotional states. It also includes a framework for speaker diarization to track and label different speakers within a single recording. The project's ca

    SenseVoice is an open-source speech recognition model and framework that supports multilingual transcription and speaker diarization, though it focuses more on speech understanding models than a traditional streaming library.

    PythonAutomatic Speech RecognitionSpeaker Diarization
    GitHub पर देखें↗7,536
  • jamiepine/voiceboxjamiepine का अवतार

    jamiepine/voicebox

    30,041GitHub पर देखें↗

    Voicebox is a local speech processing system that provides text-to-speech generation, speech-to-text transcription, and voice cloning. It utilizes local machine learning inference and GPU acceleration to process audio and text data without relying on external API calls. The project features a voice cloning toolkit for creating synthetic profiles from audio samples and a timeline-based voice editor for composing multi-character conversations. It also includes an AI voice management API that allows external applications and AI agents to programmatically manage voice profiles and generate speech

    Voicebox is a local speech processing system that provides speech-to-text transcription and text-to-speech generation, fitting the requirement for an offline speech recognition library despite its broader focus on voice cloning and editing.

    TypeScriptLocal Inference EnginesAudio TranscriptionLocal AI Inference
    GitHub पर देखें↗30,041
  • const-me/whisperConst-me का अवतार

    Const-me/Whisper

    10,489GitHub पर देखें↗

    Whisper is a high-performance speech-to-text inference engine that uses graphics hardware shaders to accelerate the transcription of spoken audio into written text. It implements a GPU-accelerated automatic speech recognition framework specifically designed to run Whisper models. The system focuses on high-speed processing for both recorded audio files and live microphone streams. It utilizes voice activity detection to analyze raw audio in real time, triggering the inference engine only when human speech is detected. The engine covers a broad range of capabilities including real-time audio

    This repository provides a high-performance, GPU-accelerated speech-to-text inference engine supporting live audio transcription and automatic speech recognition, though it focuses more on raw C++ performance and hardware acceleration than a comprehensive multi-language library wrapper.

    C++Automatic Speech RecognitionAudio Transcription
    GitHub पर देखें↗10,489
  • k2-fsa/sherpa-ncnnk2-fsa का अवतार

    k2-fsa/sherpa-ncnn

    1,743GitHub पर देखें↗

    Sherpa-ncnn is an edge-based speech recognition and synthesis engine designed to run neural network models locally on mobile, embedded, and desktop hardware. It provides a cross-platform framework for offline speech-to-text transcription and text-to-speech synthesis, ensuring that all audio processing occurs on-device without requiring an internet connection or external cloud services. The project distinguishes itself through its use of the ncnn inference engine, which is optimized for low-latency execution on resource-constrained devices. It incorporates on-device model quantization to reduc

    Sherpa-ncnn is a cross-platform speech recognition framework optimized for offline, edge-based execution, providing the exact real-time streaming transcription and local processing capabilities sought by this search.

    C++Mobile and Edge AIReal-Time Speech TranscriptionInference Engines
    GitHub पर देखें↗1,743
  • facebookresearch/omnilingual-asrfacebookresearch का अवतार

    facebookresearch/omnilingual-asr

    2,671GitHub पर देखें↗

    Omnilingual-ASR is a multilingual automatic speech recognition framework and toolkit designed to transcribe audio across 1,600 languages. It provides a complete pipeline for converting speech to text, including a toolkit for fine-tuning pre-trained speech models to specific languages or datasets using custom training recipes. The system supports zero-shot speech recognition, allowing the model to predict text in unseen languages without extensive training data. It further enables few-shot language guidance through in-context examples and uses language codes to constrain transcription output t

    Omnilingual-ASR is a multilingual automatic speech recognition framework that provides robust model fine-tuning and cross-lingual transcription capabilities, though it lacks direct streaming real-time transcription out of the box.

    PythonAutomatic Speech RecognitionSpeech Transcription
    GitHub पर देखें↗2,671
  • nvidia-nemo/nemoNVIDIA-NeMo का अवतार

    NVIDIA-NeMo/NeMo

    17,389GitHub पर देखें↗

    NeMo is a comprehensive framework designed for the development, training, and deployment of large-scale conversational and generative artificial intelligence models. It provides an integrated platform for building multimodal systems, encompassing speech processing, language modeling, and reinforcement learning alignment. The framework is built to handle the entire lifecycle of AI development, from data curation and model pretraining to production-ready service deployment. The platform distinguishes itself through advanced distributed training capabilities, including tensor and pipeline parall

    NeMo is a comprehensive conversational AI framework providing speech-to-text, automatic speech recognition, and speaker diarization capabilities, making it a robust toolkit for building real-time transcription pipelines.

    PythonSpeaker DiarizationSpeech TranscriptionAutomatic Speech Recognition Toolkits
    GitHub पर देखें↗17,389
  • koljab/realtimesttKoljaB का अवतार

    KoljaB/RealtimeSTT

    9,477GitHub पर देखें↗

    RealtimeSTT is a local speech-to-text engine and real-time automatic speech recognition server. It utilizes transformer-based recognition and omnilingual pipelines to convert live audio streams into text, providing a WebSocket-based streaming API for raw PCM audio transmission. The project is distinguished by a dual-backend transcription pipeline that uses a lightweight engine for immediate partial suggestions and a heavier model for final high-accuracy results. It includes a wake word detection system to trigger recording and employs a shared-resource inference model to distribute heavy spee

    RealtimeSTT is a Python speech recognition framework providing local streaming transcription and wake word detection, though it is missing some secondary capabilities like custom vocabulary support in its core interface.

    PythonAutomatic Speech RecognitionSpeaker DiarizationAudio Transcription
    GitHub पर देखें↗9,477
  • uberi/speech_recognitionUberi का अवतार

    Uberi/speech_recognition

    8,973GitHub पर देखें↗

    This project is a Python speech recognition library that serves as a unified interface for converting spoken audio into text. It functions as a bridge between Python applications and a variety of speech-to-text engines, providing a consistent way to interact with both local and cloud-based recognition services. The library distinguishes itself as a multi-engine transcription tool, wrapping diverse online APIs and offline recognition backends into a standardized format. This allows for interchangeable recognition engines and supports multilingual audio transcription through various language pa

    This Python library provides a unified wrapper interface for both online and offline speech-to-text engines, making it a solid toolkit for speech recognition even though it relies on external backends for core transcription processing.

    PythonMultilingual SupportSpeech Recognition Libraries
    GitHub पर देखें↗8,973
  • ufal/whisper_streamingufal का अवतार

    ufal/whisper_streaming

    3,642GitHub पर देखें↗

    Whisper streaming is an automated speech recognition engine designed to convert live audio into text. It functions as a network-based transcription server that accepts raw audio data from remote clients and returns incremental text results in real-time. The system distinguishes itself through its ability to process audio streams incrementally, allowing for immediate transcription and translation as speech is captured. It incorporates voice activity detection to isolate human speech from background noise and utilizes sliding-window buffering to manage incoming audio segments, ensuring that pro

    Whisper streaming is an automated speech recognition engine that provides real-time incremental transcription over a network server, making it well-suited for live audio-to-text pipelines even though it operates as a network service rather than an embeddable client-side library.

    PythonReal-Time Audio TranscribersIncremental Inference StreamingReal-Time Speech Transcription
    GitHub पर देखें↗3,642
  • abus-aikorea/voice-proabus-aikorea का अवतार

    abus-aikorea/voice-pro

    6,255GitHub पर देखें↗

    Voice Pro is a comprehensive speech and audio processing toolkit that combines text-to-speech synthesis, voice cloning, speech recognition, and translation capabilities into a single application. At its core, the project enables users to generate natural-sounding speech from text, clone voices from short audio samples without requiring prior training data, and perform real-time speech translation across over 100 languages. The platform distinguishes itself through its integrated multimedia workflow, allowing users to download YouTube videos, extract audio, separate voice tracks, generate word

    Voice Pro is an audio and speech processing application that provides speech recognition and real-time transcription, though it is delivered as an end-user application rather than a developer library or framework.

    PythonSpeech-to-Text Engines
    GitHub पर देखें↗6,255
  • espnet/espnetespnet का अवतार

    espnet/espnet

    9,861GitHub पर देखें↗

    ESPnet is a comprehensive speech processing toolkit and PyTorch-based trainer designed for building end-to-end speech recognition, synthesis, and translation models. It provides a structured framework for developing automatic speech recognition systems using transducer and encoder-decoder architectures, alongside engines for text-to-speech synthesis and speech translation pipelines. The project distinguishes itself through a recipe-based workflow execution system that ensures experimental reproducibility by running standardized sequences of scripts for data preparation and model training. It

    This repository provides a comprehensive speech processing toolkit and framework for building end-to-end automatic speech recognition models, though it is oriented toward training and experimentation rather than acting as a lightweight drop-in real-time transcription library.

    PythonAutomatic Speech RecognitionSpeech Recognition Models
    GitHub पर देखें↗9,861
  • cmusphinx/pocketsphinxcmusphinx का अवतार

    cmusphinx/pocketsphinx

    4,276GitHub पर देखें↗

    PocketSphinx is an offline speech recognition engine that converts raw audio from files or live microphone streams into written text without requiring a network connection. It functions as a speech-to-text library, a real-time transcription engine, and a voice command processor, capable of detecting and transcribing spoken commands from continuous audio streams with configurable acoustic and language models. The engine uses weighted finite-state transducers to represent acoustic, phonetic, and language models as a single search graph for efficient decoding. It employs fixed-point acoustic mod

    PocketSphinx is an offline speech recognition library that handles real-time transcription and streaming audio, though it lacks modern built-in speaker diarization and extensive multilingual neural models found in newer engines.

    CLive Stream TranscribersSpeech Recognition EnginesSpeech to Text Transcription
    GitHub पर देखें↗4,276
  • julius-speech/juliusjulius-speech का अवतार

    julius-speech/julius

    1,927GitHub पर देखें↗

    Julius is a high-performance, open-source speech recognition engine designed for large vocabulary continuous speech recognition. It functions as a comprehensive framework utilizing Hidden Markov Model-based acoustic modeling and N-gram language models to convert live or recorded audio into text. The engine is built to support real-time streaming and provides a network-accessible service that allows external applications to manage recognition sessions and receive transcription results through programmatic commands. The engine distinguishes itself through its modular architecture and support fo

    Julius is an open-source speech recognition engine written in C that supports real-time continuous speech recognition and streaming audio processing, matching your search for a recognition framework even if it relies on older Hidden Markov Model techniques.

    CReal-Time Speech ProcessingSpeech Recognition EnginesSpeech Transcription Engines
    GitHub पर देखें↗1,927
  • jamsch/expo-speech-recognitionjamsch का अवतार

    jamsch/expo-speech-recognition

    541GitHub पर देखें↗

    Expo Speech Recognition is a cross-platform mobile module that converts live microphone audio and pre-recorded files into text using native speech engines. It provides offline speech recognition capabilities by downloading and verifying local speech models to enable on-device processing without an active network connection. The library includes session lifecycle management to start, stop, or abort recording, alongside real-time spoken language detection with confidence scoring. It emits volume change events for metering interfaces, handles audio session configuration and routing, and persist

    This is a React Native speech recognition library for mobile apps that handles transcription and voice input, though it lacks the broader ecosystem features of larger server-side engines.

    TypeScriptSpeech Recognition Libraries
    GitHub पर देखें↗541
  • zzw922cn/automatic_speech_recognitionzzw922cn का अवतार

    zzw922cn/Automatic_Speech_Recognition

    2,834GitHub पर देखें↗

    This project is a machine learning toolkit designed for the development, training, and deployment of automatic speech recognition engines. It provides a comprehensive framework for converting spoken audio into written text, specifically supporting models trained on Mandarin and English datasets. The library utilizes an end-to-end neural architecture that processes raw audio input directly into character sequences, bypassing the need for intermediate linguistic alignment. It incorporates signal processing techniques to transform sound waves into numerical spectrograms and feature vectors, whic

    This project is a machine learning toolkit and framework designed for automatic speech recognition and speech-to-text conversion, though it focuses more on end-to-end model training and development rather than out-of-the-box real-time streaming inference.

    PythonSpeech Recognition EnginesAudio TranscriptionsSpeech-to-Text Models
    GitHub पर देखें↗2,834
  • buriburisuri/speech-to-text-wavenetburiburisuri का अवतार

    buriburisuri/speech-to-text-wavenet

    4,007GitHub पर देखें↗

    This project is a deep learning framework designed for end-to-end speech-to-text transcription. It utilizes the WaveNet neural network architecture to process spoken audio input and generate written text transcripts, leveraging connectionist temporal classification to map variable-length audio sequences to character-level outputs. The system distinguishes itself through a comprehensive training pipeline that supports distributed execution across multiple graphics processing units. It includes specialized utilities for audio data augmentation and the transformation of raw audio files into opti

    This project is a deep learning framework for end-to-end speech-to-text transcription utilizing WaveNet, making it the right kind of tool for automatic speech recognition, though it focuses heavily on training workflows rather than real-time streaming out of the box.

    PythonNeural Audio TranscribersSpeech Transcription EnginesConnectionist Temporal Classification
    GitHub पर देखें↗4,007
टॉप 10 की एक नज़र में तुलना करें
रिपॉजिटरीस्टार्सभाषालाइसेंसअंतिम पुश
alphacep/vosk-api14.9KJupyter NotebookApache-2.04 जून 2026
modelscope/funasr18.5KPythonMIT23 जून 2026
ggml-org/whisper.cpp50.8KC++MIT16 जून 2026
ggerganov/whisper.cpp50.8KC++MIT17 जून 2026
paddlepaddle/paddlespeech12.6KPythonApache-2.021 जून 2026
ahmetoner/whisper-asr-webservice3.3KPythonMIT23 नव॰ 2025
nvidia/nemo17.4KPythonApache-2.017 जून 2026
pannous/tensorflow-speech-recognition2.2KPythonNOASSERTION17 जन॰ 2024
m-bain/whisperx20.2KPythonbsd-2-clause19 फ़र॰ 2026
kaldi-asr/kaldi15.4KShellNOASSERTION22 सित॰ 2025

Related searches

  • स्ट्रीमिंग स्पीच रिकग्निशन के लिए लाइब्रेरी
  • ऑफलाइन स्पीच रिकग्निशन के लिए इंजन
  • रियल-टाइम वॉइस एजेंट्स के लिए फ्रेमवर्क
  • स्पीच सिंथेसिस और रिकग्निशन के लिए एक ओपन सोर्स टूल
  • an open source real time voice changer
  • Voice activity detection
  • सेल्फ-होस्टेड टेक्स्ट-टू-स्पीच इंजन
  • Voice and audio processing