awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
mozilla avatar

mozilla/DeepSpeechArchived

0
View on GitHub↗
26,748 نجوم·4,086 تفرعات·C++·MPL-2.0·10 مشاهدات

DeepSpeech

DeepSpeech هو إطار عمل مفتوح المصدر لتحويل الكلام إلى نص ومحرك تعلم آلي مصمم لتحويل الصوت المنطوق إلى نص مكتوب محلياً على الجهاز. يوفر التعرف على الكلام على الجهاز والذي يعمل دون الحاجة إلى اتصال بالإنترنت بخوادم خارجية.

يدعم النظام نسخ الكلام في الوقت الفعلي عبر مجموعة متنوعة من منصات الأجهزة، بدءاً من أجهزة الكمبيوتر ذات اللوحة الواحدة وأجهزة الحافة وصولاً إلى خوادم GPU. وهذا يسمح بتحليل الصوت ومعالجته مباشرة على الأجهزة المحلية.

Features

  • Local Speech-to-Text - Provides a machine learning engine for the on-device conversion of spoken audio into text without internet access.
  • Real-Time Transcription - Provides instantaneous conversion of live audio streams into text transcripts across various hardware platforms.
  • On-Device Inference Engines - Ships a runtime optimized for executing speech recognition models locally on edge hardware to ensure privacy and low latency.
  • Speech Recognition - Implements tools and models for converting spoken language into text locally on a device.
  • Speech-to-Text Engines - Implements a high-performance engine for converting spoken audio into written text using local machine learning models.
  • Speech-to-Text Modeling Toolkits - Provides a toolkit for training and deploying models that convert audio signals into written text.
  • Speech-to-Text Frameworks - Offers an open-source framework for building and deploying embedded voice recognition models on diverse hardware.
  • Embedded Voice Processing - Enables the integration of speech-to-text capabilities directly into edge devices like Raspberry Pi.
  • Natural Language Processing - TensorFlow implementation of DeepSpeech architecture.
  • Speech and Audio Models - Open-source speech-to-text engine for mobile and edge.
  • Speech Processing - Pretrained automatic speech recognition engine.
  • Acoustic User Interface - Open-source speech-to-text engine using machine learning.
  • واجهات المستخدم الصوتية - محرك مفتوح المصدر لتحويل الكلام إلى نص يعتمد على التعلم العميق.
  • Audio Processing - Embedded speech-to-text engine using deep learning.

سجل النجوم

مخطط تاريخ النجوم لـ mozilla/deepspeechمخطط تاريخ النجوم لـ mozilla/deepspeech

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ DeepSpeech

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع DeepSpeech.
  • alphacep/vosk-apiالصورة الرمزية لـ alphacep

    alphacep/vosk-api

    14,853عرض على GitHub↗

    Vosk is an offline speech-to-text engine and API that converts spoken audio into text locally on a device. It provides a cross-platform speech toolkit with language bindings for integrating voice recognition into server environments, Android, iOS, and Raspberry Pi. The project includes a speaker identification tool to distinguish between different voices and an acoustic model trainer for building custom neural network models. These training tools enable speech feature extraction and model accuracy evaluation to improve recognition for specialized domains. The system supports real-time audio

    Jupyter Notebookandroidasrdeep-learning
    عرض على GitHub↗14,853
  • k2-fsa/sherpa-onnxالصورة الرمزية لـ k2-fsa

    k2-fsa/sherpa-onnx

    13,017عرض على GitHub↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    C++aarch64androidarm32
    عرض على GitHub↗13,017
  • pipecat-ai/pipecatالصورة الرمزية لـ pipecat-ai

    pipecat-ai/pipecat

    12,846عرض على GitHub↗

    Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag

    Pythonaichatbot-frameworkchatbots
    عرض على GitHub↗12,846
  • facebookresearch/wav2letterالصورة الرمزية لـ facebookresearch

    facebookresearch/wav2letter

    6,444عرض على GitHub↗

    wav2letter is an automatic speech recognition toolkit and deep learning framework designed to convert audio speech signals into written text. It functions as a distributed training system and an inference engine for building and deploying neural network architectures. The system enables the training of large-scale speech models across multiple compute nodes using custom architecture files and structured recipes. It includes an inference engine that allows these trained models to be executed within Python workflows to transform audio sequences into text. The framework covers the full speech r

    C++
    عرض على GitHub↗6,444
عرض جميع البدائل الـ 30 لـ DeepSpeech→

الأسئلة الشائعة

ما هي وظيفة mozilla/deepspeech؟

DeepSpeech هو إطار عمل مفتوح المصدر لتحويل الكلام إلى نص ومحرك تعلم آلي مصمم لتحويل الصوت المنطوق إلى نص مكتوب محلياً على الجهاز. يوفر التعرف على الكلام على الجهاز والذي يعمل دون الحاجة إلى اتصال بالإنترنت بخوادم خارجية.

ما هي الميزات الرئيسية لـ mozilla/deepspeech؟

الميزات الرئيسية لـ mozilla/deepspeech هي: Local Speech-to-Text, Real-Time Transcription, On-Device Inference Engines, Speech Recognition, Speech-to-Text Engines, Speech-to-Text Modeling Toolkits, Speech-to-Text Frameworks, Embedded Voice Processing.

ما هي البدائل مفتوحة المصدر لـ mozilla/deepspeech؟

تشمل البدائل مفتوحة المصدر لـ mozilla/deepspeech: alphacep/vosk-api — Vosk is an offline speech-to-text engine and API that converts spoken audio into text locally on a device. It provides… k2-fsa/sherpa-onnx — Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device… pipecat-ai/pipecat — Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech… facebookresearch/wav2letter — wav2letter is an automatic speech recognition toolkit and deep learning framework designed to convert audio speech… sevask/ecoute — Ecoute is a live transcription tool that provides real-time transcripts for both the user's microphone input (You) and… koljab/realtimestt — RealtimeSTT is a local speech-to-text engine and real-time automatic speech recognition server. It utilizes…