awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
spotify avatar

spotify/basic-pitch

0
View on GitHub↗
5,207 نجوم·467 تفرعات·Python·Apache-2.0·12 مشاهداتbasicpitch.io↗

Basic Pitch

Basic-pitch هو ناسخ صوتي للشبكة العصبية وكاشف طبقة صوت متعدد الأصوات. يعمل كمحول صوت إلى MIDI يحول تسجيلات الصوت متعددة الأصوات إلى أحداث ملاحظات MIDI وبيانات انحناء طبقة الصوت.

يحافظ النظام على التعبير الموسيقي من خلال تتبع تقلبات التردد المستمرة لتحويل الانزلاقات والاهتزاز إلى أحداث انحناء طبقة صوت MIDI. يستخدم محرك استدلال قابلاً للتوصيل يسمح بتهيئة وقت تشغيل النموذج بناءً على نظام التشغيل أو احتياجات تسريع الأجهزة.

يوفر المشروع واجهة سطر أوامر لمعالجة الصوت المجمعة وواجهة برمجية لدمج النسخ واستخراج أحداث الملاحظات في برمجيات مخصصة. يمكن تصدير نتائج النسخ كملفات MIDI، ومخرجات نموذج خام، وجداول بيانات أحداث الملاحظات.

Features

  • Audio Transcription Models - Uses a deep learning model to detect multiple simultaneous musical pitches from raw audio waveforms.
  • Automated Music Transcription - Uses neural networks to programmatically turn audio files into MIDI data and note events.
  • Neural Audio Transcribers - Uses a machine learning system to detect musical pitches in audio files to automate the transcription process.
  • AI Audio Analysis - Processes audio data using AI to extract multiple simultaneous musical notes for transcription.
  • Audio-to-MIDI Converters - Transforms polyphonic audio recordings into MIDI note events and pitch bend data using neural networks.
  • Audio-to-MIDI Pipelines - Implements a pipeline that sequentially processes raw audio through a neural network to produce MIDI files.
  • Audio-to-MIDI Transcription - Converts polyphonic audio recordings into MIDI files to capture notes and pitch bends for music production.
  • Pitch Bend Detection - Tracks continuous frequency fluctuations to convert glides and vibrato into MIDI pitch bend events.
  • Polyphonic Audio-to-MIDI Transcription - Transforms polyphonic audio files into MIDI format with pitch bend detection to preserve musical expression.
  • Polyphonic Pitch Detectors - Identifies multiple simultaneous notes and frequency changes within an audio signal.
  • Polyphonic Pitch Tracking - Analyzes complex audio signals to separate and identify individual overlapping notes across a wide frequency range.
  • CLI Transcription Tools - Provides a command-line tool for transcribing audio files directly into MIDI data from the terminal.
  • Audio Command-Line Tools - Provides terminal-based utilities for processing audio files into MIDI transcriptions.
  • Audio Transcription APIs - Provides a programmatic interface for integrating audio-to-MIDI transcription into custom software applications.
  • CLI Data Processors - Provides a command-line tool that transforms audio files into structured MIDI notation data.
  • Audio Processing Automation - Exposes the model runtime through a terminal interface for batch processing and automated file conversion.

سجل النجوم

مخطط تاريخ النجوم لـ spotify/basic-pitchمخطط تاريخ النجوم لـ spotify/basic-pitch

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Basic Pitch

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Basic Pitch.
  • music-and-culture-technology-lab/omnizartالصورة الرمزية لـ Music-and-Culture-Technology-Lab

    Music-and-Culture-Technology-Lab/omnizart

    1,915عرض على GitHub↗

    Omnizart is a deep learning framework designed for automatic music transcription and music information retrieval. It functions as a toolkit for analyzing polyphonic audio recordings to extract structured musical information, including notes, chord progressions, drum events, and rhythmic patterns. The system provides a modular pipeline that orchestrates the entire lifecycle of audio analysis, from initial feature extraction and data preparation to model inference. Users can apply pre-trained models to transcribe audio directly or utilize the included utilities to train and fine-tune neural net

    Pythonbeat-trackingchorddrum-transcription
    عرض على GitHub↗1,915
  • thewh1teagle/vibeالصورة الرمزية لـ thewh1teagle

    thewh1teagle/vibe

    5,298عرض على GitHub↗

    Vibe is a cross-platform transcription tool that converts spoken audio into text by running Whisper neural models directly on your device, with no cloud dependency. It can transcribe audio from files, microphones, system output, and network streams, and supports both batch processing of multiple files and real-time captioning from continuous input. Beyond basic transcription, Vibe identifies and labels different speakers through speaker diarization, and offers a choice of Command-Line Interface or HTTP API for automated and remote workflows. It also includes plugins to export transcripts to c

    TypeScriptaicross-platformdesktop
    عرض على GitHub↗5,298
  • buriburisuri/speech-to-text-wavenetالصورة الرمزية لـ buriburisuri

    buriburisuri/speech-to-text-wavenet

    4,007عرض على GitHub↗

    This project is a deep learning framework designed for end-to-end speech-to-text transcription. It utilizes the WaveNet neural network architecture to process spoken audio input and generate written text transcripts, leveraging connectionist temporal classification to map variable-length audio sequences to character-level outputs. The system distinguishes itself through a comprehensive training pipeline that supports distributed execution across multiple graphics processing units. It includes specialized utilities for audio data augmentation and the transformation of raw audio files into opti

    Python
    عرض على GitHub↗4,007
  • crmne/ruby_llmالصورة الرمزية لـ crmne

    crmne/ruby_llm

    3,566عرض على GitHub↗

    ruby_llm is an LLM integration framework and AI agent orchestrator designed to connect applications to multiple large language model providers through a unified interface. It serves as a toolkit for building autonomous assistants with custom personas, managing structured output via JSON schemas, and implementing vector embedding engines for semantic search. The project distinguishes itself as an observability suite and multimodal toolkit. It provides specialized capabilities for tracking token usage, calculating model costs, and tracing workflows via OpenTelemetry, while supporting the proces

    Rubyaianthropicchatgpt
    عرض على GitHub↗3,566
عرض جميع البدائل الـ 27 لـ Basic Pitch→

الأسئلة الشائعة

ما هي وظيفة spotify/basic-pitch؟

Basic-pitch هو ناسخ صوتي للشبكة العصبية وكاشف طبقة صوت متعدد الأصوات. يعمل كمحول صوت إلى MIDI يحول تسجيلات الصوت متعددة الأصوات إلى أحداث ملاحظات MIDI وبيانات انحناء طبقة الصوت.

ما هي الميزات الرئيسية لـ spotify/basic-pitch؟

الميزات الرئيسية لـ spotify/basic-pitch هي: Audio Transcription Models, Automated Music Transcription, Neural Audio Transcribers, AI Audio Analysis, Audio-to-MIDI Converters, Audio-to-MIDI Pipelines, Audio-to-MIDI Transcription, Pitch Bend Detection.

ما هي البدائل مفتوحة المصدر لـ spotify/basic-pitch؟

تشمل البدائل مفتوحة المصدر لـ spotify/basic-pitch: music-and-culture-technology-lab/omnizart — Omnizart is a deep learning framework designed for automatic music transcription and music information retrieval. It… thewh1teagle/vibe — Vibe is a cross-platform transcription tool that converts spoken audio into text by running Whisper neural models… buriburisuri/speech-to-text-wavenet — This project is a deep learning framework designed for end-to-end speech-to-text transcription. It utilizes the… jaakkopasanen/autoeq — AutoEq is a command-line tool that generates headphone equalization settings by comparing frequency response… crmne/ruby_llm — ruby_llm is an LLM integration framework and AI agent orchestrator designed to connect applications to multiple large… argmaxinc/whisperkit.