awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
ace-step avatar

ace-step/ACE-Step-1.5

0
View on GitHub↗
6,002 نجوم·675 تفرعات·Python·mit·13 مشاهداتacemusic.ai↗

ACE Step 1.5

ACE Step 1.5 is a local text-to-music generation and audio editing system that runs on consumer hardware. It transforms plain-language descriptions into full-length songs with lyrics, and can edit existing audio through cover generation, vocal removal, track separation, and selective repainting. The system supports multilingual prompts and lyrics in over 50 languages, and provides precise control over musical structure including duration, BPM, key, and time signature.

The project distinguishes itself through a dual-stream diffusion architecture that processes separate latent streams for vocals and instruments, synchronized through cross-attention layers during denoising. It enables style personalization through lightweight LoRA adapters that can be trained from a few songs in about one hour, and supports batch generation of up to eight songs simultaneously. The system can generate complete songs in under ten seconds on a standard consumer GPU while using less than four gigabytes of video memory.

The software is accessible through multiple interfaces including a Gradio web UI, a REST API, a CLI wizard, and a VST3 plugin for direct integration into digital audio workstations. It also includes a pre-trained source separation pipeline for isolating vocal and instrumental stems from mixed audio.

Features

  • Text-to-Music Engines - An engine that transforms text prompts into full-length songs with lyrics, supporting multilingual input and precise control over musical structure.
  • Local Generative Music Systems - Generating full songs from text prompts on consumer hardware without cloud dependency, supporting multilingual lyrics and style control.
  • Text-to-Music Generators - Suno transforms a plain-language description into a full-length song, handling composition, lyrics, and style from a single user query.
  • Compositional Parameter Controllers - Suno specifies duration, BPM, key, time signature, and lyrics in 50+ languages to guide the generated composition.
  • Audio Source Separation Models - Removing vocals from songs or isolating instrumental tracks from uploaded audio files for remixing and editing.
  • Source Separation Tools - Separates mixed audio into vocal and instrumental stems using a pre-trained source separation model before applying targeted editing or conversion operations.
  • Dual-Stream Diffusion Architectures - Generates music by processing separate latent streams for vocals and instruments that are synchronised through cross-attention layers during the denoising process.
  • Latent Diffusion Models - Encodes text prompts into a compressed latent space where a diffusion model iteratively denoises random noise into structured audio guided by cross-attention to the text embedding.
  • Audio Editing - Suno performs cover generation, selective repainting, vocal-to-BGM conversion, and track separation on uploaded audio files.
  • AI-Assisted Audio Editors - A platform that performs cover generation, vocal removal, track separation, and selective repainting on uploaded audio files.
  • Multilingual Music Generation - Suno follows a text description accurately regardless of the language used, supporting multilingual music generation and editing.
  • Programmatic Audio Editing - Editing and remixing existing audio through cover generation, selective repainting, and track separation via API or CLI.
  • Local Web Interfaces - Suno starts a Gradio web UI, REST API, CLI wizard, or VST3 plugin for interactive or programmatic music generation.
  • Music Generation Batch Processors - An engine that produces up to eight songs simultaneously from text prompts to accelerate creative workflows.
  • Batch Generation Pipelines - Suno produces up to eight songs simultaneously to accelerate creative workflows and experimentation.
  • Multilingual Prompt Systems - A system that follows prompts in over 50 languages to generate and edit music, including lyrics and structural parameters.
  • LoRA Style Adapters - Training a lightweight LoRA adapter from a few songs to capture and reproduce a user's unique musical style.
  • Diffusion Model Adaptations - Applies low-rank adaptation matrices to the diffusion model's cross-attention layers, enabling personalised style transfer with minimal parameter updates and fast training.
  • LoRA Training - Suno trains a lightweight adapter from a few songs to capture a user's unique style, completing training in about one hour.
  • Parallel Diffusion Generation - Generates multiple songs simultaneously by running independent diffusion processes in parallel on the GPU, maximising throughput for batch workflows.
  • Low-VRAM Music Generation - Suno generates complete songs in under ten seconds on a standard consumer GPU while using less than four gigabytes of video memory.
  • Audio Track Repainting Tools - Suno replaces the vocal, instrumental, or background elements of an existing track while preserving the original structure and timing.
  • Batch Generation Workflows - Producing multiple song variations simultaneously from text prompts to accelerate creative experimentation and workflow.
  • Vocal-to-Instrumental Converters - Suno removes the vocal track from a song and replaces it with a purely instrumental arrangement derived from the original.
  • VST3 Plugin Packaging - Packages the generation engine as a VST3 audio plugin, enabling direct integration into digital audio workstations for real-time music production workflows.
  • Gradio Interfaces - Serves the model through a Gradio-based interactive UI and a RESTful API endpoint, allowing both manual and programmatic access to generation and editing features.

سجل النجوم

مخطط تاريخ النجوم لـ ace-step/ace-step-1.5مخطط تاريخ النجوم لـ ace-step/ace-step-1.5

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة ace-step/ace-step-1.5؟

ACE Step 1.5 is a local text-to-music generation and audio editing system that runs on consumer hardware. It transforms plain-language descriptions into full-length songs with lyrics, and can edit existing audio through cover generation, vocal removal, track separation, and selective repainting. The system supports multilingual prompts and lyrics in over 50 languages, and provides precise control over musical structure including duration, BPM, key, and time signature.

ما هي الميزات الرئيسية لـ ace-step/ace-step-1.5؟

الميزات الرئيسية لـ ace-step/ace-step-1.5 هي: Text-to-Music Engines, Local Generative Music Systems, Text-to-Music Generators, Compositional Parameter Controllers, Audio Source Separation Models, Source Separation Tools, Dual-Stream Diffusion Architectures, Latent Diffusion Models.

ما هي البدائل مفتوحة المصدر لـ ace-step/ace-step-1.5؟

تشمل البدائل مفتوحة المصدر لـ ace-step/ace-step-1.5: fspecii/ace-step-ui — ace-step-ui is an AI music production workspace and interface for generating, editing, and organizing synthetic audio… ace-step/ace-step — ACE-Step is a high-fidelity audio synthesis system and diffusion model designed to generate music and vocals from text… facebookresearch/demucs — Demucs is a deep learning stem splitter and AI music de-mixing software used to isolate vocals and instruments from a… anjok07/ultimatevocalremovergui — Ultimate Vocal Remover is a desktop application designed for AI-driven audio source separation. It utilizes deep… multimodal-art-projection/yue — YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open. aigc-audio/audiogpt — AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural…

بدائل مفتوحة المصدر لـ ACE Step 1.5

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع ACE Step 1.5.
  • fspecii/ace-step-uiالصورة الرمزية لـ fspecii

    fspecii/ace-step-ui

    4,138عرض على GitHub↗

    ace-step-ui is an AI music production workspace and interface for generating, editing, and organizing synthetic audio tracks and vocals. It provides a technical control panel for managing prompts, seeds, and style parameters to produce high-quality audio. The project includes a digital audio workstation interface for trimming and fading files, alongside an audio stem separation tool that splits mixed tracks into individual components such as drums, bass, and vocals. It also features a music video creator for generating visual content and procedural album art to accompany generated music. The

    JavaScriptace-stepaiai-music
    عرض على GitHub↗4,138
  • ace-step/ace-stepالصورة الرمزية لـ ace-step

    ace-step/ACE-Step

    4,088عرض على GitHub↗

    ACE-Step is a high-fidelity audio synthesis system and diffusion model designed to generate music and vocals from text descriptions. It functions as a music generator and vocal synthesizer, using a diffusion transformer decoder to produce audio across various languages and genres. The project provides tools for text-guided audio editing, including the ability to extend the duration of tracks, regenerate specific song segments, and perform latent-space audio inpainting to modify lyrics or styles. It also includes a framework for audio style fine-tuning using low-rank adaptation to adapt vocal

    Python
    عرض على GitHub↗4,088
  • facebookresearch/demucsالصورة الرمزية لـ facebookresearch

    facebookresearch/demucs

    10,236عرض على GitHub↗

    Demucs is a deep learning stem splitter and AI music de-mixing software used to isolate vocals and instruments from a single audio file. It functions as a PyTorch audio source separation tool that splits mixed tracks into individual stems such as drums, bass, and vocals. The system is a hybrid spectrogram waveform separator that combines spectral and waveform analysis. This approach allows the software to process audio in both frequency and time domains to achieve high-fidelity source separation. The tool provides capabilities for audio source separation, including acapella track extraction

    Python
    عرض على GitHub↗10,236
  • anjok07/ultimatevocalremoverguiالصورة الرمزية لـ Anjok07

    Anjok07/ultimatevocalremovergui

    23,673عرض على GitHub↗

    Ultimate Vocal Remover is a desktop application designed for AI-driven audio source separation. It utilizes deep learning models to isolate vocals, drums, and other individual instruments from mixed audio files, providing a utility for professional production and creative editing workflows. The software distinguishes itself by leveraging GPU-accelerated tensor computation to perform complex signal processing tasks, significantly reducing the time required for high-fidelity audio extraction. It incorporates a modular plugin architecture that integrates external utilities to support a wide rang

    Pythonaudioinstrumentalkaraoke
    عرض على GitHub↗23,673
عرض جميع البدائل الـ 30 لـ ACE Step 1.5→