For ein Stem-Splitter, the strongest matches are anjok07/ultimatevocalremovergui (Ultimate Vocal Remover is a desktop application that uses), facebookresearch/demucs (Demucs is a deep learning-powered stem splitter that separates) and deezer/spleeter (Spleeter is an AI-powered audio source separation library that). ace-step/ace-step-1.5 and tsurumeso/vocal-remover round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Open-Source-Tools, die Machine-Learning-Modelle nutzen, um Gesang und einzelne Instrumentenspuren aus Audioaufnahmen zu isolieren.
Ultimate Vocal Remover is a desktop application designed for AI-driven audio source separation. It utilizes deep learning models to isolate vocals, drums, and other individual instruments from mixed audio files, providing a utility for professional production and creative editing workflows. The software distinguishes itself by leveraging GPU-accelerated tensor computation to perform complex signal processing tasks, significantly reducing the time required for high-fidelity audio extraction. It incorporates a modular plugin architecture that integrates external utilities to support a wide rang
Ultimate Vocal Remover is a desktop application that uses deep learning to isolate vocals, drums, and other instruments from mixed audio, directly matching the need for AI-driven stem separation and vocal removal, with batch processing and GPU acceleration — exactly the kind of tool described.
Demucs is a deep learning stem splitter and AI music de-mixing software used to isolate vocals and instruments from a single audio file. It functions as a PyTorch audio source separation tool that splits mixed tracks into individual stems such as drums, bass, and vocals. The system is a hybrid spectrogram waveform separator that combines spectral and waveform analysis. This approach allows the software to process audio in both frequency and time domains to achieve high-fidelity source separation. The tool provides capabilities for audio source separation, including acapella track extraction
Demucs is a deep learning-powered stem splitter that separates vocals and instruments from mixed audio, covering vocal removal and multi-stem separation via command line, though it lacks a built-in graphical user interface.
Spleeter is an AI audio source separation library and deep learning toolkit designed to split mixed music files into individual audio stems, such as vocals and drums. It provides a suite of pretrained models for isolating different instruments and voices from a recording. The toolkit includes capabilities for training and evaluating custom audio separation models using labeled datasets and configuration files. It also features utilities for measuring model performance by comparing separation outputs against reference datasets. The system manages audio processing through spectral representati
Spleeter is an AI-powered audio source separation library that isolates vocals, drums, bass, and other stems from music tracks, matching your need for vocal removal and multi-stem separation, though it lacks a built-in graphical interface.
ACE Step 1.5 is a local text-to-music generation and audio editing system that runs on consumer hardware. It transforms plain-language descriptions into full-length songs with lyrics, and can edit existing audio through cover generation, vocal removal, track separation, and selective repainting. The system supports multilingual prompts and lyrics in over 50 languages, and provides precise control over musical structure including duration, BPM, key, and time signature. The project distinguishes itself through a dual-stream diffusion architecture that processes separate latent streams for vocal
ACE Step 1.5 is an AI-powered local audio editing system that can perform vocal removal and multi-stem separation alongside music generation; it includes a Gradio web GUI and batch processing workflows, making it a solid fit for your source separation needs even though it is not a dedicated stem-splitting tool.
Vocal Remover is a deep learning application designed for audio source separation. It functions as a command-line utility that decomposes complex audio signals into individual components, specifically isolating vocals and instrumental tracks from mixed recordings. The software utilizes a symmetric encoder-decoder neural network architecture to process audio spectrograms. By applying learned magnitude masks to the original signal phase, the system reconstructs output audio while maintaining temporal coherence. It supports both the execution of pre-trained models for track extraction and the tr
This tool uses deep neural networks to separate vocals from music, fitting the core need of vocal removal, but it does not support multi-stem separation (e.g., drums, bass) as requested.
Vocal-separate is an audio processing tool designed to isolate vocal and instrumental tracks from audio and video files. It functions as a local artificial intelligence engine that performs source separation directly on the user's machine, ensuring data privacy by eliminating the need for external server connectivity. The system provides a browser-based control interface for managing media uploads and monitoring processing tasks. To handle intensive signal decomposition, it utilizes hardware-accelerated tensor processing, which offloads complex mathematical calculations to dedicated graphics
This local web-based tool uses deep-learning models to separate vocals and background music into 2, 4, or 5 stems, directly matching the request for an open-source audio source separation tool with vocal removal and multi-stem separation, though batch processing isn’t explicitly mentioned.
This project is a deep learning toolkit designed for audio source separation and music information retrieval. It provides a framework for decomposing polyphonic audio signals into distinct components, such as vocals, drums, and bass, by processing raw waveforms through neural network architectures. The library enables users to train custom separation models or fine-tune existing ones to improve accuracy on specific audio datasets. It supports the entire model lifecycle, including the conversion of raw audio into structured, indexed formats to optimize data loading and training efficiency. Th
This repository provides a PyTorch-based tool to separate audio recordings into individual sources like vocals and instruments, making it a direct open-source candidate for your use case, though it likely lacks a graphical interface and explicit batch processing features.