9 repository-uri
Diffusion architectures that use flow-matching for more efficient noise-to-image transformation.
Distinct from Image Diffusion Models: Specifically focuses on flow-matching as an alternative to standard iterative denoising diffusion.
Explore 9 awesome GitHub repositories matching artificial intelligence & ml · Flow-Matching Frameworks. Refine with filters or upvote what's useful.
Flux is a diffusion model inference engine designed for text-to-image generation and image-to-image manipulation. It provides a system for executing open-weight models to transform natural language descriptions into visual imagery or to modify existing images. The project distinguishes itself through a flow-matching framework for image generation and a structural image controller. This controller allows for guided synthesis by using depth maps and Canny edge detection to constrain the geometry and composition of the output. The toolkit covers a broad range of image editing capabilities, incl
Utilizes a flow-matching framework to generate high-quality images more efficiently than standard diffusion.
Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al
Implements a flow matching objective to train models on continuous latents extracted from audio compressors.
F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s
Uses a flow matching engine and diffusion transformers to generate fluent synthetic speech.
This project is a machine learning research automation system designed to manage the full research lifecycle, from idea discovery to final paper submission. It utilizes markdown-based skill templates to execute autonomous research tasks and manage iterative loops of deep review and experimentation. The system distinguishes itself through integrated capabilities for academic communication and integrity auditing. It can automate the generation of LaTeX papers, conference slide decks, and evidence-grounded peer review rebuttals. To ensure rigor, it employs cross-model review routing and adversar
Transforms noise into clean embeddings using flow matching for continuous text generation.
AI NovelGenerator is a tool for generating long-form fiction using large language models. It functions as a narrative architect and writing assistant, automating the creation of multi-chapter novels while managing the overall story structure and character tracking. The project distinguishes itself through a semantic context retrieval system and an AI story consistency checker. These tools use semantic search to recall specific story details from previous chapters and scan generated text for plot contradictions or behavioral inconsistencies. The system covers a full narrative lifecycle, inclu
Provides automated scanning of generated text to identify logical plot contradictions and character inconsistencies.
Acest proiect este un framework neuronal de text-to-speech și un model PyTorch conceput pentru a sintetiza vorbirea umană. Convertește textul scris în audio sintetic prin prezicerea mel spectrogramelor, care servesc drept reprezentare intermediară pentru generarea vocii. Sistemul include un model de condiționare pentru WaveNet pentru a asigura o ieșire audio cu sunet natural. Oferă un framework de antrenare distribuită care utilizează procesarea multi-GPU și precizia mixtă automată pentru a optimiza viteza de antrenare și a reduce utilizarea memoriei. Proiectul acoperă întregul pipeline de sinteză neuronală a vorbirii, de la antrenarea modelului folosind seturi de date de text și audio până la generarea vocilor artificiale. Utilizează un encoder-decoder convoluțional și atenție secvență-la-secvență pentru a mapa caracteristicile lingvistice la cadre acustice.
Provides a comprehensive neural engine for training speech models and generating synthetic audio.
Acest proiect este un framework de modele generative bazat pe PyTorch, conceput pentru a transforma zgomotul în distribuții complexe de date prin învățarea câmpurilor vectoriale și a căilor de probabilitate. Acesta servește drept toolkit multimodal pentru generarea de text și imagini sintetice prin fluxuri de probabilitate. Biblioteca se distinge prin suportul pentru integrări continue, discrete și pe varietăți Riemanniene. Acest lucru permite framework-ului să gestioneze o varietate de tipuri de date, inclusiv date categorice prin „discrete-state flow matching” și spații non-euclidiene prin integrare pe varietăți Riemanniene. Toolkit-ul acoperă întregul pipeline generativ, incluzând definirea căilor de probabilitate, regresia câmpurilor vectoriale și utilizarea solverelor de ecuații diferențiale pentru eșantionarea datelor. Aceste capabilități permit antrenarea și inferența modelelor generative capabile să creeze conținut sintetic pe mai multe modalități.
Provides a PyTorch-based library for implementing continuous and discrete flow matching algorithms to train generative models.
ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,
ExecuTorch continues text generation from a specific point in the cache, enabling stateful continuation.
Hunyuan3D-2.1 is a generative 3D framework and image-to-3D pipeline that transforms single 2D images into textured 3D geometries. It functions as an asset generator that produces high-quality 3D meshes and textures using a flow-matching system. The project includes a specialized synthesizer for creating photorealistic textures with physically based rendering properties. These tools allow for the simulation of metallic reflections and light interactions on generated models. The system covers 3D asset pipeline automation through a sequence of shape generation and mesh refinement. It also provi
Utilizes a flow-matching pipeline to transform Gaussian noise into initial 3D asset shapes.