9 Repos
Diffusion architectures that use flow-matching for more efficient noise-to-image transformation.
Distinct from Image Diffusion Models: Specifically focuses on flow-matching as an alternative to standard iterative denoising diffusion.
Explore 9 awesome GitHub repositories matching artificial intelligence & ml · Flow-Matching Frameworks. Refine with filters or upvote what's useful.
Flux is a diffusion model inference engine designed for text-to-image generation and image-to-image manipulation. It provides a system for executing open-weight models to transform natural language descriptions into visual imagery or to modify existing images. The project distinguishes itself through a flow-matching framework for image generation and a structural image controller. This controller allows for guided synthesis by using depth maps and Canny edge detection to constrain the geometry and composition of the output. The toolkit covers a broad range of image editing capabilities, incl
Utilizes a flow-matching framework to generate high-quality images more efficiently than standard diffusion.
Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al
Implements a flow matching objective to train models on continuous latents extracted from audio compressors.
F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s
Uses a flow matching engine and diffusion transformers to generate fluent synthetic speech.
This project is a machine learning research automation system designed to manage the full research lifecycle, from idea discovery to final paper submission. It utilizes markdown-based skill templates to execute autonomous research tasks and manage iterative loops of deep review and experimentation. The system distinguishes itself through integrated capabilities for academic communication and integrity auditing. It can automate the generation of LaTeX papers, conference slide decks, and evidence-grounded peer review rebuttals. To ensure rigor, it employs cross-model review routing and adversar
Transforms noise into clean embeddings using flow matching for continuous text generation.
AI NovelGenerator ist ein Tool zur Generierung von Belletristik in Romanlänge unter Verwendung von Large Language Models. Es fungiert als narrativer Architekt und Schreibassistent, der die Erstellung von Romanen mit mehreren Kapiteln automatisiert, während er die gesamte Story-Struktur und das Charakter-Tracking verwaltet. Das Projekt zeichnet sich durch ein semantisches Kontext-Retrieval-System und einen KI-Story-Konsistenz-Checker aus. Diese Tools nutzen semantische Suche, um spezifische Story-Details aus vorherigen Kapiteln abzurufen und generierten Text auf Plot-Widersprüche oder Verhaltensinkonsistenzen zu scannen. Das System deckt den gesamten narrativen Lebenszyklus ab, einschließlich des Entwurfs des Story-Fundaments, Worldbuilding und der Planung der Romanstruktur. Es nutzt eine mehrstufige Pipeline zum Entwerfen kohärenter Kapitel und integriert eine kreative Workflow-Workbench zur Verwaltung von Einstellungen und Korrekturlesen.
Provides automated scanning of generated text to identify logical plot contradictions and character inconsistencies.
Dieses Projekt ist ein neuronales Text-to-Speech-Framework und ein PyTorch-Modell, das darauf ausgelegt ist, menschliche Sprache zu synthetisieren. Es konvertiert geschriebenen Text in synthetisches Audio durch die Vorhersage von Mel-Spektrogrammen, die als Zwischenrepräsentation für die Stimmgenerierung dienen. Das System enthält ein Konditionierungsmodell für WaveNet, um eine natürlich klingende Audioausgabe sicherzustellen. Es bietet ein verteiltes Trainings-Framework, das Multi-GPU-Verarbeitung und automatische Mixed-Precision nutzt, um die Trainingsgeschwindigkeit zu optimieren und den Speicherverbrauch zu reduzieren. Das Projekt deckt die gesamte Pipeline der neuronalen Sprachsynthese ab, vom Modelltraining unter Verwendung von Text- und Audiodatensätzen bis zur Generierung künstlicher Stimmen. Es verwendet einen konvolutionalen Encoder-Decoder und Sequence-to-Sequence-Attention, um sprachliche Merkmale auf akustische Frames abzubilden.
Provides a comprehensive neural engine for training speech models and generating synthetic audio.
Dieses Projekt ist ein auf PyTorch basierendes Framework für generative Modelle, das darauf ausgelegt ist, Rauschen in komplexe Datenverteilungen zu transformieren, indem es Vektorfelder und Wahrscheinlichkeitspfade erlernt. Es dient als multimodales generatives Toolkit zur Erzeugung synthetischer Texte und Bilder durch erlernte Wahrscheinlichkeitsflüsse. Die Bibliothek zeichnet sich durch die Unterstützung von kontinuierlichen, diskreten und Riemannschen Mannigfaltigkeits-Integrationen aus. Dadurch kann das Framework eine Vielzahl von Datentypen verarbeiten, einschließlich kategorialer Daten mittels Discrete-State Flow Matching und nicht-euklidischer Räume durch Riemannsche Mannigfaltigkeits-Integration. Das Toolkit deckt die gesamte generative Pipeline ab, einschließlich der Definition von Wahrscheinlichkeitspfaden, Vektorfeld-Regression und der Verwendung von Differentialgleichungslösern für das Data-Sampling. Diese Funktionen ermöglichen das Training und die Inferenz generativer Modelle, die in der Lage sind, synthetische Inhalte über mehrere Modalitäten hinweg zu generieren.
Provides a PyTorch-based library for implementing continuous and discrete flow matching algorithms to train generative models.
ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,
ExecuTorch continues text generation from a specific point in the cache, enabling stateful continuation.
Hunyuan3D-2.1 is a generative 3D framework and image-to-3D pipeline that transforms single 2D images into textured 3D geometries. It functions as an asset generator that produces high-quality 3D meshes and textures using a flow-matching system. The project includes a specialized synthesizer for creating photorealistic textures with physically based rendering properties. These tools allow for the simulation of metallic reflections and light interactions on generated models. The system covers 3D asset pipeline automation through a sequence of shape generation and mesh refinement. It also provi
Utilizes a flow-matching pipeline to transform Gaussian noise into initial 3D asset shapes.