26 repositorios
Generative models that create images by iteratively refining noise into structured visual patterns.
Explore 26 awesome GitHub repositories matching artificial intelligence & ml · Image Diffusion Models. Refine with filters or upvote what's useful.
Stable Diffusion is a generative machine learning pipeline that synthesizes high-resolution visual content by performing iterative denoising within a compressed latent space. By mapping natural language embeddings into pixel outputs through conditioned probabilistic processes, the framework enables the generation of images from text prompts and the transformation of existing visual inputs based on semantic instructions. The architecture utilizes a modular execution environment that decouples model loading, scheduler logic, and inference components to support diverse hardware configurations. I
Creates structured visual patterns by iteratively refining noise through a specialized generative machine learning pipeline.
Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr
Produces images from text prompts using large-scale diffusion models.
Flux is a diffusion model inference engine designed for text-to-image generation and image-to-image manipulation. It provides a system for executing open-weight models to transform natural language descriptions into visual imagery or to modify existing images. The project distinguishes itself through a flow-matching framework for image generation and a structural image controller. This controller allows for guided synthesis by using depth maps and Canny edge detection to constrain the geometry and composition of the output. The toolkit covers a broad range of image editing capabilities, incl
Utilizes a flow-matching framework to generate high-quality images more efficiently than standard diffusion.
Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al
Implements a flow matching objective to train models on continuous latents extracted from audio compressors.
Lama Cleaner is an AI-powered image editing application focused on inpainting, object removal, and generative filling. It provides a suite of tools for erasing unwanted elements from photos and filling the resulting gaps using generative artificial intelligence. The project includes specialized capabilities for image outpainting to extend borders, background removal through object segmentation, and face restoration to fix visual defects. It also features an image upscaler to increase resolution and clarity via super-resolution AI, as well as a Stable Diffusion-based editor for replacing speci
Implements image diffusion models to iteratively refine noise into coherent pixels for filling and extending images.
F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s
Uses a flow matching engine and diffusion transformers to generate fluent synthetic speech.
This project is a machine learning research automation system designed to manage the full research lifecycle, from idea discovery to final paper submission. It utilizes markdown-based skill templates to execute autonomous research tasks and manage iterative loops of deep review and experimentation. The system distinguishes itself through integrated capabilities for academic communication and integrity auditing. It can automate the generation of LaTeX papers, conference slide decks, and evidence-grounded peer review rebuttals. To ensure rigor, it employs cross-model review routing and adversar
Transforms noise into clean embeddings using flow matching for continuous text generation.
Z-Image is an AI image editing engine and generation framework designed for photorealistic synthesis and the refinement of diffusion models. It functions as a multilingual text-to-image renderer and a system for training custom foundation models to generate and edit visuals using natural language instructions. The project distinguishes itself through a reasoning-based prompt enhancer that expands simple descriptions into detailed visual instructions using a structured reasoning chain. It also features specialized capabilities for rendering high-quality Chinese and English typography within ge
Provides a toolkit for refining image generation models to improve specific visual capabilities through unified development bases.
This is a PyTorch implementation of a text-to-image model designed for synthesizing high-fidelity images from natural language descriptions. It utilizes a diffusion image generator to transform latent embeddings into visual data through an iterative denoising process. The system employs a two-stage latent mapping process, using a CLIP-based latent prior to map text embeddings to image embeddings before decoding them into pixels. It features a cascading diffusion decoder that produces high-resolution imagery by passing low-resolution outputs through a sequence of models at increasing scales.
Implements a generative model that creates high-fidelity images through an iterative denoising process.
RoomGPT is a generative AI image processor designed to transform photographs of existing rooms into redesigned interior spaces. It functions as an AI interior design generator and room visualizer that applies new styles and layouts to uploaded images using machine learning models. The system utilizes diffusion-based image transformation and prompt-template engineering to modify visual environments and generate home decor visualizations. These capabilities allow for the creation of diverse interior design variations based on specific style prompts. The infrastructure includes client-side imag
Employs generative image diffusion models to transform existing room photos into new interior design layouts.
Implementation of Denoising Diffusion Probabilistic Model in Pytorch
Generates images by iteratively denoising random noise through a learned reverse diffusion process.
Facechain is a generative AI toolchain and portrait generator designed to create personalized synthetic identities and consistent digital portraits. It provides a pipeline for training and refining diffusion models to produce subject-driven image synthesis from reference photos. The project focuses on digital twin generation, enabling the creation of a personalized model from a single image to maintain identity consistency across various poses and artistic styles. It utilizes identity fusion and similarity sorting to balance facial accuracy with stylized visual effects. The toolkit covers a
Uses image diffusion models to iteratively refine random noise into high-quality synthetic portraits.
IC-Light is a diffusion-based image editor and generative tool designed for controlling the illumination of foreground subjects. It functions as an image relighting system that uses latent diffusion models to modify lighting effects on isolated subjects. The project provides two primary methods for lighting control: text-based relighting, which uses descriptive prompts and lighting directions, and background-based relighting, which conditions the foreground lighting to match the visual properties of a provided background image. Beyond illumination, the system includes a surface normal estima
Employs image diffusion models to synthesize lighting and color details while maintaining the original image structure.
This is a classifier-guided diffusion framework for high-fidelity image generation. It implements a cascaded diffusion pipeline that chains a base diffusion model with a dedicated upsampler to progressively increase image resolution in stages, and uses classifier-guided diffusion sampling to steer the reverse diffusion process toward higher-quality outputs. The framework provides tools for training diffusion models from scratch using distributed processes with gradient accumulation, as well as training classifier models that provide gradient-based guidance during sampling. It supports both un
Generates high-fidelity images by sampling from a diffusion model, optionally guided by a classifier for improved quality.
AI NovelGenerator es una herramienta para generar ficción de formato largo utilizando modelos de lenguaje extensos. Funciona como un arquitecto narrativo y asistente de escritura, automatizando la creación de novelas de múltiples capítulos mientras gestiona la estructura general de la historia y el seguimiento de personajes. El proyecto se distingue por un sistema de recuperación de contexto semántico y un verificador de consistencia de historia por IA. Estas herramientas utilizan búsqueda semántica para recordar detalles específicos de la historia de capítulos anteriores y escanear el texto generado en busca de contradicciones en la trama o inconsistencias de comportamiento. El sistema cubre un ciclo de vida narrativo completo, incluyendo el diseño de la base de la historia, la construcción del mundo y la planificación de la estructura de la novela. Utiliza una canalización de múltiples etapas para redactar capítulos coherentes e incorpora un banco de trabajo de flujo de trabajo creativo para gestionar configuraciones y corrección de pruebas.
Provides automated scanning of generated text to identify logical plot contradictions and character inconsistencies.
Este proyecto es un framework de texto a voz neuronal y modelo de PyTorch diseñado para sintetizar voz humana. Convierte texto escrito en audio sintético prediciendo espectrogramas de mel, que sirven como una representación intermedia para la generación de voz. El sistema incluye un modelo de acondicionamiento para WaveNet para asegurar una salida de audio de sonido natural. Proporciona un framework de entrenamiento distribuido que utiliza procesamiento multi-GPU y precisión mixta automática para optimizar la velocidad de entrenamiento y reducir el uso de memoria. El proyecto cubre todo el pipeline de síntesis de voz neuronal, desde el entrenamiento del modelo utilizando conjuntos de datos de texto y audio hasta la generación de voces artificiales. Emplea un codificador-decodificador convolucional y atención de secuencia a secuencia para mapear características lingüísticas a marcos acústicos.
Provides a comprehensive neural engine for training speech models and generating synthetic audio.
AnyText es un framework de síntesis visual de texto y modelo de difusión latente diseñado para generar y editar texto dentro de imágenes. Funciona como un generador de texto de difusión multilingüe que combina datos de glifos y trazos en características de imagen latentes para asegurar una colocación y renderizado preciso de los caracteres. El sistema permite la modificación o reemplazo de caracteres y palabras existentes dentro de imágenes mientras preserva el contexto visual circundante. Soporta la creación de efectos de texto estilizados mediante el uso de un pipeline de fusión de pesos que combina pesos de modelos especializados y capas de adaptación para expandir las capacidades lingüísticas y estéticas. El framework cubre una gama de capacidades que incluyen generación de texto visual multilingüe, personalización de apariencia de texto para fuentes y colores, y entrenamiento de modelos de texto a imagen. También incluye herramientas de evaluación de calidad para cuantificar la precisión del texto visual y la fidelidad de la imagen utilizando métricas de distancia y precisión.
Implements a latent diffusion model that iteratively refines noise to generate high-fidelity visual text within images.
Este proyecto es un framework de modelos generativos basado en PyTorch, diseñado para transformar ruido en distribuciones de datos complejas mediante el aprendizaje de campos vectoriales y trayectorias de probabilidad. Funciona como un kit de herramientas generativo multimodal para producir texto e imágenes sintéticas a través de flujos de probabilidad aprendidos. La librería se distingue por su soporte para integraciones en variedades continuas, discretas y de Riemann. Esto permite que el framework maneje una variedad de tipos de datos, incluyendo datos categóricos mediante el emparejamiento de flujos de estado discreto y espacios no euclidianos mediante la integración en variedades de Riemann. El kit de herramientas cubre todo el pipeline generativo, incluyendo la definición de trayectorias de probabilidad, regresión de campos vectoriales y el uso de solvers de ecuaciones diferenciales para el muestreo de datos. Estas capacidades permiten el entrenamiento e inferencia de modelos generativos capaces de crear contenido sintético en múltiples modalidades.
Provides a PyTorch-based library for implementing continuous and discrete flow matching algorithms to train generative models.
Este proyecto es un recurso educativo integral y un curso para construir redes neuronales usando PyTorch. Cubre los bloques de construcción fundamentales del deep learning, incluyendo la manipulación de tensores, la diferenciación automática y la construcción de componentes modulares de redes neuronales. El repositorio sirve como guía técnica para varios dominios especializados. Proporciona detalles de implementación para tareas de visión artificial como clasificación de imágenes, detección de objetos y segmentación semántica, así como flujos de trabajo de procesamiento de lenguaje natural que involucran transformers, redes recurrentes y modelos generativos. Además, incluye una referencia para IA generativa, centrándose específicamente en la síntesis de imágenes mediante modelos de difusión y redes adversarias. El material se extiende a pipelines de optimización y despliegue de modelos. Cubre técnicas para reducir el tamaño del modelo y aumentar la velocidad de inferencia mediante cuantización y la exportación de modelos a formatos como ONNX y TensorRT. Otras áreas de capacidad incluyen ingeniería de datos para carga paralela, evaluación de modelos mediante métricas personalizadas y el despliegue de modelos de lenguaje grandes (LLM) de código abierto. El proyecto se entrega principalmente como una serie de Jupyter Notebooks.
Implements generative models that produce images by iteratively refining Gaussian noise.
ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,
ExecuTorch continues text generation from a specific point in the cache, enabling stateful continuation.