9 Repos
Improving visual fidelity and resolution using a two-stage structural-to-texture inference paradigm.
Distinct from Video Generation: Specific to the multi-stage structural and texture refinement process of generative models
Explore 9 awesome GitHub repositories matching artificial intelligence & ml · Multi-Stage Refinement. Refine with filters or upvote what's useful.
Sana is a framework for high-resolution image and video synthesis based on a linear diffusion transformer. It provides a toolkit for the training, fine-tuning, and execution of text-to-image and text-to-video models, as well as a video generative world model capable of simulating physical environments with precise spatial control. The project is distinguished by its use of linear complexity layers to handle high resolutions and its support for long-form, minute-length video generation in real time. It implements a two-stage inference paradigm that separates structural generation from visual t
Implements a two-stage inference paradigm to improve visual quality and resolution of generated videos.
MochiDiffusion is a local client for Stable Diffusion that functions as an AI image generation studio. It provides a workspace for performing text-to-image, image-to-image, and inpainting tasks, enabling the production of high-resolution images offline using local hardware and neural engine acceleration. The project includes a local model manager for importing, organizing, and converting machine learning models into compatible formats for offline execution. It features a ControlNet integration tool to guide structural composition and spatial layout, alongside a dedicated image upscaler that u
Improves image quality by applying a second diffusion pass using a specialized refiner model.
Dieses Projekt ist ein KI-gestütztes visuelles Canvas- und kollaboratives Whiteboard-Framework. Es fungiert als anpassbare Vektor-Zeichen-Engine und als Werkzeug zur Umwandlung handgezeichneter Interface-Skizzen und Wireframes in funktionalen Code mittels künstlicher Intelligenz. Das System zeichnet sich durch die Integration von KI-Agenten aus, die visuelle Diagramme direkt auf dem Canvas lesen, modifizieren und generieren können. Zudem bietet es einen knotenbasierten Workflow-Editor zum Aufbau von Automatisierungspipelines und Datenverarbeitungsflüssen durch die Verbindung multimodaler Komponenten. Die Plattform deckt ein breites Spektrum an Funktionen ab, einschließlich Echtzeit-Multiplayer-Kollaboration mit User-Presence-Tracking, einem unendlichen Canvas mit GPU-beschleunigtem Rendering und einer umfassenden Suite an Werkzeugen zur Objektmanipulation und -ausrichtung. Zudem implementiert sie Web-Accessibility-Standards und bietet eine skriptfähige Schnittstelle zur Definition benutzerdefinierter Formen und programmatischer Canvas-Steuerelemente.
Transforms sketches into code through a conversational loop that iteratively refines the generated layout.
vibesdk is an agentic software development platform and framework designed to coordinate autonomous agents that write, debug, and refine full-stack applications from natural language. It serves as a cloud-native application orchestrator and an LLM-powered code generation framework that converts prompts into functional code through iterative conversations and multi-phase agent behaviors. The project distinguishes itself by providing a complete toolchain for building AI development platforms. This includes the ability to integrate various model providers, construct custom LLM toolkits, and mana
Enables iterative modification of generated UI code through conversational feedback loops using text and image messages.
EchoMimic V2 ist eine KI-Video-Generierungs-Pipeline und ein Computer-Vision-Animationsmodell, das darauf ausgelegt ist, synthetische menschliche Animationen zu produzieren. Es fungiert als generatives Framework, das Halbkörper-Videos erstellt, indem ein statisches Referenzbild mit Posenbewegungen abgeglichen wird, die aus einem treibenden Video extrahiert wurden. Das System nutzt einen diffusionsbasierten Generierungsprozess in Kombination mit latenter Raumkompression und einem temporalen Aufmerksamkeitsmechanismus, um flüssige Übergänge zwischen Frames zu gewährleisten. Es wahrt die konsistente Identität einer Person durch referenzbasiertes Encoding und steuert die räumliche Platzierung mittels posengesteuerter Bewegungskonditionierung. Das Projekt enthält Funktionen zur mehrstufigen Bildverfeinerung, um Gesichtsdetails und Schärfe zu verbessern. Zudem bietet es Tools zur Vorbereitung von Animationsdatensätzen, einschließlich des Herunterladens und der Vorverarbeitung von Videodaten in Formate, die für Modelltraining und Inferenz erforderlich sind.
Employs a multi-stage refinement process to enhance facial details and overall sharpness of generated frames.
DiffBIR ist ein auf Diffusion basierendes Framework zur Bildrestaurierung für die blinde Bildrekonstruktion. Es nutzt generative Diffusions-Priors, um hochwertige Bilder aus Quellen mit unbekannten oder komplexen Degradierungen wiederherzustellen, ohne dass explizite Degradierungsmodelle erforderlich sind. Das System enthält spezialisierte Modelle für die Gesichtsrestaurierung, die die Wiederherstellung von Gesichtszügen, Texturen und Hintergründen in beschädigten Porträts ermöglichen. Um hochauflösende Ausgaben auf Hardware mit begrenztem Speicher zu unterstützen, verwendet es einen Kachel-Upscaler, der Bilder während des Samplings in kleinere Patches unterteilt. Das Framework umfasst eine mehrstufige Restaurierungspipeline und generatives Bild-Upscaling. Es bietet Funktionen für das Training von Restaurierungsmodellen und die Anwendung spezialisierter Gewichte zur Optimierung der Verbesserung für spezifische Szenen.
Employs a multi-stage pipeline of specialized models to iteratively remove image artifacts and refine details.
Presenton is an AI-powered presentation engine and API designed to transform natural language prompts, uploaded documents, and structured data into professional slide decks. It functions as a generation service that leverages large language models to automate the creation of outlines, slide content, and visual assets. The system is distinguished by its support for both cloud-based and self-hosted infrastructure, allowing for the integration of local language models and image generators to ensure data privacy. It implements a Model Context Protocol server, enabling external AI agents to trigge
Enables the refinement of slide structural and aesthetic layouts through natural language instructions.
ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat
Employs a multi-stage refinement process to recover fine visual details and increase resolution.
vibe-vibe is an LLM agent engineering framework and toolchain optimizer designed for orchestrating multi-agent systems. It serves as a comprehensive guide and methodology for transforming conceptual ideas into deployed applications through agentic software engineering. The project focuses on the orchestration of specialized AI agent roles with defined collaboration boundaries and iterative feedback loops. It provides frameworks for toolchain optimization, including the selection and evaluation of protocols that extend model capabilities and the design of standardized tool interfaces. The sys
Translates aesthetic preferences into concrete instructions for AI to refine layout and typography.