awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

9 repository-uri

Awesome GitHub RepositoriesMulti-Stage Refinement

Improving visual fidelity and resolution using a two-stage structural-to-texture inference paradigm.

Distinct from Video Generation: Specific to the multi-stage structural and texture refinement process of generative models

Explore 9 awesome GitHub repositories matching artificial intelligence & ml · Multi-Stage Refinement. Refine with filters or upvote what's useful.

Awesome Multi-Stage Refinement GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • nvlabs/sanaAvatar NVlabs

    NVlabs/Sana

    8,310Vezi pe GitHub↗

    Sana is a framework for high-resolution image and video synthesis based on a linear diffusion transformer. It provides a toolkit for the training, fine-tuning, and execution of text-to-image and text-to-video models, as well as a video generative world model capable of simulating physical environments with precise spatial control. The project is distinguished by its use of linear complexity layers to handle high resolutions and its support for long-form, minute-length video generation in real time. It implements a two-stage inference paradigm that separates structural generation from visual t

    Implements a two-stage inference paradigm to improve visual quality and resolution of generated videos.

    Python
    Vezi pe GitHub↗8,310
  • mochidiffusion/mochidiffusionAvatar MochiDiffusion

    MochiDiffusion/MochiDiffusion

    7,895Vezi pe GitHub↗

    MochiDiffusion is a local client for Stable Diffusion that functions as an AI image generation studio. It provides a workspace for performing text-to-image, image-to-image, and inpainting tasks, enabling the production of high-resolution images offline using local hardware and neural engine acceleration. The project includes a local model manager for importing, organizing, and converting machine learning models into compatible formats for offline execution. It features a ControlNet integration tool to guide structural composition and spatial layout, alongside a dedicated image upscaler that u

    Improves image quality by applying a second diffusion pass using a specialized refiner model.

    Swiftaneappleapple-silicon
    Vezi pe GitHub↗7,895
  • tldraw/draw-a-uiAvatar tldraw

    tldraw/draw-a-ui

    5,445Vezi pe GitHub↗

    Acest proiect este o pânză vizuală (canvas) bazată pe AI și un framework de whiteboard colaborativ. Acesta funcționează ca un motor de desen vectorial personalizabil și un instrument pentru convertirea schițelor de interfață desenate de mână și a wireframe-urilor în cod funcțional folosind inteligența artificială. Sistemul se distinge prin integrarea agenților AI care pot citi, modifica și genera diagrame vizuale direct pe canvas. Oferă, de asemenea, un editor de flux de lucru bazat pe noduri pentru construirea de pipeline-uri de automatizare și fluxuri de procesare a datelor prin conectarea componentelor multimodale. Platforma acoperă o gamă largă de capabilități, inclusiv colaborare multiplayer în timp real cu urmărirea prezenței utilizatorilor, un canvas infinit cu randare accelerată GPU și o suită cuprinzătoare de instrumente de manipulare și aliniere a obiectelor. De asemenea, implementează standarde de accesibilitate web și oferă o interfață scriptabilă pentru definirea formelor personalizate și a controalelor programatice pentru canvas.

    Transforms sketches into code through a conversational loop that iteratively refines the generated layout.

    TypeScript
    Vezi pe GitHub↗5,445
  • cloudflare/vibesdkAvatar cloudflare

    cloudflare/vibesdk

    5,094Vezi pe GitHub↗

    vibesdk este o platformă și un framework de dezvoltare software agentic conceput pentru a coordona agenți autonomi care scriu, depanează și rafinează aplicații full-stack din limbaj natural. Servește ca un orchestrator de aplicații cloud-native și un framework de generare de cod bazat pe LLM care convertește prompturile în cod funcțional prin conversații iterative și comportamente ale agenților în mai multe faze. Proiectul se remarcă prin furnizarea unui toolchain complet pentru construirea platformelor de dezvoltare AI. Aceasta include capacitatea de a integra diverși furnizori de modele, de a construi toolkit-uri LLM personalizate și de a gestiona întregul ciclu de viață al aplicațiilor generate de AI printr-un toolchain de deployment serverless și un SDK TypeScript programatic. Platforma acoperă o gamă largă de capabilități, inclusiv orchestrarea sandbox-urilor AI pentru execuție izolată și previzualizări live, sisteme de fișiere virtuale susținute de Git pentru urmărirea versiunilor și deployment automat în cloud pe platforme serverless. De asemenea, încorporează sisteme pentru gestionarea schemelor de baze de date, criptarea ierarhică a secretelor și sincronizarea stării în timp real prin WebSockets. Utilizatorii pot gestiona fluxurile de lucru ale proiectului printr-o interfață de linie de comandă sau programatic folosind SDK-ul furnizat.

    Enables iterative modification of generated UI code through conversational feedback loops using text and image messages.

    TypeScript
    Vezi pe GitHub↗5,094
  • antgroup/echomimic_v2Avatar antgroup

    antgroup/echomimic_v2

    4,597Vezi pe GitHub↗

    EchoMimic V2 este un pipeline de generare video AI și un model de animație computer vision conceput pentru a produce animații umane sintetice. Funcționează ca un framework generativ care creează videoclipuri semi-corp prin alinierea unei imagini de referință statice cu mișcările de postură extrase dintr-un videoclip de conducere (driving video). Sistemul utilizează un proces de generare bazat pe difuzie combinat cu compresia spațiului latent și un mecanism de atenție temporală pentru a asigura tranziții line între cadre. Menține identitatea consistentă a persoanei prin codificare bazată pe referință și ghidează plasarea spațială prin condiționarea mișcării bazată pe postură. Proiectul include capabilități pentru rafinarea imaginilor în mai multe etape pentru a îmbunătăți detaliile faciale și claritatea. Oferă, de asemenea, instrumente pentru pregătirea seturilor de date de animație, inclusiv descărcarea și preprocesarea datelor video în formatele necesare pentru antrenarea modelului și inferență.

    Employs a multi-stage refinement process to enhance facial details and overall sharpness of generated frames.

    Pythonaudio-driven-body-animationaudio-driven-portrait-animationsaudio-driven-talking-face
    Vezi pe GitHub↗4,597
  • xpixelgroup/diffbirAvatar XPixelGroup

    XPixelGroup/DiffBIR

    4,087Vezi pe GitHub↗

    DiffBIR is a diffusion-based image restoration framework designed for blind image reconstruction. It utilizes generative diffusion priors to recover high-quality images from sources with unknown or complex degradations without requiring explicit degradation models. The system includes specialized models for face restoration, enabling the recovery of facial landmarks, textures, and backgrounds in degraded portraits. To support high-resolution outputs on hardware with limited memory, it employs a tiled image upscaler that divides images into smaller patches during sampling. The framework cover

    Employs a multi-stage pipeline of specialized models to iteratively remove image artifacts and refine details.

    Python
    Vezi pe GitHub↗4,087
  • presenton/presentonAvatar presenton

    presenton/presenton

    4,042Vezi pe GitHub↗

    Presenton is an AI-powered presentation engine and API designed to transform natural language prompts, uploaded documents, and structured data into professional slide decks. It functions as a generation service that leverages large language models to automate the creation of outlines, slide content, and visual assets. The system is distinguished by its support for both cloud-based and self-hosted infrastructure, allowing for the integration of local language models and image generators to ensure data privacy. It implements a Model Context Protocol server, enabling external AI agents to trigge

    Enables the refinement of slide structural and aesthetic layouts through natural language instructions.

    TypeScriptai-agentai-presentationapi
    Vezi pe GitHub↗4,042
  • lightricks/comfyui-ltxvideoAvatar Lightricks

    Lightricks/ComfyUI-LTXVideo

    3,840Vezi pe GitHub↗

    ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat

    Employs a multi-stage refinement process to recover fine visual details and increase resolution.

    Pythoncomfyuidiffusion-modelsdit
    Vezi pe GitHub↗3,840
  • datawhalechina/vibe-vibeAvatar datawhalechina

    datawhalechina/vibe-vibe

    3,126Vezi pe GitHub↗

    vibe-vibe is an LLM agent engineering framework and toolchain optimizer designed for orchestrating multi-agent systems. It serves as a comprehensive guide and methodology for transforming conceptual ideas into deployed applications through agentic software engineering. The project focuses on the orchestration of specialized AI agent roles with defined collaboration boundaries and iterative feedback loops. It provides frameworks for toolchain optimization, including the selection and evaluation of protocols that extend model capabilities and the design of standardized tool interfaces. The sys

    Translates aesthetic preferences into concrete instructions for AI to refine layout and typography.

    agentagentic-aiai
    Vezi pe GitHub↗3,126
  1. Home
  2. Artificial Intelligence & ML
  3. Video Generation
  4. Multi-Stage Refinement

Explorează sub-etichetele

  • AI Layout Refinement2 sub-tag-uriProcesses for manually refining and adjusting AI-generated structural layouts in a visual editor. **Distinct from Multi-Stage Refinement:** Focuses on the human-in-the-loop refinement of AI-generated infographics rather than multi-stage model inference.
  • Image Restoration PipelinesSequential processing stages designed to iteratively remove artifacts and refine visual details in images. **Distinct from Multi-Stage Refinement:** Focuses on image restoration sequences rather than video structural-to-texture refinement