awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

8 Repos

Awesome GitHub RepositoriesReference-Conditioned Generation

Generating images using a combination of text prompts and external visual reference files.

Distinct from Image Generation: Specifically addresses using reference files to modify styles or backgrounds within the generation process.

Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Reference-Conditioned Generation. Refine with filters or upvote what's useful.

Awesome Reference-Conditioned Generation GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • anil-matcha/open-higgsfield-aiAvatar von Anil-matcha

    Anil-matcha/Open-Higgsfield-AI

    20,529Auf GitHub ansehen↗

    Open-Higgsfield-AI is a generative AI content studio and visual workflow orchestrator. It provides a unified interface for creating photorealistic images and videos, utilizing a node-based editor to chain multiple image, video, and audio models into automated content pipelines. The system functions as an AI video animation tool and local GPU inference engine, allowing users to run generative models on local hardware or remote servers. It includes specialized capabilities for audio-driven lip synchronization and cinematic camera controls to adjust virtual lens and focal settings. The platform

    Maintains a local cache of uploaded visual assets to serve as conditional inputs for generative models.

    JavaScriptai-art-generatorai-image-generationai-video-generation
    Auf GitHub ansehen↗20,529
  • facebookresearch/parlaiAvatar von facebookresearch

    facebookresearch/ParlAI

    10,625Auf GitHub ansehen↗

    ParlAI is a conversational AI research framework designed for training, evaluating, and sharing dialogue models using a unified interface for datasets and agents. It functions as a PyTorch-based training platform and a dialogue data collection system, providing a centralized model zoo for the distribution of versioned pretrained agents. The project distinguishes itself through a knowledge-grounded retrieval system that combines dense and sparse indexing to ground responses in external information. It also provides a comprehensive infrastructure for gathering human-AI interaction data via inte

    Generates conversational responses conditioned on image inputs for multimodal dialogue.

    Python
    Auf GitHub ansehen↗10,625
  • lllyasviel/ic-lightAvatar von lllyasviel

    lllyasviel/IC-Light

    8,445Auf GitHub ansehen↗

    IC-Light is a diffusion-based image editor and generative tool designed for controlling the illumination of foreground subjects. It functions as an image relighting system that uses latent diffusion models to modify lighting effects on isolated subjects. The project provides two primary methods for lighting control: text-based relighting, which uses descriptive prompts and lighting directions, and background-based relighting, which conditions the foreground lighting to match the visual properties of a provided background image. Beyond illumination, the system includes a surface normal estima

    Provides reference-conditioned generation to match foreground lighting with the visual properties of a background image.

    Python
    Auf GitHub ansehen↗8,445
  • humanaigc/emoAvatar von HumanAIGC

    HumanAIGC/EMO

    7,616Auf GitHub ansehen↗

    EMO ist ein KI-Porträt-Animator und Audio-zu-Video-Diffusionsmodell, das entwickelt wurde, um ausdrucksstarke Talking-Head-Videos zu generieren. Es verwandelt ein einzelnes statisches Porträtbild und eine Audiospur in ein synchronisiertes Video einer sprechenden Person. Das System konzentriert sich auf die Synthese digitaler Menschen und erzeugt hochauflösende Gesichtsbewegungen und emotionale Signale. Es synchronisiert Lippenbewegungen und Gesichtsausdrücke mit gesprochenen Sprachaufnahmen, um realistische Porträt-Animationen zu erstellen. Das Framework nutzt einen Diffusionsprozess und einen Cross-Modal-Alignment-Mechanismus, um das Timing zwischen Audiosignalen und visuellen Landmarks sicherzustellen. Es verwendet referenzbasierte Bildkonditionierung, um die Identitätskonsistenz zu wahren, sowie eine zeitliche Konsistenzschicht, um flüssige Bewegungen zwischen den Frames zu gewährleisten.

    Uses a single static portrait image as a reference to maintain identity consistency across frames.

    Auf GitHub ansehen↗7,616
  • zejun-yang/aniportraitAvatar von Zejun-Yang

    Zejun-Yang/AniPortrait

    5,020Auf GitHub ansehen↗

    AniPortrait is an AI video synthesis pipeline designed to generate photorealistic speaking portraits and facial animations. It functions as a talking head generator and audio-driven animator that synchronizes lip movements, expressions, and head poses to speech or reference video sources. The system includes a facial expression transfer tool for reenacting movements from a source video onto a static reference image. It utilizes a latent diffusion model with reference-based image conditioning to maintain visual identity and consistency across generated frames. The pipeline covers audio-to-exp

    Employs reference-conditioned generation to maintain visual identity and consistency using a source portrait image.

    Python
    Auf GitHub ansehen↗5,020
  • thudm/visualglm-6bAvatar von THUDM

    THUDM/VisualGLM-6B

    4,157Auf GitHub ansehen↗

    VisualGLM-6B ist ein zweisprachiges, multimodales Large Language Model und Vision-Language-Modell für Konversationsaufgaben und visuelles Verständnis. Es fungiert als zweisprachiges KI-Modell, das Antworten sowohl auf Chinesisch als auch auf Englisch verarbeiten und generieren kann. Das System ist ein quantisiertes Large Language Model, das 4-Bit- und 8-Bit-Präzision unterstützt, um den Speicherbedarf und die Hardwareanforderungen bei der lokalen Bereitstellung zu reduzieren. Es ist zudem ein parameter-effizientes Fine-Tuning-Modell, das Gewichtsanpassungen ermöglicht, um das System ohne vollständiges Retraining an spezifische nachgelagerte Aufgaben anzupassen. Das Projekt deckt multimodale Konversations-KI und bildbasierte Dialoge ab und ermöglicht die Analyse visueller Inhalte für Aufgaben des visuellen Verständnisses in mehreren Sprachen. Zu den Funktionen gehören Modell-Präzisionsquantisierung und domänenspezifisches Fine-Tuning für spezialisierte Anwendungen.

    Generates conversational responses that answer questions and describe visual content from images.

    Python
    Auf GitHub ansehen↗4,157
  • nunchaku-ai/comfyui-nunchakuAvatar von nunchaku-ai

    nunchaku-ai/ComfyUI-nunchaku

    2,901Auf GitHub ansehen↗

    ComfyUI-nunchaku is a 4-bit diffusion inference engine and a set of nodes for running low-precision quantized diffusion models within ComfyUI visual workflows. It provides a backend that reduces memory overhead and increases generation speed for transformer models. The project includes specialized tools for identity-preserving generation and an image-to-image guidance toolkit that uses depth maps and reference images. It also features a multimodal visual question answering implementation and a utility for merging multiple quantized model files into single unified files. The engine covers a b

    Creates new images from a reference source utilizing a vision-language model.

    Pythoncomfyuidiffusionflux
    Auf GitHub ansehen↗2,901
  • tencent-hunyuan/hunyuanimage-3.0Avatar von Tencent-Hunyuan

    Tencent-Hunyuan/HunyuanImage-3.0

    2,862Auf GitHub ansehen↗

    HunyuanImage-3.0 is a diffusion-based text-to-image tool and large language model image generator designed for creating high-fidelity, photorealistic visual content. It functions as an image-to-image synthesis framework and a multimodal visual reasoning engine. The system includes a prompt refinement system that automatically rewrites sparse user inputs into detailed descriptions to improve output precision. It also employs a reasoning chain architecture to analyze image inputs and prompts, decomposing complex editing tasks into structured sub-tasks. The project covers a range of synthesis c

    Generates new images by combining text prompts with reference files to modify styles or replace backgrounds.

    Pythonimage-generationnative-multimodal-model
    Auf GitHub ansehen↗2,862
  1. Home
  2. Artificial Intelligence & ML
  3. Image Generation
  4. Reference-Conditioned Generation

Unter-Tags erkunden

  • Image-Grounded Dialogue GeneratorsGenerates conversational responses that reference visual content from images. **Distinct from Reference-Conditioned Generation:** Distinct from Reference-Conditioned Generation: generates dialogue grounded in images, not images from references.
  • Reference Asset CachingLocal storage and management of visual assets used as conditions for AI generation. **Distinct from Reference-Conditioned Generation:** Focuses on the local caching and gallery management of reference images rather than the generation process itself.