8 रिपॉजिटरी
Generating images using a combination of text prompts and external visual reference files.
Distinct from Image Generation: Specifically addresses using reference files to modify styles or backgrounds within the generation process.
Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Reference-Conditioned Generation. Refine with filters or upvote what's useful.
Open-Higgsfield-AI is a generative AI content studio and visual workflow orchestrator. It provides a unified interface for creating photorealistic images and videos, utilizing a node-based editor to chain multiple image, video, and audio models into automated content pipelines. The system functions as an AI video animation tool and local GPU inference engine, allowing users to run generative models on local hardware or remote servers. It includes specialized capabilities for audio-driven lip synchronization and cinematic camera controls to adjust virtual lens and focal settings. The platform
Maintains a local cache of uploaded visual assets to serve as conditional inputs for generative models.
ParlAI is a conversational AI research framework designed for training, evaluating, and sharing dialogue models using a unified interface for datasets and agents. It functions as a PyTorch-based training platform and a dialogue data collection system, providing a centralized model zoo for the distribution of versioned pretrained agents. The project distinguishes itself through a knowledge-grounded retrieval system that combines dense and sparse indexing to ground responses in external information. It also provides a comprehensive infrastructure for gathering human-AI interaction data via inte
Generates conversational responses conditioned on image inputs for multimodal dialogue.
IC-Light is a diffusion-based image editor and generative tool designed for controlling the illumination of foreground subjects. It functions as an image relighting system that uses latent diffusion models to modify lighting effects on isolated subjects. The project provides two primary methods for lighting control: text-based relighting, which uses descriptive prompts and lighting directions, and background-based relighting, which conditions the foreground lighting to match the visual properties of a provided background image. Beyond illumination, the system includes a surface normal estima
Provides reference-conditioned generation to match foreground lighting with the visual properties of a background image.
EMO is an AI portrait animator and audio-to-video diffusion model designed to generate expressive talking head videos. It transforms a single static portrait image and an audio track into a synchronized video of a person speaking. The system focuses on digital human synthesis, producing high-fidelity facial movements and emotional cues. It synchronizes lip movements and facial gestures to match spoken voice recordings to create realistic portrait animations. The framework utilizes a diffusion process and a cross-modal alignment mechanism to ensure timing between audio signals and visual land
Uses a single static portrait image as a reference to maintain identity consistency across frames.
AniPortrait एक AI वीडियो सिंथेसिस पाइपलाइन है जिसे फ़ोटोरियलिस्टिक स्पीकिंग पोर्ट्रेट और चेहरे के एनिमेशन जनरेट करने के लिए डिज़ाइन किया गया है। यह एक टॉकिंग हेड जनरेटर और ऑडियो-संचालित एनिमेटर के रूप में कार्य करता है जो होंठों की गतिविधियों, अभिव्यक्तियों और सिर के पोज़ को भाषण या संदर्भ वीडियो स्रोतों के साथ सिंक्रोनाइज़ करता है। सिस्टम में एक स्थिर संदर्भ छवि पर स्रोत वीडियो से गतिविधियों को फिर से लागू करने के लिए एक चेहरे की अभिव्यक्ति स्थानांतरण टूल शामिल है। यह जनरेट किए गए फ़्रेमों में दृश्य पहचान और स्थिरता बनाए रखने के लिए संदर्भ-आधारित छवि कंडीशनिंग के साथ एक लेटेंट डिफ्यूजन मॉडल का उपयोग करता है। पाइपलाइन ऑडियो-टू-एक्सप्रेशन मैपिंग, पोज़-गाइडेड मोशन कंट्रोल और फ़ोटोरियलिस्टिक वीडियो सिंथेसिस को कवर करती है। यह जनरेशन प्रक्रिया में तेज़ी लाने और कुल रेंडरिंग समय को कम करने के लिए फ़्रेम इंटरपोलेशन अपसैंपलिंग को शामिल करती है।
Employs reference-conditioned generation to maintain visual identity and consistency using a source portrait image.
VisualGLM-6B is a bilingual multimodal large language model and vision-language model designed for conversational tasks and visual understanding. It functions as a bilingual AI model capable of processing and generating responses in both Chinese and English. The system is a quantized large language model supporting 4-bit and 8-bit precision to reduce memory usage and hardware requirements during local deployment. It is also a parameter-efficient fine-tuning model, allowing for weight adjustments to adapt the system to specific downstream tasks without full retraining. The project covers mult
Generates conversational responses that answer questions and describe visual content from images.
ComfyUI-nunchaku is a 4-bit diffusion inference engine and a set of nodes for running low-precision quantized diffusion models within ComfyUI visual workflows. It provides a backend that reduces memory overhead and increases generation speed for transformer models. The project includes specialized tools for identity-preserving generation and an image-to-image guidance toolkit that uses depth maps and reference images. It also features a multimodal visual question answering implementation and a utility for merging multiple quantized model files into single unified files. The engine covers a b
Creates new images from a reference source utilizing a vision-language model.
HunyuanImage-3.0 is a diffusion-based text-to-image tool and large language model image generator designed for creating high-fidelity, photorealistic visual content. It functions as an image-to-image synthesis framework and a multimodal visual reasoning engine. The system includes a prompt refinement system that automatically rewrites sparse user inputs into detailed descriptions to improve output precision. It also employs a reasoning chain architecture to analyze image inputs and prompts, decomposing complex editing tasks into structured sub-tasks. The project covers a range of synthesis c
Generates new images by combining text prompts with reference files to modify styles or replace backgrounds.