awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

8 مستودعات

Awesome GitHub RepositoriesReference-Conditioned Generation

Generating images using a combination of text prompts and external visual reference files.

Distinct from Image Generation: Specifically addresses using reference files to modify styles or backgrounds within the generation process.

Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Reference-Conditioned Generation. Refine with filters or upvote what's useful.

Awesome Reference-Conditioned Generation GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • anil-matcha/open-higgsfield-aiالصورة الرمزية لـ Anil-matcha

    Anil-matcha/Open-Higgsfield-AI

    20,529عرض على GitHub↗

    Open-Higgsfield-AI is a generative AI content studio and visual workflow orchestrator. It provides a unified interface for creating photorealistic images and videos, utilizing a node-based editor to chain multiple image, video, and audio models into automated content pipelines. The system functions as an AI video animation tool and local GPU inference engine, allowing users to run generative models on local hardware or remote servers. It includes specialized capabilities for audio-driven lip synchronization and cinematic camera controls to adjust virtual lens and focal settings. The platform

    Maintains a local cache of uploaded visual assets to serve as conditional inputs for generative models.

    JavaScriptai-art-generatorai-image-generationai-video-generation
    عرض على GitHub↗20,529
  • facebookresearch/parlaiالصورة الرمزية لـ facebookresearch

    facebookresearch/ParlAI

    10,625عرض على GitHub↗

    ParlAI is a conversational AI research framework designed for training, evaluating, and sharing dialogue models using a unified interface for datasets and agents. It functions as a PyTorch-based training platform and a dialogue data collection system, providing a centralized model zoo for the distribution of versioned pretrained agents. The project distinguishes itself through a knowledge-grounded retrieval system that combines dense and sparse indexing to ground responses in external information. It also provides a comprehensive infrastructure for gathering human-AI interaction data via inte

    Generates conversational responses conditioned on image inputs for multimodal dialogue.

    Python
    عرض على GitHub↗10,625
  • lllyasviel/ic-lightالصورة الرمزية لـ lllyasviel

    lllyasviel/IC-Light

    8,445عرض على GitHub↗

    IC-Light is a diffusion-based image editor and generative tool designed for controlling the illumination of foreground subjects. It functions as an image relighting system that uses latent diffusion models to modify lighting effects on isolated subjects. The project provides two primary methods for lighting control: text-based relighting, which uses descriptive prompts and lighting directions, and background-based relighting, which conditions the foreground lighting to match the visual properties of a provided background image. Beyond illumination, the system includes a surface normal estima

    Provides reference-conditioned generation to match foreground lighting with the visual properties of a background image.

    Python
    عرض على GitHub↗8,445
  • humanaigc/emoالصورة الرمزية لـ HumanAIGC

    HumanAIGC/EMO

    7,616عرض على GitHub↗

    EMO هو نموذج ذكاء اصطناعي لتحريك الصور الشخصية ونموذج انتشار من الصوت إلى الفيديو، مصمم لتوليد مقاطع فيديو تعبيرية لأشخاص يتحدثون. يقوم بتحويل صورة شخصية ثابتة ومقطع صوتي إلى فيديو متزامن لشخص يتحدث. يركز النظام على تركيب الشخصيات الرقمية، وينتج حركات وجه وإشارات عاطفية عالية الدقة. يقوم بمزامنة حركات الشفاه وإيماءات الوجه لتطابق التسجيلات الصوتية لإنشاء رسوم متحركة واقعية للصور الشخصية. يستخدم إطار العمل عملية انتشار وآلية محاذاة متعددة الوسائط لضمان التوقيت بين الإشارات الصوتية ومعالم الوجه. كما يستخدم تكييف الصور القائم على المرجع للحفاظ على اتساق الهوية وطبقة اتساق زمنية لضمان سلاسة الحركة بين الإطارات.

    Uses a single static portrait image as a reference to maintain identity consistency across frames.

    عرض على GitHub↗7,616
  • zejun-yang/aniportraitالصورة الرمزية لـ Zejun-Yang

    Zejun-Yang/AniPortrait

    5,020عرض على GitHub↗

    AniPortrait هو خط أنابيب لتوليف الفيديو بالذكاء الاصطناعي مصمم لإنشاء صور شخصية ناطقة واقعية ورسوم متحركة للوجه. يعمل كمولد للرؤوس المتحدثة ورسوم متحركة مدفوعة بالصوت تقوم بمزامنة حركات الشفاه، والتعبيرات، ووضعيات الرأس مع الكلام أو مصادر الفيديو المرجعية. يتضمن النظام أداة لنقل تعبيرات الوجه لإعادة تمثيل الحركات من فيديو مصدر على صورة مرجعية ثابتة. يستخدم نموذج انتشار كامن مع تكييف الصورة القائم على المرجع للحفاظ على الهوية البصرية والاتساق عبر الإطارات المولدة. يغطي خط الأنابيب تعيين الصوت إلى التعبير، والتحكم في الحركة الموجه بالوضعية، وتوليف الفيديو الواقعي. يدمج النظام زيادة دقة الإطارات (Upsampling) لتسريع عملية التوليد وتقليل وقت العرض الإجمالي.

    Employs reference-conditioned generation to maintain visual identity and consistency using a source portrait image.

    Python
    عرض على GitHub↗5,020
  • thudm/visualglm-6bالصورة الرمزية لـ THUDM

    THUDM/VisualGLM-6B

    4,157عرض على GitHub↗

    VisualGLM-6B هو نموذج لغوي كبير متعدد الوسائط ثنائي اللغة ونموذج رؤية-لغة مصمم لمهام المحادثة والفهم البصري. يعمل كنموذج ذكاء اصطناعي ثنائي اللغة قادر على معالجة وتوليد الاستجابات باللغتين الصينية والإنجليزية. النظام عبارة عن نموذج لغوي كبير مكمم (quantized) يدعم دقة 4-بت و8-بت لتقليل استخدام الذاكرة ومتطلبات الأجهزة أثناء النشر المحلي. وهو أيضاً نموذج فعال في ضبط المعلمات (parameter-efficient fine-tuning)، مما يسمح بتعديلات الأوزان لتكييف النظام مع مهام محددة دون الحاجة لإعادة التدريب الكامل. يغطي المشروع الذكاء الاصطناعي للمحادثة متعدد الوسائط والحوار القائم على الصور، مما يتيح تحليل المحتوى البصري لأداء مهام الفهم البصري عبر لغات متعددة. تشمل قدراته تكميم دقة النموذج والضبط الدقيق الخاص بالمجال للتطبيقات المتخصصة.

    Generates conversational responses that answer questions and describe visual content from images.

    Python
    عرض على GitHub↗4,157
  • nunchaku-ai/comfyui-nunchakuالصورة الرمزية لـ nunchaku-ai

    nunchaku-ai/ComfyUI-nunchaku

    2,901عرض على GitHub↗

    ComfyUI-nunchaku is a 4-bit diffusion inference engine and a set of nodes for running low-precision quantized diffusion models within ComfyUI visual workflows. It provides a backend that reduces memory overhead and increases generation speed for transformer models. The project includes specialized tools for identity-preserving generation and an image-to-image guidance toolkit that uses depth maps and reference images. It also features a multimodal visual question answering implementation and a utility for merging multiple quantized model files into single unified files. The engine covers a b

    Creates new images from a reference source utilizing a vision-language model.

    Pythoncomfyuidiffusionflux
    عرض على GitHub↗2,901
  • tencent-hunyuan/hunyuanimage-3.0الصورة الرمزية لـ Tencent-Hunyuan

    Tencent-Hunyuan/HunyuanImage-3.0

    2,862عرض على GitHub↗

    HunyuanImage-3.0 is a diffusion-based text-to-image tool and large language model image generator designed for creating high-fidelity, photorealistic visual content. It functions as an image-to-image synthesis framework and a multimodal visual reasoning engine. The system includes a prompt refinement system that automatically rewrites sparse user inputs into detailed descriptions to improve output precision. It also employs a reasoning chain architecture to analyze image inputs and prompts, decomposing complex editing tasks into structured sub-tasks. The project covers a range of synthesis c

    Generates new images by combining text prompts with reference files to modify styles or replace backgrounds.

    Pythonimage-generationnative-multimodal-model
    عرض على GitHub↗2,862
  1. Home
  2. Artificial Intelligence & ML
  3. Image Generation
  4. Reference-Conditioned Generation

استكشف الوسوم الفرعية

  • Image-Grounded Dialogue GeneratorsGenerates conversational responses that reference visual content from images. **Distinct from Reference-Conditioned Generation:** Distinct from Reference-Conditioned Generation: generates dialogue grounded in images, not images from references.
  • Reference Asset CachingLocal storage and management of visual assets used as conditions for AI generation. **Distinct from Reference-Conditioned Generation:** Focuses on the local caching and gallery management of reference images rather than the generation process itself.