5 repository-uri
Training processes that associate image data with corresponding text descriptions to improve prompt adherence.
Distinct from Text Model Training: Focuses on image-text pair association for generative models rather than general text-only model training.
Explore 5 awesome GitHub repositories matching artificial intelligence & ml · Caption-Based Training. Refine with filters or upvote what's useful.
This project is a toolkit for fine-tuning and managing text-to-image diffusion models. It focuses on low-rank adaptation to create small, portable weight files that customize model styles and behaviors without modifying the entire base model. The project provides specialized utilities for model distillation using singular value decomposition to extract adapters from fully trained models, as well as tools for blending and merging multiple adapters through weight interpolation. It includes capabilities for subject inversion and pivotal tuning to increase the visual fidelity of specific identiti
Implements training capabilities that link images with text descriptions to improve generation accuracy.
BLIP is a vision-language model framework that combines contrastive, matching, and language modeling objectives to align images with text. Built on a multimodal encoder-decoder architecture, it supports distributed data-parallel training with cosine learning rate scheduling and sliding-window metric tracking for training stability. The framework provides capabilities for image captioning, visual question answering, and cross-modal retrieval, scoring semantic alignment between images and text through learned embeddings. It includes toolkits for fine-tuning pre-trained models on custom datasets
Trains vision-language models to generate descriptive captions for images using paired image-caption datasets.
Neuraltalk is an automated image captioning system that generates natural language descriptions for images. It utilizes a deep learning model that integrates a pretrained convolutional neural network for visual feature extraction with a recurrent neural network decoder to produce text sequences. The project provides a full workflow for training and evaluating captioning models, including weight optimization via backpropagation and gradient descent. It includes tools for measuring caption accuracy by comparing generated text against reference descriptions. The system covers data preprocessing
Optimizes model parameters to predict sentence descriptions by associating image features with ground-truth text.
Kolors este o implementare de model generativ pentru sintetizarea imaginilor fotorealiste din descrieri în limbaj natural și referințe vizuale. Utilizează un framework de tip latent diffusion model pentru a produce imagini de înaltă fidelitate, operând într-un spațiu latent comprimat pentru a îmbunătăți eficiența și calitatea generării. Sistemul funcționează ca un generator de imagini multilingv, interpretând prompt-uri text în mai multe limbi pentru a produce rezultate vizuale corecte din punct de vedere semantic. Include un pipeline personalizat de antrenare a modelelor care utilizează adaptarea low-rank (LoRA) pentru a învăța modelul subiecte specifice sau stiluri artistice dintr-un set mic de imagini. Proiectul acoperă o gamă largă de capabilități de sinteză și editare a imaginilor, inclusiv transformări text-to-image și image-to-image. Oferă instrumente pentru controlul layout-ului spațial prin hărți de adâncime sau de postură, injectarea identității vizuale pentru consistență estetică și inpainting bazat pe măști pentru a reconstrui sau modifica regiuni specifice ale imaginii. Implementarea include utilitare pentru evaluarea calității imaginilor, punctând imaginile generate pe baza metricilor de preferință umană pentru calitatea estetică și semantică.
Provides the ability to interpret text prompts in multiple languages to produce semantically accurate visual outputs.
This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal language models. It provides a comprehensive suite for vision-language pretraining, enabling the development of models that map images and text into shared latent spaces. The framework is distinguished by its capabilities in high-fidelity image generation and multimodal research, utilizing normalizing flows and variational autoencoders to produce images from text prompts or class labels. It supports the development of both generative and contrastive models, allowing for a wide
Maps images and text into a shared space using captioning-based pretraining and self-supervised losses.