awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

27 repositorios

Awesome GitHub RepositoriesVision-Language Training

Specialized training workflows for models that process both visual and textual data in query-response formats.

Distinct from Vision Model Training: Specifically addresses the intersection of vision and language (VLM) rather than general vision-only models

Explore 27 awesome GitHub repositories matching artificial intelligence & ml · Vision-Language Training. Refine with filters or upvote what's useful.

Awesome Vision-Language Training GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • microsoft/bringing-old-photos-back-to-lifeAvatar de microsoft

    microsoft/Bringing-Old-Photos-Back-to-Life

    15,691Ver en GitHub↗

    This project is a deep learning image restoration tool designed to remove scratches, fading, and noise from aged photographs and film. It utilizes generative adversarial networks for image translation, alongside specialized networks for face enhancement and video colorization. The system distinguishes itself through a combination of latent-space domain mapping and progressive face enhancement to recover blurred or missing high-frequency facial details. For video content, it employs a colorization framework that uses optical flow and temporal guidance to propagate color from selected keyframes

    Identifies and labels scratched areas in old photos to generate paired data for training restoration models.

    Pythongansgenerative-adversarial-networkimage-manipulation
    Ver en GitHub↗15,691
  • mlfoundations/open_clipAvatar de mlfoundations

    mlfoundations/open_clip

    13,935Ver en GitHub↗

    Open CLIP is an open source framework for training and deploying Contrastive Language-Image Pre-training models. It serves as a vision-language training framework and multimodal embedding engine that maps images and text into a shared vector space for similarity searches and zero-shot classification. The project provides a toolkit for distributed training of contrastive models and includes an image-to-text generative model for producing natural language descriptions. It supports custom text encoder integration and utilizes teacher-student model distillation to transfer knowledge from large pr

    Provides a comprehensive framework for training contrastive models that align visual and textual data.

    Pythoncomputer-visioncontrastive-lossdeep-learning
    Ver en GitHub↗13,935
  • liheyoung/depth-anythingAvatar de LiheYoung

    LiheYoung/Depth-Anything

    8,124Ver en GitHub↗

    Depth-Anything is a monocular depth estimation foundation model that produces dense per-pixel depth maps from a single RGB image. It is built on a DINOv2 Vision Transformer encoder backbone and trained on 62 million unlabeled images using a teacher-student pseudo-labeling framework, enabling robust generalization across diverse scenes without task-specific training. The model outputs both relative depth maps, which capture the ordering of scene points, and metric depth maps with real-world units after fine-tuning on datasets like NYUv2 or KITTI. The project distinguishes itself through its ab

    Ships a fine-tuning framework for adapting the pretrained depth model to custom datasets and downstream tasks.

    Pythondepth-estimationimage-synthesismetric-depth-estimation
    Ver en GitHub↗8,124
  • paddlepaddle/larkAvatar de PaddlePaddle

    PaddlePaddle/LARK

    7,717Ver en GitHub↗

    LARK is a development toolkit for training, fine-tuning, and deploying large language models and multimodal models based on PaddlePaddle. It functions as a comprehensive framework that includes an LLM training orchestrator, an inference server, and a multimodal model framework for processing text, image, and video inputs. The project features a retrieval-augmented generation system for building conversational applications that integrate web search and private knowledge bases. It provides specific capabilities for multimodal reasoning and complex logic, enabling the extraction of structured da

    Optimizes multimodal training using specialized data processing for images and video in query-response formats.

    Python
    Ver en GitHub↗7,717
  • thudm/cogvlmAvatar de THUDM

    THUDM/CogVLM

    6,742Ver en GitHub↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Adapts pretrained vision-language models to custom domains using LoRA and other fine-tuning methods.

    Python
    Ver en GitHub↗6,742
  • foundationvision/bytetrackAvatar de FoundationVision

    FoundationVision/ByteTrack

    6,492Ver en GitHub↗

    ByteTrack is a multi-object tracking framework that implements the ByteTrack algorithm, an ECCV 2022 method designed to recover occluded objects and reduce trajectory fragmentation. The core innovation of the project is its association algorithm, which processes every detection box—including low-confidence ones—by using separate high and low score thresholds, Kalman filter motion prediction, and Hungarian algorithm matching to produce consistent object identities across video frames. The project distinguishes itself by its comprehensive approach to handling occlusions and fragmented trajector

    Provides a pipeline to fine-tune pretrained detectors on custom multi-object tracking datasets.

    Pythondeploymentmulti-object-trackingpytorch
    Ver en GitHub↗6,492
  • qwenlm/qwen-vlAvatar de QwenLM

    QwenLM/Qwen-VL

    6,535Ver en GitHub↗

    Adapts pretrained vision-language models to custom tasks using full-parameter, LoRA, or Q-LoRA methods.

    Pythonlarge-language-modelsvision-language-model
    Ver en GitHub↗6,535
  • ailab-cvc/yolo-worldAvatar de AILab-CVC

    AILab-CVC/YOLO-World

    6,425Ver en GitHub↗

    YOLO-World is a vision-language framework and open-vocabulary object detection model. It identifies objects in images and video based on free-form text prompts without requiring predefined category labels. The system enables the identification of arbitrary objects by fusing image features with text embeddings. It includes a specialized tool for automated image labeling, which generates bounding box annotations for custom datasets using text-based prompts. The project provides a deployment pipeline for converting models into quantized ONNX and TFLite formats, supporting real-time inference on

    Adapts pre-trained vision-language models to custom domains using specialized fine-tuning methods.

    Python
    Ver en GitHub↗6,425
  • jingyaogong/minimind-vAvatar de jingyaogong

    jingyaogong/minimind-v

    6,431Ver en GitHub↗

    Provides an open-source framework for building and fine-tuning small vision-language models.

    Pythonartificial-intelligencechatgptvision-language-model
    Ver en GitHub↗6,431
  • om-ai-lab/vlm-r1Avatar de om-ai-lab

    om-ai-lab/VLM-R1

    5,991Ver en GitHub↗

    VLM-R1 es un modelo de razonamiento visión-lenguaje y framework de IA corpórea (embodied AI) diseñado para mapear entradas visuales e instrucciones de lenguaje en puntos de navegación físicos y acciones robóticas. Funciona como un optimizador de políticas multimodal y un detector de vocabulario abierto capaz de localizar objetos basados en descripciones arbitrarias en lenguaje natural. El sistema se distingue por el uso de razonamiento de cadena de pensamiento (chain-of-thought) y aprendizaje por refuerzo para resolver tareas visuales y espaciales complejas. Utiliza un sistema de memoria semántica de video, que emplea una caché visual para mantener un historial de video en vivo para interacciones de baja latencia y razonamiento temporal continuo. El framework cubre una amplia gama de capacidades, incluyendo el mapeo de puntos de navegación monoculares para robótica, la localización de tokens de región para identificación de objetos y el ajuste fino supervisado basado en políticas para la estabilidad del razonamiento multimodal. También admite detección de vocabulario abierto, comprensión de expresiones de referencia y la extracción de características de objetos detalladas mediante recuperación de prompts visuales. El proyecto está implementado en Python y admite inferencia en hardware Ascend.

    Trains and fine-tunes vision-language models to improve multimodal reasoning stability.

    Python
    Ver en GitHub↗5,991
  • ofa-sys/chinese-clipAvatar de OFA-Sys

    OFA-Sys/Chinese-CLIP

    5,942Ver en GitHub↗

    Chinese-CLIP es un framework multimodal y modelo de visión-lenguaje diseñado para la recuperación intermodal y la generación de representaciones utilizando texto e imágenes en chino. Emplea una arquitectura de aprendizaje contrastivo para mapear datos visuales y textuales en un espacio vectorial compartido para cálculos de similitud. El sistema permite la búsqueda bidireccional, facilitando la recuperación de texto a imagen e imagen a texto. También proporciona clasificación de imágenes zero-shot, que identifica objetos dentro de imágenes sin requerir entrenamiento específico para la tarea. El proyecto incluye herramientas para el ajuste fino (fine-tuning) de modelos preentrenados en conjuntos de datos especializados mediante entrenamiento distribuido y aprendizaje contrastivo. También proporciona utilidades para exportar pesos de modelos a formatos optimizados para aumentar la velocidad de inferencia en entornos de producción.

    Provides tools to adapt pre-trained vision-language models to specific datasets using contrastive learning.

    Jupyter Notebook
    Ver en GitHub↗5,942
  • facebookresearch/mmfAvatar de facebookresearch

    facebookresearch/mmf

    5,635Ver en GitHub↗

    MMF is a modular framework for building, training, and evaluating vision-and-language models. It provides a configuration-driven experiment system where model, dataset, and training parameters are defined through composable YAML files, alongside a curated model zoo of pretrained checkpoints for state-of-the-art multimodal architectures. The framework includes a multimodal dataset loader that downloads, processes, and batches vision-and-language data, and a vision-language model trainer supporting distributed training, mixed precision, and checkpoint-based resumption. The framework distinguish

    Trains multimodal models on specified datasets using provided configurations and saves trained weights.

    Pythoncaptioningdeep-learningdialog
    Ver en GitHub↗5,635
  • salesforce/blipAvatar de salesforce

    salesforce/BLIP

    5,676Ver en GitHub↗

    BLIP is a vision-language model framework that combines contrastive, matching, and language modeling objectives to align images with text. Built on a multimodal encoder-decoder architecture, it supports distributed data-parallel training with cosine learning rate scheduling and sliding-window metric tracking for training stability. The framework provides capabilities for image captioning, visual question answering, and cross-modal retrieval, scoring semantic alignment between images and text through learned embeddings. It includes toolkits for fine-tuning pre-trained models on custom datasets

    Provides an open-source framework for training, fine-tuning, and evaluating vision-language models on custom image-text datasets.

    Jupyter Notebookimage-captioningimage-text-retrievalvision-and-language-pre-training
    Ver en GitHub↗5,676
  • karpathy/neuraltalk2Avatar de karpathy

    karpathy/neuraltalk2

    5,588Ver en GitHub↗

    Neuraltalk2 es un sistema de visión de aprendizaje profundo diseñado para el subtitulado automático de imágenes. Construido con PyTorch, utiliza una arquitectura híbrida que combina un codificador de red neuronal convolucional con un decodificador de red neuronal recurrente para generar descripciones textuales a partir de entradas visuales. El proyecto cuenta con una tubería de entrenamiento acelerada por GPU capaz de distribuir cargas de trabajo a través de múltiples unidades de procesamiento gráfico mediante distribución multiproceso. Admite la generación de descripciones tanto para archivos de imagen estáticos como para flujos de video en tiempo real. El framework incluye capacidades para el ajuste fino del codificador, muestreo de texto mediante búsqueda de haz (beam search) con control de temperatura y el uso de métricas de lenguaje estándar de la industria para evaluar la precisión y fluidez de los subtítulos. También proporciona utilidades para el preprocesamiento de conjuntos de datos, persistencia de puntos de control del modelo y exportación de predicciones a archivos JSON estructurados. La implementación se proporciona como un Jupyter Notebook.

    Provides industry-standard metrics to evaluate the accuracy and fluency of generated image captions.

    Jupyter Notebook
    Ver en GitHub↗5,588
  • rllm-org/rllmAvatar de rllm-org

    rllm-org/rllm

    5,641Ver en GitHub↗

    rllm is an asynchronous reinforcement learning framework for training language agents. It provides a unified pipeline that runs the same agent code for both evaluation and training, automatically capturing traces for gradient computation. The framework supports distributed reinforcement learning across multiple GPUs and nodes using pluggable backends, and executes agents in isolated sandboxes—either locally or in the cloud—for safe and scalable rollout collection. It trains agents built with LangGraph, SmolAgents, OpenAI Agents SDK, or custom frameworks without requiring core logic changes. T

    Supports multimodal models like Qwen2-VL and Qwen3-VL by processing image inputs alongside text during training.

    Pythonagent-frameworkagentic-workflowcoding-agent
    Ver en GitHub↗5,641
  • openvla/openvlaAvatar de openvla

    openvla/openvla

    5,305Ver en GitHub↗

    OpenVLA is a vision-language-action model and framework designed for general-purpose robotic manipulation. It provides a robotic policy training framework and a control inference engine that map visual and textual inputs to robotic control actions, enabling zero-shot instruction following on hardware. The project includes a robotics dataset pipeline for standardizing diverse trajectory data and managing dataset mixtures. It supports large-scale model training through distributed GPU compute and sharded data parallelism, alongside parameter-efficient adaptation for fine-tuning models to new ta

    Provides a framework for adapting pretrained vision-language-action models to new tasks using parameter-efficient fine-tuning.

    Python
    Ver en GitHub↗5,305
  • hiyouga/easyr1Avatar de hiyouga

    hiyouga/EasyR1

    5,034Ver en GitHub↗

    EasyR1 es un sistema de entrenamiento distribuido y framework de aprendizaje por refuerzo para modelos de lenguaje y visión-lenguaje de gran escala. Funciona como un entrenador multimodal y una implementación de un pipeline de Proximal Policy Optimization diseñado para refinar las capacidades de razonamiento y percepción de modelos que procesan tanto texto como imágenes. El sistema se especializa en distribuir cargas de trabajo de aprendizaje por refuerzo a través de múltiples nodos de cómputo para gestionar altos requisitos de memoria. Optimiza el uso del hardware mediante entrenamiento sin padding y fine-tuning para ajustar modelos grandes en las unidades de procesamiento gráfico (GPU) disponibles. El framework cubre el aprendizaje por refuerzo y la orquestación de modelos de recompensa, incluyendo flujos de trabajo de aprendizaje por refuerzo a partir de retroalimentación humana (RLHF). Su superficie técnica incluye paralelismo de datos distribuido, entrenamiento de precisión híbrida y pipelines de entrada multimodal para datos intercalados de texto e imagen. El proyecto incluye utilidades para la recuperación de estado basada en checkpoints y se integra con herramientas de registro externas para rastrear el progreso del entrenamiento y las métricas de rendimiento.

    Runs reinforcement learning pipelines to improve reasoning and perception in models processing both text and images.

    Python
    Ver en GitHub↗5,034
  • huggingface/nanovlmAvatar de huggingface

    huggingface/nanoVLM

    4,917Ver en GitHub↗

    nanoVLM es un framework de entrenamiento y kit de herramientas para modelos pequeños de visión-lenguaje. Proporciona un entorno basado en PyTorch para entrenar y ajustar modelos con el fin de asociar entradas de imagen con descripciones textuales y generar respuestas en lenguaje natural. El proyecto incluye una herramienta de versionado de modelos en la nube para guardar y cargar pesos de modelos en repositorios centralizados, sincronizando activos entre entornos. También cuenta con una suite de evaluación dedicada para medir la precisión y fiabilidad de los modelos de visión-lenguaje frente a datasets de tareas estándar. El framework cubre la planificación de recursos de GPU mediante la medición del consumo de VRAM y gestiona la estabilidad del entrenamiento con persistencia de estado basada en checkpoints y gestión de memoria por lotes.

    Provides a comprehensive framework for training and fine-tuning small vision-language models.

    Python
    Ver en GitHub↗4,917
  • mlfoundations/open_flamingoAvatar de mlfoundations

    mlfoundations/open_flamingo

    4,107Ver en GitHub↗

    Open Flamingo es un framework de entrenamiento de modelos de lenguaje multimodal de gran tamaño diseñado para integrar codificadores de visión preentrenados con modelos de lenguaje. Implementa una arquitectura de visión-lenguaje que utiliza capas de atención cruzada para procesar secuencias intercaladas de imágenes y texto. El sistema se caracteriza por sus capacidades de aprendizaje multimodal few-shot, permitiendo al modelo adaptarse a nuevas tareas visuales utilizando un pequeño conjunto de ejemplos de imagen-texto proporcionados en el prompt. Admite aprendizaje en contexto y generación de texto multimodal para tareas como respuesta a preguntas visuales y subtitulado. El framework incluye un entrenador de modelos distribuido que emplea paralelismo de datos y checkpointing de gradiente para la optimización de memoria a través de múltiples GPUs. También proporciona utilidades para la carga de datasets multimodales fragmentados, evaluación de modelos paralelizada e infraestructura para alojar modelos a gran escala para inferencia.

    Assembles a unified architecture by integrating and tuning weights from specialized pretrained vision and language models.

    Pythoncomputer-visiondeep-learningflamingo
    Ver en GitHub↗4,107
  • starsfieldai/r1-vAvatar de StarsfieldAI

    StarsfieldAI/R1-V

    4,060Ver en GitHub↗

    R1-V es un conjunto de herramientas para el desarrollo de modelos multimodales, que proporciona un entorno de entrenamiento de bajo coste diseñado para optimizar los bucles de razonamiento y retroalimentación de modelos grandes de visión-lenguaje. Integra un framework de entrenamiento, pipelines de ajuste fino (fine-tuning) y herramientas de evaluación de rendimiento. El proyecto cuenta con un framework de aprendizaje por refuerzo que mejora el razonamiento visual y la generalización al recompensar las salidas correctas basadas en verificación visual. También incluye un pipeline de ajuste fino supervisado para personalizar modelos de visión-lenguaje en tareas específicas mediante datasets etiquetados y archivos de configuración. La suite abarca herramientas de evaluación de razonamiento visual y datasets específicos para evaluar el rendimiento del modelo en tareas de conteo y geometría.

    Provides tools for adapting pretrained vision-language models to custom tasks using supervised training scripts.

    Python
    Ver en GitHub↗4,060
Ant.12Siguiente
  1. Home
  2. Artificial Intelligence & ML
  3. Model Training Frameworks
  4. Vision Model Training
  5. Vision-Language Training

Explorar subetiquetas

  • Captioning Metric EvaluatorsRuns inference on validation sets using trained vision-language models and reports standard captioning metrics. **Distinct from Vision-Language Training:** Distinct from Vision-Language Training: focuses on post-training evaluation with captioning metrics, not the training workflow itself.
  • Cost-Efficient TrainingTraining methodologies optimized for affordable compute resources to enable rapid experimentation. **Distinct from Vision-Language Training:** Focuses on resource efficiency and cost reduction rather than the general training workflow of VLMs.
  • From-Scratch TrainingsBuilding a multimodal model that processes images and text together by adding a visual encoder and projection layer to a small language model backbone. **Distinct from Vision-Language Training:** Distinct from Vision-Language Training: specifically trains from scratch rather than fine-tuning a pretrained model.
  • NLVR2 Training WorkflowsTraining workflows for vision-language models on the NLVR2 dataset using paired images and text to perform visual reasoning tasks. **Distinct from Vision-Language Training:** Distinct from Vision-Language Training: specifically targets the NLVR2 dataset for visual reasoning, not general vision-language training.
  • Supervised Scratch DetectionsDefect detection methods that utilize labeled paired data to identify scratched areas in images. **Distinct from From-Scratch Trainings:** Specifically covers supervised labeling for scratch detection, unlike the generic training candidates provided.
  • Training FrameworksOpen-source frameworks for training and fine-tuning small vision-language models from scratch or from pretrained components. **Distinct from Vision-Language Training:** Distinct from Vision-Language Training: specifically provides a framework for training and fine-tuning, not just the training workflow itself.
  • Vision-Language Fine-Tunings2 sub-etiquetasAdapting pretrained vision-language models to custom tasks using full-parameter, LoRA, or Q-LoRA training methods. **Distinct from Vision-Language Training:** Distinct from Vision-Language Training: focuses on fine-tuning pretrained models rather than training from scratch.
  • Vision-Language Pretraining1 sub-etiquetaFrameworks for the initial training of models that map visual and textual data into a shared latent space. **Distinct from Vision-Language Fine-Tunings:** Focuses on the pretraining phase (contrastive/generative) rather than the fine-tuning phase of pretrained models.