awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to x-plug/mplug-owl

Projects sharing features with MPLUG Owl

30 open-source projects similar to x-plug/mplug-owl, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • open-compass/vlmevalkitopen-compass avatar

    open-compass/VLMEvalKit

    3,824View on GitHub↗

    VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex

    Pythonchatgptclaudeclip
    View on GitHub↗3,824
  • alenai97/micevalalenai97 avatar

    alenai97/MiCEval

    6View on GitHub↗

    An automatic evaluation framework for Multimodal Chain-of-Thought.

    Python
    View on GitHub↗6
  • baaivision/emu3.5baaivision avatar

    baaivision/Emu3.5

    1,526View on GitHub↗

    Native Multimodal Models are World Learners

    Python
    View on GitHub↗1,526
  • bradyfu/video-mmeBradyFU avatar

    BradyFU/Video-MME

    779View on GitHub↗

    ✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    View on GitHub↗779
  • bytedance/lynx-llmbytedance avatar

    bytedance/lynx-llm

    272View on GitHub↗

    paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/

    Pythonresearch
    View on GitHub↗272

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • cmmmu-benchmark/cmmmuCMMMU-Benchmark avatar

    CMMMU-Benchmark/CMMMU

    48View on GitHub↗

    🌐 Homepage | 🤗 Paper | 📖 arXiv | 🤗 Dataset | 🏆 EvalAI | GitHub

    Python
    View on GitHub↗48
  • damo-nlp-sg/m3examDAMO-NLP-SG avatar

    DAMO-NLP-SG/M3Exam

    105View on GitHub↗

    Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"

    Pythonai-educationchatgptevaluation
    View on GitHub↗105
  • dcdmllm/cheetahDCDmllm avatar

    DCDmllm/Cheetah

    354View on GitHub↗

    Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

    Python
    View on GitHub↗354
  • deepseek-ai/deepseek-vl2deepseek-ai avatar

    deepseek-ai/DeepSeek-VL2

    5,302View on GitHub↗

    DeepSeek-VL2 is a multimodal large language model and vision-language system designed to analyze visual scenes and generate descriptive text. It functions as a visual question answering and visual grounding model, capable of extracting information from documents and locating specific objects or regions within images based on textual descriptions. The project utilizes a mixture-of-experts architecture to process combined image and text inputs. It is optimized for inference through incremental prefilling, which reduces the GPU memory requirements on hardware. The model covers multimodal data a

    Python
    View on GitHub↗5,302
  • freedomintelligence/mllm-benchFreedomIntelligence avatar

    FreedomIntelligence/MLLM-Bench

    76View on GitHub↗

    MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

    Python
    View on GitHub↗76
  • fuxiaoliu/lrv-instructionFuxiaoLiu avatar

    FuxiaoLiu/LRV-Instruction

    297View on GitHub↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Pythonchatgptevaluationevaluation-metrics
    View on GitHub↗297
  • fuxiaoliu/mmcFuxiaoLiu avatar

    FuxiaoLiu/MMC

    95View on GitHub↗

    NAACL 2024 MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning

    Pythonarxivbenchmarkchart
    View on GitHub↗95
  • gzcch/bingogzcch avatar

    gzcch/Bingo

    55View on GitHub↗

    Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution

    View on GitHub↗55
  • haotian-liu/llavahaotian-liu avatar

    haotian-liu/LLaVA

    24,465View on GitHub↗

    LLaVA is a multimodal large language model architecture designed to process and interpret both image and text inputs to generate natural language responses. It functions as a research-oriented platform for visual instruction tuning, providing a framework to align language models with human intent through training on diverse datasets of paired images and text queries. The system distinguishes itself through a specialized vision-language training pipeline that connects visual data to language models using projection layers and instruction-based fine-tuning. It supports distributed inference by

    Pythonchatbotchatgptfoundation-models
    View on GitHub↗24,465
  • hypjudy/sparklesHYPJUDY avatar

    HYPJUDY/Sparkles

    45View on GitHub↗

    Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models

    Python
    View on GitHub↗45
  • inst-it/inst-itinst-it avatar

    inst-it/inst-it

    40View on GitHub↗

    NeurIPS 2025 The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning"

    Pythoninstruction-tuninglarge-multimodal-modelsmultimodal
    View on GitHub↗40
  • katha-ai/velocitikatha-ai avatar

    katha-ai/VELOCITI

    8View on GitHub↗

    VELOCITI Benchmark Evaluation and Visualisation Code

    Pythonartificial-intelligenceawesome-listbenchmark
    View on GitHub↗8
  • lerogo/mmgenbenchlerogo avatar

    lerogo/MMGenBench

    119View on GitHub↗

    Official repository of MMGenBench

    Pythonllms-benchmarkingmllmmmgenbench
    View on GitHub↗119
  • lightchen233/m3cotLightChen233 avatar

    LightChen233/M3CoT

    91View on GitHub↗

    @Author: Qiguang Chen @LastEditors: Qiguang Chen @Date: 2024-05-23 20:24:16 @LastEditTime: 2024-05-26 18:09:00 @Description: -->

    Python
    View on GitHub↗91
  • llyx97/tempcompassllyx97 avatar

    llyx97/TempCompass

    132View on GitHub↗

    ACL 2024 Findings "TempCompass: Do Video LLMs Really Understand Videos?", Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, Lu Hou

    Pythonevaluationtemporal-perceptionvideo-llms
    View on GitHub↗132
  • mathvision-cuhk/mathvisionmathvision-cuhk avatar

    mathvision-cuhk/MathVision

    139View on GitHub↗

    NeurIPS 2024 MATH-Vision dataset and code to measure multimodal mathematical reasoning capabilities.

    Python
    View on GitHub↗139
  • microsoft/unilmmicrosoft avatar

    microsoft/unilm

    22,030View on GitHub↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Pythonbeitbeit-3bitnet
    View on GitHub↗22,030
  • ncsoft/idkncsoft avatar

    ncsoft/idk

    6View on GitHub↗

    Official implementation of "Visually Dehallucinative Instruction Generation: Know What You Don't Know"

    View on GitHub↗6
  • nvlabs/vilaNVlabs avatar

    NVlabs/VILA

    3,819View on GitHub↗

    VILA is a vision-language model integration that combines a visual encoder with a large language model to process images and text in a shared space. Its primary purpose is to enable the generation of natural language explanations and detailed text summaries of images and videos based on user prompts. The project utilizes a multi-stage alignment pipeline to synchronize visual and textual embeddings through sequential pretraining and supervised fine-tuning. To support deployment on desktop and edge hardware, it employs quantized low-precision inference to reduce model weights to 4-bit precision

    Python
    View on GitHub↗3,819
  • open-compass/mmbenchopen-compass avatar

    open-compass/MMBench

    303View on GitHub↗

    Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"

    View on GitHub↗303
  • openbmb/minicpm-oOpenBMB avatar

    OpenBMB/MiniCPM-o

    23,850View on GitHub↗

    MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out

    Pythonminicpmminicpm-vmulti-modal
    View on GitHub↗23,850
  • opengvlab/ask-anythingOpenGVLab avatar

    OpenGVLab/Ask-Anything

    3,341View on GitHub↗

    CVPR2024 HighlightVideoChatGPT ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.

    Pythonbig-modelcaptioning-videoschat
    View on GitHub↗3,341
  • opengvlab/internvlOpenGVLab avatar

    OpenGVLab/InternVL

    10,061View on GitHub↗

    InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling

    Pythongptgpt-4ogpt-4v
    View on GitHub↗10,061
  • opengvlab/llama-adapterOpenGVLab avatar

    OpenGVLab/LLaMA-Adapter

    5,921View on GitHub↗

    LLaMA-Adapter is a parameter-efficient fine-tuning framework designed to adapt large language models using a minimal set of trainable parameters. It functions as an instruction tuning tool and a multimodal adapter, allowing pre-trained models to follow human instructions and process non-textual data. The project specializes in the integration of image, video, audio, and sensor data into language models for cross-modal understanding. It enables the customization of LLaMA models through the use of lightweight adapters, which allows for the extraction and storage of learned weights independently

    Python
    View on GitHub↗5,921
  • opengvlab/multi-modality-arenaOpenGVLab avatar

    OpenGVLab/Multi-Modality-Arena

    558View on GitHub↗
    Pythonchatchatbotchatgpt
    View on GitHub↗558