awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
OpenGVLab avatar

OpenGVLab/InternVL

0
View on GitHub↗
10,061 نجوم·783 تفرعات·Python·MIT·10 مشاهداتinternvl.readthedocs.io/en/latest↗

InternVL

InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions.

The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling for video understanding and provides zero-shot capabilities for image classification and multilingual cross-modal retrieval.

The framework covers a broad range of capabilities including optical character recognition, object localization, and semantic image segmentation. It supports distributed multimodal training and fine-tuning via low-rank adaptation, as well as performance optimizations such as weight quantization and model distillation.

Deployment is supported through an OpenAI-compatible REST interface, a web-based chat interface, and a command-line interface with multi-GPU layer distribution.

Features

  • Vision-Language Models - Fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning.
  • Multimodal Inference - Provides a framework for multimodal inference, processing images and text to generate descriptions or answer questions.
  • Chain-of-Thought Modules - Provides step-by-step logical derivations to solve complex mathematical and spatial problems in visual contexts.
  • Visual Mathematical Reasoning - Applies chain-of-thought reasoning to solve complex quantitative problems based on visual and textual inputs.
  • Representative Frame Sampling - Supports temporal frame sampling to understand events within video content.
  • Image Tiling - Uses dynamic tiling to divide high-resolution images into adaptive segments, supporting detailed analysis up to 4K resolution.
  • Dynamic Tiling - Uses dynamic tiling to maintain detail for images up to 4K resolution.
  • Multimodal Dialogue and Interaction - Supports interactive conversations that use one or more images as visual context for the dialogue.
  • Visual - Responds to natural language questions about images using chart analysis, document extraction, and external knowledge.
  • Adapter Fine-Tuning - Implements Low-Rank Adaptation to update a small subset of parameters for memory-efficient model adaptation.
  • OpenAI-Compatible APIs - Exposes model inference via standard OpenAI-compatible HTTP endpoints for external client integration.
  • Text-Based Object Localization - Identifies the bounding box coordinates of specific regions in an image based on text descriptions.
  • Conversation State Managers - Maintains interaction history and context across a sequence of multimodal exchanges to support follow-up questions.
  • Distributed Training - Supports distributed multimodal fine-tuning across multiple compute nodes and GPUs to optimize performance.
  • Domain-Specific Reasoning Evaluation - Evaluates high-level reasoning performance in specialized fields including academic papers and scientific diagrams.
  • Hallucination Detection - Implements mechanisms to detect and score the tendency of models to produce descriptions of non-existent objects.
  • Model Fine-Tuning - Provides procedures for adapting pre-trained models to specific datasets or tasks using parameter fine-tuning.
  • Model Distillation Frameworks - Runs distilled versions of multimodal models to reduce memory requirements for consumer-grade hardware.
  • Multi-GPU Distribution - Splits model layers across multiple GPUs to execute parameters exceeding single-device memory capacity.
  • Model Quantization - Applies eight-bit quantization to lower the memory footprint during the inference process.
  • Long Multimodal Contexts - Handles extended sequences of images and text using flexible position encoding to maintain long-term context.
  • Optical Character Recognition - Recognizes and extracts text content from images using end-to-end optical character recognition.
  • Parameter Efficient Fine-Tuning - Supports low-rank adaptation (LoRA) to efficiently fine-tune pre-trained weights on multimodal datasets.
  • Precision Quantization - Compresses model weights to four-bit or eight-bit precision to reduce memory footprint and accelerate inference.
  • Preference Optimization - Implements preference optimization to align model reasoning by learning relative quality between response pairs.
  • Weight Quantization - The project uses four-bit quantization to accelerate inference speed and reduce the total amount of graphics memory required.
  • Synthetic Reasoning Data Generators - Generates synthetic reasoning data pairs by sampling negative responses to improve model logical capabilities.
  • Training Configurations - Manages training data distribution through sampling frequency, image tiling, and conditional augmentation.
  • Visual Data Reasoning Evaluation - Evaluates logical and arithmetic reasoning based on visual data representations like charts and diagrams.
  • Visual Mathematical Reasoning Evaluation - Tests the model's ability to solve mathematical problems and logical puzzles presented within visual contexts.
  • Visual Question Answering Evaluation - Tests the ability to answer questions based on images, incorporating external knowledge and specialized content.
  • Visual Spatial Reasoning Evaluation - Tests the understanding of physical environments, object locations, and spatial relations in real-world scenes.
  • Image Captioning - Generates descriptive text summaries for images in a zero-shot manner without requiring task-specific training.
  • Video Understanding - Provides benchmarks and capabilities for assessing temporal comprehension and sequential visual data processing in long-form videos.
  • OCR Accuracy Evaluators - Tests the accuracy of recognizing and extracting text from documents, infographics, and handwritten math expressions.
  • Multimodal Training Data Formatters - Implements JSONL-based data formatting to support text, single-image, multi-image, and video inputs for training.
  • Multi-GPU Layer Distribution - Splits model layers across multiple graphics cards to execute large models that exceed single-device memory.
  • Inference Load Balancers - Distributes model workers across multiple GPUs to balance inference load and optimize performance.
  • Comparative Visual Analysis - Enables the processing of multiple images in one prompt for comparative and collective visual analysis.
  • Video Analysis and Processing - Analyzes sequences of video frames to understand temporal content and answer questions about the video.
  • 3D Spatial Reasoning - Perceives and reasons about three-dimensional visual information to understand complex spatial relationships.
  • Vision-Language Model Benchmarking - Measures accuracy across vision-language benchmarks using standardized evaluation kits and chain-of-thought prompting.
  • Web Chat Interfaces - Provides a browser-based chat interface by launching a coordinated controller, worker, and server.
  • Frontier Reasoning Models - Multimodal reasoning model with advanced training recipes.
  • Multimodal Foundation Models - Scalable vision-language foundation model for generic tasks.
  • Multimodal LLM Models - Multimodal model achieving high performance on multidisciplinary benchmarks.
  • Multimodal Models - Large-scale multimodal model for visual and textual reasoning.
  • Vision Language Models - Advanced multimodal series using a scalable ViT-MLP-LLM architecture.

سجل النجوم

مخطط تاريخ النجوم لـ opengvlab/internvlمخطط تاريخ النجوم لـ opengvlab/internvl

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة opengvlab/internvl؟

InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions.

ما هي الميزات الرئيسية لـ opengvlab/internvl؟

الميزات الرئيسية لـ opengvlab/internvl هي: Vision-Language Models, Multimodal Inference, Chain-of-Thought Modules, Visual Mathematical Reasoning, Representative Frame Sampling, Image Tiling, Dynamic Tiling, Multimodal Dialogue and Interaction.

ما هي البدائل مفتوحة المصدر لـ opengvlab/internvl؟

تشمل البدائل مفتوحة المصدر لـ opengvlab/internvl: hiyouga/llama-efficient-tuning — This project is a fine-tuning framework and training pipeline designed to optimize and adapt large language and vision… intel/ipex-llm — Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning… sgl-project/sglang — Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It… zhaochenyang20/awesome-ml-sys-tutorial — This project provides a comprehensive technical guide and framework for engineering large-scale machine learning… hiyouga/llama-factory — LLaMA-Factory is a comprehensive suite for dataset preparation, model fine-tuning, memory optimization, and… openrlhf/openrlhf — OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across…

بدائل مفتوحة المصدر لـ InternVL

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع InternVL.
  • hiyouga/llama-efficient-tuningالصورة الرمزية لـ hiyouga

    hiyouga/LLaMA-Efficient-Tuning

    72,239عرض على GitHub↗

    This project is a fine-tuning framework and training pipeline designed to optimize and adapt large language and vision models. It provides a specialized toolkit for parameter-efficient tuning and supervised learning, serving as both a trainer for multimodal models and a deployment tool for serving fine-tuned models via high-performance inference engines. The framework focuses on reducing memory and compute requirements by updating a small subset of model parameters. It supports a wide range of adaptation strategies, including vision-language model training to align text, image, video, and aud

    Python
    عرض على GitHub↗72,239
  • intel/ipex-llmالصورة الرمزية لـ intel

    intel/ipex-llm

    8,836عرض على GitHub↗

    Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning on Intel CPUs, GPUs, and NPUs. It provides a distributed inference engine for scaling models across multiple accelerators, a multimodal model runtime for vision and speech tasks, and a low-bit model quantization tool for converting weights into INT4, FP8, and GGUF formats. The project features a parameter-efficient finetuning framework that enables model adaptation using QLoRA and DPO on Intel hardware. It distinguishes itself by providing specialized optimizations for Intel XP

    Python
    عرض على GitHub↗8,836
  • sgl-project/sglangالصورة الرمزية لـ sgl-project

    sgl-project/sglang

    29,079عرض على GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Pythonattentionblackwellcuda
    عرض على GitHub↗29,079
  • zhaochenyang20/awesome-ml-sys-tutorialالصورة الرمزية لـ zhaochenyang20

    zhaochenyang20/Awesome-ML-SYS-Tutorial

    5,371عرض على GitHub↗

    This project provides a comprehensive technical guide and framework for engineering large-scale machine learning systems. It covers the full lifecycle of model development, focusing on the infrastructure and computational principles required to build, train, and serve generative AI models across distributed GPU clusters. The repository distinguishes itself by offering deep-dive tutorials and implementation strategies for complex system challenges. It emphasizes high-performance architectural primitives, such as collective communication orchestration, distributed tensor sharding, and static gr

    Python
    عرض على GitHub↗5,371
  • عرض جميع البدائل الـ 30 لـ InternVL→