awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
OpenGVLab avatar

OpenGVLab/InternVL-U

0
View on GitHub↗
291 stele·16 fork-uri·Python·MIT·9 vizualizări

InternVL U

InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image editing into a single framework.

Features

  • Unified Models - Unified vision-language model for diverse tasks.
  • Unified Multimodal Models - Unified vision-language model for multimodal tasks.

Istoric stele

Graficul istoricului de stele pentru opengvlab/internvl-uGraficul istoricului de stele pentru opengvlab/internvl-u

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Întrebări frecvente

Ce face opengvlab/internvl-u?

InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image editing into a single framework.

Care sunt principalele funcționalități ale opengvlab/internvl-u?

Principalele funcționalități ale opengvlab/internvl-u sunt: Unified Models, Unified Multimodal Models.

Care sunt câteva alternative open-source pentru opengvlab/internvl-u?

Alternativele open-source pentru opengvlab/internvl-u includ: bytedance/lance — A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing. alpha-vllm/lumina-dimoo — Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding. byteflow-ai/tokenflow — [CVPR 2025] 🔥 Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation". deepseek-ai/janus — Janus is a multimodal large language model and unified framework that integrates visual understanding and image… facebookresearch/tuna-2 — Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation. lehduong/onediffusion.

Alternative open-source pentru InternVL U

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu InternVL U.
  • bytedance/lanceAvatar bytedance

    bytedance/Lance

    1,250Vezi pe GitHub↗

    A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

    Python
    Vezi pe GitHub↗1,250
  • alpha-vllm/lumina-dimooAvatar Alpha-VLLM

    Alpha-VLLM/Lumina-DiMOO

    1,001Vezi pe GitHub↗

    Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

    Python
    Vezi pe GitHub↗1,001
  • byteflow-ai/tokenflowAvatar ByteFlow-AI

    ByteFlow-AI/TokenFlow

    465Vezi pe GitHub↗

    CVPR 2025 🔥 Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".

    Python
    Vezi pe GitHub↗465
  • deepseek-ai/janusAvatar deepseek-ai

    deepseek-ai/Janus

    17,746Vezi pe GitHub↗

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Pythonany-to-anyfoundation-modelsllm
    Vezi pe GitHub↗17,746
Vezi toate cele 14 alternative pentru InternVL U→