awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
mbzuai-oryx avatar

mbzuai-oryx/Video-ChatGPT

0
View on GitHub↗
1,504 星标·129 分支·Python·CC-BY-4.0·5 次浏览mbzuai-oryx.github.io/Video-ChatGPT↗

Video ChatGPT

[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

Features

  • Multimodal Datasets - Quantitative evaluation framework for video-based dialogue.
  • Video Understanding Models - Detailed video understanding via vision-language integration.
  • Pre-training Datasets - High-quality video instruction dataset for detailed understanding.

Star 历史

mbzuai-oryx/video-chatgpt 的 Star 历史图表mbzuai-oryx/video-chatgpt 的 Star 历史图表

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Start searching with AI

Video ChatGPT 的开源替代方案

相似的开源项目,按与 Video ChatGPT 的功能重合度排序。
  • luodian/otterLuodian 的头像

    Luodian/Otter

    3,410在 GitHub 上查看↗

    🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

    Python
    在 GitHub 上查看↗3,410
  • nvidia/cosmosNVIDIA 的头像

    NVIDIA/cosmos

    10,494在 GitHub 上查看↗

    Cosmos is an open platform of world models, datasets, and tools for building physical AI systems such as robots and autonomous vehicles. It provides video generation and video understanding models that can generate synthetic videos and world simulations from text, image, video, or action inputs, and analyze videos to produce captions, event timestamps, spatial bounding boxes, and next-action predictions. The platform includes a world simulation generator that produces images, videos, synchronized audio, and action-conditioned rollouts for synthetic data, alongside a visual content analyzer th

    Jupyter Notebook
    在 GitHub 上查看↗10,494
  • llava-vl/llava-nextLLaVA-VL 的头像

    LLaVA-VL/LLaVA-NeXT

    4,695在 GitHub 上查看↗

    LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images and video sequences to generate text. It functions as a visual language model that combines vision encoders with language models to perform complex reasoning, question answering, and video understanding. The system is capable of analyzing high-resolution images and temporal video frames to describe events, summarize actions, and reason across multiple visual inputs. It supports the interpretation of documents and charts, spatial environment analysis, and the generation of desc

    Python
    在 GitHub 上查看↗4,695
  • openbmb/minicpm-vOpenBMB 的头像

    OpenBMB/MiniCPM-V

    25,653在 GitHub 上查看↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Python
    在 GitHub 上查看↗25,653
查看 Video ChatGPT 的所有 30 个替代方案→

常见问题解答

mbzuai-oryx/video-chatgpt 是做什么的?

[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

mbzuai-oryx/video-chatgpt 的主要功能有哪些?

mbzuai-oryx/video-chatgpt 的主要功能包括:Multimodal Datasets, Video Understanding Models, Pre-training Datasets。

mbzuai-oryx/video-chatgpt 有哪些开源替代品?

mbzuai-oryx/video-chatgpt 的开源替代品包括: luodian/otter — 🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT… qwenlm/qwen2-vl — Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text,… llava-vl/llava-next — LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images… nvidia/cosmos — Cosmos is an open platform of world models, datasets, and tools for building physical AI systems such as robots and… openbmb/minicpm-v — MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic… plexpt/chatgpt-corpus — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of…