SciGraphQA: Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
Data and code for NeurIPS 2022 Paper "Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering".
ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
[AINL 2023] IMAD: IMage Augmented multi-modal Dialogue
vityavitalich/imad 的主要功能包括:Multimodal Datasets。
vityavitalich/imad 的开源替代品包括: findalexli/scigraphqa — SciGraphQA: Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs. luodian/otter — 🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT… lupantech/scienceqa — Data and code for NeurIPS 2022 Paper "Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question… mbzuai-oryx/video-chatgpt — [ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos.… zeroqiaoba/explainable-multimodal-emotion-reasoning — EMER, OV-MER (ICML25), AffectGPT (ICML25, Oral), EmoPrefer (ICLR26).