awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to lupantech/scienceqa

Open-source alternatives to ScienceQA

19 open-source projects similar to lupantech/scienceqa, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best ScienceQA alternative.

  • zeroqiaoba/explainable-multimodal-emotion-reasoningzeroQiaoba 的头像

    zeroQiaoba/Explainable-Multimodal-Emotion-Reasoning

    396在 GitHub 上查看↗

    EMER, OV-MER (ICML25), AffectGPT (ICML25, Oral), EmoPrefer (ICLR26)

    Python
    在 GitHub 上查看↗396
  • om-ai-lab/vlm-r1om-ai-lab 的头像

    om-ai-lab/VLM-R1

    5,991在 GitHub 上查看↗

    VLM-R1 is a reasoning vision-language model and embodied AI framework designed to map visual inputs and language instructions into physical navigation waypoints and robotic actions. It functions as a multimodal policy optimizer and an open vocabulary detector capable of locating objects based on arbitrary natural language descriptions. The system distinguishes itself through the use of chain-of-thought reasoning and reinforcement learning to solve complex visual and spatial tasks. It utilizes a video semantic memory system, which employs a visual cache to maintain a history of live video for

    Python
    在 GitHub 上查看↗5,991
  • phodal/prompt-patternsphodal 的头像

    phodal/prompt-patterns

    3,096在 GitHub 上查看↗

    Prompt patterns is a framework for organizing AI-driven system design through structured prompt engineering and domain-driven development methodologies. It provides a library of standardized interaction strategies designed to improve the consistency, accuracy, and logical reasoning of large language model outputs. By applying these patterns, users can translate complex business scenarios into structured domain models and technical specifications. The project distinguishes itself by integrating domain-driven design principles directly into the prompting workflow. It utilizes techniques such as

    chatgptgithub-copilotprompt-engineering
    在 GitHub 上查看↗3,096
  • qwenlm/qwen2-vlQwenLM 的头像

    QwenLM/Qwen2-VL

    19,404在 GitHub 上查看↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Jupyter Notebook
    在 GitHub 上查看↗19,404

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • jacoblee93/fully-local-pdf-chatbotjacoblee93 的头像

    jacoblee93/fully-local-pdf-chatbot

    1,813在 GitHub 上查看↗

    This project is a private document analysis tool that enables conversational interaction with PDF files by executing all language model inference and processing entirely on the local machine. By running models directly within the browser or local environment, it ensures that sensitive user data remains offline and inaccessible to external servers or third-party cloud providers. The system utilizes retrieval augmented generation to provide context-aware answers, supported by local document text extraction and vector embedding indexing. This architecture allows for semantic search and informati

    TypeScript
    在 GitHub 上查看↗1,813
  • pandabearlab/prompt-tutorialPandaBearLab 的头像

    PandaBearLab/prompt-tutorial

    1,330在 GitHub 上查看↗

    This project serves as an educational resource and guide for prompt engineering, providing a structured methodology for interacting with large language models. It focuses on teaching core strategies to improve the reliability, accuracy, and consistency of model outputs across a variety of natural language processing tasks. The framework emphasizes the use of standardized templates and logical decomposition to manage complex instructions. By implementing techniques such as few-shot context injection, iterative refinement, and delimiter-based segmentation, the project demonstrates how to guide

    在 GitHub 上查看↗1,330
  • ggg0919/cantorggg0919 的头像

    ggg0919/cantor

    90在 GitHub 上查看↗

    Project Page | Paper

    HTML
    在 GitHub 上查看↗90
  • luodian/otterLuodian 的头像

    Luodian/Otter

    3,410在 GitHub 上查看↗

    🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

    Python
    在 GitHub 上查看↗3,410
  • amazon-science/mm-cotamazon-science 的头像

    amazon-science/mm-cot

    3,990在 GitHub 上查看↗

    This project is a multimodal large language model reasoning framework designed to train and evaluate models in performing chain-of-thought reasoning across text and image data. It provides a reasoning engine and training system that enable vision-language models to generate step-by-step logical rationales and final answers for complex queries. The framework utilizes a two-stage training pipeline that decouples the generation of logical justifications from final answer inference. It transforms visual data into descriptive text through image captioning and uses vision-transformer feature extrac

    Python
    在 GitHub 上查看↗3,990
  • shikras/shikrashikras 的头像

    shikras/shikra

    813在 GitHub 上查看↗

    Shikra : Unleashing Multimodal LLM’s Referential Dialogue Magic

    Python
    在 GitHub 上查看↗813
  • soolab/ddcotSooLab 的头像

    SooLab/DDCOT

    48在 GitHub 上查看↗

    NeurIPS 2023DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

    Python
    在 GitHub 上查看↗48
  • ttengwang/caption-anythingttengwang 的头像

    ttengwang/Caption-Anything

    1,774在 GitHub 上查看↗

    Caption-Anything is a versatile tool combining image segmentation, visual captioning, and ChatGPT, generating tailored captions with diverse controls for user preferences. https://huggingface.co/spaces/TencentARC/Caption-Anything https://huggingface.co/spaces/VIPLab/Caption-Anything

    Python
    在 GitHub 上查看↗1,774
  • vityavitalich/imadVityaVitalich 的头像

    VityaVitalich/IMAD

    4在 GitHub 上查看↗

    AINL 2023 IMAD: IMage Augmented multi-modal Dialogue

    Python
    在 GitHub 上查看↗4
  • mbzuai-oryx/video-chatgptmbzuai-oryx 的头像

    mbzuai-oryx/Video-ChatGPT

    1,504在 GitHub 上查看↗

    ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

    Pythonchatbotclipgpt-4
    在 GitHub 上查看↗1,504
  • chancharikmitra/ccotchancharikmitra 的头像

    chancharikmitra/CCoT

    143在 GitHub 上查看↗

    CVPR 2024 Official Code for the Paper "Compositional Chain-of-Thought Prompting for Large Multimodal Models"

    Python
    在 GitHub 上查看↗143
  • dannyrose30/vcotD

    dannyrose30/VCOT

    0在 GitHub 上查看↗
    在 GitHub 上查看↗0
  • deepcs233/visual-cotdeepcs233 的头像

    deepcs233/Visual-CoT

    445在 GitHub 上查看↗

    Neurips'24 Spotlight Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

    Python
    在 GitHub 上查看↗445
  • dongyh20/insight-vdongyh20 的头像

    dongyh20/Insight-V

    238在 GitHub 上查看↗

    CVPR2025 Highlight Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

    Python
    在 GitHub 上查看↗238
  • findalexli/scigraphqafindalexli 的头像

    findalexli/SciGraphQA

    43在 GitHub 上查看↗

    SciGraphQA: Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs

    Jupyter Notebookdatasetsllmsynthetic-data
    在 GitHub 上查看↗43