awesome-repositories.comCategoriesBlog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
VT-NLP avatar

VT-NLP/MultiInstruct

0
View on GitHub↗

MultiInstruct

MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Features

  • Pre-training Datasets - Benchmark dataset for multimodal zero-shot learning.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI
135 stars·4 forks·Python·Apache-2.0·9 views

Star history

Star history chart for vt-nlp/multiinstructStar history chart for vt-nlp/multiinstruct

Frequently asked questions

What does vt-nlp/multiinstruct do?

MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

What are the main features of vt-nlp/multiinstruct?

The main features of vt-nlp/multiinstruct are: Pre-training Datasets.

What are some open-source alternatives to vt-nlp/multiinstruct?

Open-source alternatives to vt-nlp/multiinstruct include: plexpt/chatgpt-corpus — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of… fuxiaoliu/lrv-instruction — [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning. guoyang9/unk-vqa — A VQA dataset that includes unanswerable questions [TPAMI 2024]. hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models. icoz69/stablellava — Official repo for StableLLAVA. esbatmop/mnbvc — MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and…

Open-source alternatives to MultiInstruct

Similar open-source projects, ranked by how many features they share with MultiInstruct.
  • plexpt/chatgpt-corpusPlexPt avatar

    PlexPt/chatgpt-corpus

    964View on GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    awesomecorpuscorpus-data
    View on GitHub↗964
  • fuxiaoliu/lrv-instructionFuxiaoLiu avatar

    FuxiaoLiu/LRV-Instruction

    297View on GitHub↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Pythonchatgptevaluationevaluation-metrics
    View on GitHub↗297
  • guoyang9/unk-vqaguoyang9 avatar

    guoyang9/UNK-VQA

    7View on GitHub↗

    A VQA dataset that includes unanswerable questions TPAMI 2024.

    View on GitHub↗7
  • esbatmop/mnbvcesbatmop avatar

    esbatmop/MNBVC

    4,123View on GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    chinesechinese-languagechinese-nlp
    View on GitHub↗4,123
See all 27 alternatives to MultiInstruct→