This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l
ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Les fonctionnalités principales de vt-nlp/multiinstruct sont : Pre-training Datasets.
Les alternatives open-source à vt-nlp/multiinstruct incluent : plexpt/chatgpt-corpus — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of… fuxiaoliu/lrv-instruction — [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning. guoyang9/unk-vqa — A VQA dataset that includes unanswerable questions [TPAMI 2024]. hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models. icoz69/stablellava — Official repo for StableLLAVA. esbatmop/mnbvc — MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and…