28 repository-uri
Large-scale raw text corpora for foundational model training.
Explore 28 awesome GitHub repositories matching part of an awesome list · Pre-training Datasets. Refine with filters or upvote what's useful.
RedPajama-Data este un set de instrumente pentru preprocesarea seturilor de date text la scară largă utilizate pentru antrenarea modelelor de limbaj mari (LLM). Oferă un pipeline de preprocesare axat pe curățarea, deduplicarea și punctarea unor colecții masive de text pentru a asigura calitatea și diversitatea datelor. Proiectul utilizează un framework de punctare a calității documentelor care folosește machine learning și euristici statistice pentru a evalua dacă documentele sunt potrivite pentru antrenare. Include un pipeline de filtrare a seturilor de date care utilizează clasificatori și liste de blocare pentru a elimina cuvintele sau URL-urile nedorite. Sistemul dispune de un set de instrumente de deduplicare a textului care elimină conținutul redundant folosind tehnici de potrivire exactă și fuzzy. Aceste capabilități permit identificarea și eliminarea documentelor duplicate sau aproape identice dintr-un corpus.
Reproduced dataset for training large-scale language models.
MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS
Massive, diverse Chinese text corpus from internet sources.
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
Multimodal in-context instruction tuning data.
ECCV2024 Video Foundation Models & Data for Multimodal Understanding
Video-centric instruction dataset for chat-based understanding.
Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.
Biomedical instruction-following dataset for vision-language models.
NeurIPS 2023 MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs
Human motion-related instruction tuning tasks.
Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration
Multi-turn dialogue dataset for multimodal integration.
ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
High-quality video instruction dataset for detailed understanding.
COYO-700M: Large-scale Image-Text Pair Dataset
Large-scale image-text pairs for pre-training.
Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料
Cleaned 100GB Chinese corpus for pre-training and NLP tasks.
This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l
Provides a large-scale Chinese language text collection including dialogue and fiction for training large language models.
2023-06-13 Added tuned linear weights for Vicuna-7b. 2023-05-25 Our paper is available at this link. 2023-05-09 We have launched our project website. 2023-05-08 The first version of DetGPT is available now! Try our demo.
Query-answer pairs for visual detection and reasoning.
GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.
Tool-use instruction datasets for language models.
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Chinese multimodal instruction dataset for foreign language treatment.
NeurIPS 2023 Datasets and Benchmarks Track LAMM: Multi-Modal Large Language Models and Applications as AI Agents
Comprehensive dataset for language-assisted multimodal tuning.
ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Robust instruction tuning to mitigate model hallucinations.
Large-scale dataset used for training general-purpose language models.
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Benchmark dataset for multimodal zero-shot learning.
Official repo for StableLLAVA
Synthesized image-dialogue data for instruction tuning.
👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)
Time-sensitive video understanding instruction dataset.