awesome-repositories.comCategoríasBlog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

28 repositorios

Awesome GitHub RepositoriesPre-training Datasets

Large-scale raw text corpora for foundational model training.

Explore 28 awesome GitHub repositories matching part of an awesome list · Pre-training Datasets. Refine with filters or upvote what's useful.

Awesome Pre-training Datasets GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • togethercomputer/redpajama-dataAvatar de togethercomputer

    togethercomputer/RedPajama-Data

    4,947Ver en GitHub↗

    RedPajama-Data es un conjunto de herramientas para el preprocesamiento de conjuntos de datos de texto a gran escala utilizados para entrenar modelos de lenguaje grandes. Proporciona una canalización de preprocesamiento centrada en la limpieza, deduplicación y puntuación de colecciones masivas de texto para garantizar la calidad y diversidad de los datos. El proyecto utiliza un framework de puntuación de calidad de documentos que emplea aprendizaje automático y heurísticas estadísticas para evaluar si los documentos son adecuados para el entrenamiento. Incluye una canalización de filtrado de conjuntos de datos que utiliza clasificadores y listas de bloqueo para eliminar palabras o URLs no deseadas. El sistema cuenta con un conjunto de herramientas de deduplicación de texto que elimina contenido redundante utilizando técnicas de coincidencia exacta y difusa. Estas capacidades permiten la identificación y eliminación de documentos duplicados o casi idénticos en un corpus.

    Reproduced dataset for training large-scale language models.

    Python
    Ver en GitHub↗4,947
  • esbatmop/mnbvcAvatar de esbatmop

    esbatmop/MNBVC

    4,123Ver en GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    Massive, diverse Chinese text corpus from internet sources.

    chinesechinese-languagechinese-nlp
    Ver en GitHub↗4,123
  • luodian/otterAvatar de Luodian

    Luodian/Otter

    3,410Ver en GitHub↗

    🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

    Multimodal in-context instruction tuning data.

    Python
    Ver en GitHub↗3,410
  • opengvlab/internvideoAvatar de OpenGVLab

    OpenGVLab/InternVideo

    2,292Ver en GitHub↗

    ECCV2024 Video Foundation Models & Data for Multimodal Understanding

    Video-centric instruction dataset for chat-based understanding.

    Pythonaction-recognitionbenchmarkcontrastive-learning
    Ver en GitHub↗2,292
  • microsoft/llava-medAvatar de microsoft

    microsoft/LLaVA-Med

    2,214Ver en GitHub↗

    Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.

    Biomedical instruction-following dataset for vision-language models.

    Python
    Ver en GitHub↗2,214
  • openmotionlab/motiongptAvatar de OpenMotionLab

    OpenMotionLab/MotionGPT

    1,938Ver en GitHub↗

    NeurIPS 2023 MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs

    Human motion-related instruction tuning tasks.

    Python3d-generationchatgptgpt
    Ver en GitHub↗1,938
  • lyuchenyang/macaw-llmAvatar de lyuchenyang

    lyuchenyang/Macaw-LLM

    1,590Ver en GitHub↗

    Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration

    Multi-turn dialogue dataset for multimodal integration.

    Pythondeep-learninglanguage-modelmachine-learning
    Ver en GitHub↗1,590
  • mbzuai-oryx/video-chatgptAvatar de mbzuai-oryx

    mbzuai-oryx/Video-ChatGPT

    1,504Ver en GitHub↗

    ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

    High-quality video instruction dataset for detailed understanding.

    Pythonchatbotclipgpt-4
    Ver en GitHub↗1,504
  • kakaobrain/coyo-datasetAvatar de kakaobrain

    kakaobrain/coyo-dataset

    1,256Ver en GitHub↗

    COYO-700M: Large-scale Image-Text Pair Dataset

    Large-scale image-text pairs for pre-training.

    Python
    Ver en GitHub↗1,256
  • cluebenchmark/cluecorpus2020Avatar de CLUEbenchmark

    CLUEbenchmark/CLUECorpus2020

    1,012Ver en GitHub↗

    Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

    Cleaned 100GB Chinese corpus for pre-training and NLP tasks.

    albertbertchinese
    Ver en GitHub↗1,012
  • plexpt/chatgpt-corpusAvatar de PlexPt

    PlexPt/chatgpt-corpus

    964Ver en GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    Provides a large-scale Chinese language text collection including dialogue and fiction for training large language models.

    awesomecorpuscorpus-data
    Ver en GitHub↗964
  • optimalscale/detgptAvatar de OptimalScale

    OptimalScale/DetGPT

    788Ver en GitHub↗

    2023-06-13 Added tuned linear weights for Vicuna-7b. 2023-05-25 Our paper is available at this link. 2023-05-09 We have launched our project website. 2023-05-08 The first version of DetGPT is available now! Try our demo.

    Query-answer pairs for visual detection and reasoning.

    Jupyter Notebook
    Ver en GitHub↗788
  • stevengrove/gpt4toolsAvatar de StevenGrove

    StevenGrove/GPT4Tools

    771Ver en GitHub↗

    GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.

    Tool-use instruction datasets for language models.

    Python
    Ver en GitHub↗771
  • phellonchen/x-llmAvatar de phellonchen

    phellonchen/X-LLM

    318Ver en GitHub↗

    X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

    Chinese multimodal instruction dataset for foreign language treatment.

    Python
    Ver en GitHub↗318
  • openlamm/lammAvatar de OpenLAMM

    OpenLAMM/LAMM

    317Ver en GitHub↗

    NeurIPS 2023 Datasets and Benchmarks Track LAMM: Multi-Modal Large Language Models and Applications as AI Agents

    Comprehensive dataset for language-assisted multimodal tuning.

    Python
    Ver en GitHub↗317
  • fuxiaoliu/lrv-instructionAvatar de FuxiaoLiu

    FuxiaoLiu/LRV-Instruction

    297Ver en GitHub↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Robust instruction tuning to mitigate model hallucinations.

    Pythonchatgptevaluationevaluation-metrics
    Ver en GitHub↗297
  • mobvoi/seq-monkey-dataAvatar de mobvoi

    mobvoi/seq-monkey-data

    178Ver en GitHub↗

    Large-scale dataset used for training general-purpose language models.

    Ver en GitHub↗178
  • vt-nlp/multiinstructAvatar de VT-NLP

    VT-NLP/MultiInstruct

    135Ver en GitHub↗

    MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

    Benchmark dataset for multimodal zero-shot learning.

    Python
    Ver en GitHub↗135
  • icoz69/stablellavaAvatar de icoz69

    icoz69/StableLLAVA

    95Ver en GitHub↗

    Official repo for StableLLAVA

    Synthesized image-dialogue data for instruction tuning.

    Python
    Ver en GitHub↗95
  • polyu-chenlab/etbenchAvatar de PolyU-ChenLab

    PolyU-ChenLab/ETBench

    74Ver en GitHub↗

    👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)

    Time-sensitive video understanding instruction dataset.

    Python
    Ver en GitHub↗74
Ant.12Siguiente
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Pre-training Datasets