awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

28 repository-uri

Awesome GitHub RepositoriesPre-training Datasets

Large-scale raw text corpora for foundational model training.

Explore 28 awesome GitHub repositories matching part of an awesome list · Pre-training Datasets. Refine with filters or upvote what's useful.

Awesome Pre-training Datasets GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • togethercomputer/redpajama-dataAvatar togethercomputer

    togethercomputer/RedPajama-Data

    4,947Vezi pe GitHub↗

    RedPajama-Data este un set de instrumente pentru preprocesarea seturilor de date text la scară largă utilizate pentru antrenarea modelelor de limbaj mari (LLM). Oferă un pipeline de preprocesare axat pe curățarea, deduplicarea și punctarea unor colecții masive de text pentru a asigura calitatea și diversitatea datelor. Proiectul utilizează un framework de punctare a calității documentelor care folosește machine learning și euristici statistice pentru a evalua dacă documentele sunt potrivite pentru antrenare. Include un pipeline de filtrare a seturilor de date care utilizează clasificatori și liste de blocare pentru a elimina cuvintele sau URL-urile nedorite. Sistemul dispune de un set de instrumente de deduplicare a textului care elimină conținutul redundant folosind tehnici de potrivire exactă și fuzzy. Aceste capabilități permit identificarea și eliminarea documentelor duplicate sau aproape identice dintr-un corpus.

    Reproduced dataset for training large-scale language models.

    Python
    Vezi pe GitHub↗4,947
  • esbatmop/mnbvcAvatar esbatmop

    esbatmop/MNBVC

    4,123Vezi pe GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    Massive, diverse Chinese text corpus from internet sources.

    chinesechinese-languagechinese-nlp
    Vezi pe GitHub↗4,123
  • luodian/otterAvatar Luodian

    Luodian/Otter

    3,410Vezi pe GitHub↗

    🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

    Multimodal in-context instruction tuning data.

    Python
    Vezi pe GitHub↗3,410
  • opengvlab/internvideoAvatar OpenGVLab

    OpenGVLab/InternVideo

    2,292Vezi pe GitHub↗

    ECCV2024 Video Foundation Models & Data for Multimodal Understanding

    Video-centric instruction dataset for chat-based understanding.

    Pythonaction-recognitionbenchmarkcontrastive-learning
    Vezi pe GitHub↗2,292
  • microsoft/llava-medAvatar microsoft

    microsoft/LLaVA-Med

    2,214Vezi pe GitHub↗

    Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.

    Biomedical instruction-following dataset for vision-language models.

    Python
    Vezi pe GitHub↗2,214
  • openmotionlab/motiongptAvatar OpenMotionLab

    OpenMotionLab/MotionGPT

    1,938Vezi pe GitHub↗

    NeurIPS 2023 MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs

    Human motion-related instruction tuning tasks.

    Python3d-generationchatgptgpt
    Vezi pe GitHub↗1,938
  • lyuchenyang/macaw-llmAvatar lyuchenyang

    lyuchenyang/Macaw-LLM

    1,590Vezi pe GitHub↗

    Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration

    Multi-turn dialogue dataset for multimodal integration.

    Pythondeep-learninglanguage-modelmachine-learning
    Vezi pe GitHub↗1,590
  • mbzuai-oryx/video-chatgptAvatar mbzuai-oryx

    mbzuai-oryx/Video-ChatGPT

    1,504Vezi pe GitHub↗

    ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

    High-quality video instruction dataset for detailed understanding.

    Pythonchatbotclipgpt-4
    Vezi pe GitHub↗1,504
  • kakaobrain/coyo-datasetAvatar kakaobrain

    kakaobrain/coyo-dataset

    1,256Vezi pe GitHub↗

    COYO-700M: Large-scale Image-Text Pair Dataset

    Large-scale image-text pairs for pre-training.

    Python
    Vezi pe GitHub↗1,256
  • cluebenchmark/cluecorpus2020Avatar CLUEbenchmark

    CLUEbenchmark/CLUECorpus2020

    1,012Vezi pe GitHub↗

    Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

    Cleaned 100GB Chinese corpus for pre-training and NLP tasks.

    albertbertchinese
    Vezi pe GitHub↗1,012
  • plexpt/chatgpt-corpusAvatar PlexPt

    PlexPt/chatgpt-corpus

    964Vezi pe GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    Provides a large-scale Chinese language text collection including dialogue and fiction for training large language models.

    awesomecorpuscorpus-data
    Vezi pe GitHub↗964
  • optimalscale/detgptAvatar OptimalScale

    OptimalScale/DetGPT

    788Vezi pe GitHub↗

    2023-06-13 Added tuned linear weights for Vicuna-7b. 2023-05-25 Our paper is available at this link. 2023-05-09 We have launched our project website. 2023-05-08 The first version of DetGPT is available now! Try our demo.

    Query-answer pairs for visual detection and reasoning.

    Jupyter Notebook
    Vezi pe GitHub↗788
  • stevengrove/gpt4toolsAvatar StevenGrove

    StevenGrove/GPT4Tools

    771Vezi pe GitHub↗

    GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.

    Tool-use instruction datasets for language models.

    Python
    Vezi pe GitHub↗771
  • phellonchen/x-llmAvatar phellonchen

    phellonchen/X-LLM

    318Vezi pe GitHub↗

    X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

    Chinese multimodal instruction dataset for foreign language treatment.

    Python
    Vezi pe GitHub↗318
  • openlamm/lammAvatar OpenLAMM

    OpenLAMM/LAMM

    317Vezi pe GitHub↗

    NeurIPS 2023 Datasets and Benchmarks Track LAMM: Multi-Modal Large Language Models and Applications as AI Agents

    Comprehensive dataset for language-assisted multimodal tuning.

    Python
    Vezi pe GitHub↗317
  • fuxiaoliu/lrv-instructionAvatar FuxiaoLiu

    FuxiaoLiu/LRV-Instruction

    297Vezi pe GitHub↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Robust instruction tuning to mitigate model hallucinations.

    Pythonchatgptevaluationevaluation-metrics
    Vezi pe GitHub↗297
  • mobvoi/seq-monkey-dataAvatar mobvoi

    mobvoi/seq-monkey-data

    178Vezi pe GitHub↗

    Large-scale dataset used for training general-purpose language models.

    Vezi pe GitHub↗178
  • vt-nlp/multiinstructAvatar VT-NLP

    VT-NLP/MultiInstruct

    135Vezi pe GitHub↗

    MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

    Benchmark dataset for multimodal zero-shot learning.

    Python
    Vezi pe GitHub↗135
  • icoz69/stablellavaAvatar icoz69

    icoz69/StableLLAVA

    95Vezi pe GitHub↗

    Official repo for StableLLAVA

    Synthesized image-dialogue data for instruction tuning.

    Python
    Vezi pe GitHub↗95
  • polyu-chenlab/etbenchAvatar PolyU-ChenLab

    PolyU-ChenLab/ETBench

    74Vezi pe GitHub↗

    👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)

    Time-sensitive video understanding instruction dataset.

    Python
    Vezi pe GitHub↗74
Înapoi12Înainte
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Pre-training Datasets