awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

28 مستودعات

Awesome GitHub RepositoriesPre-training Datasets

Large-scale raw text corpora for foundational model training.

Explore 28 awesome GitHub repositories matching part of an awesome list · Pre-training Datasets. Refine with filters or upvote what's useful.

Awesome Pre-training Datasets GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • togethercomputer/redpajama-dataالصورة الرمزية لـ togethercomputer

    togethercomputer/RedPajama-Data

    4,947عرض على GitHub↗

    RedPajama-Data هي مجموعة أدوات لمعالجة مجموعات البيانات النصية واسعة النطاق المستخدمة لتدريب النماذج اللغوية الكبيرة. توفر خط أنابيب معالجة يركز على تنظيف، وإزالة التكرار، وتسجيل مجموعات ضخمة من النصوص لضمان جودة البيانات وتنوعها. يستخدم المشروع إطار عمل لتسجيل جودة المستندات يستخدم التعلم الآلي والاستدلالات الإحصائية لتقييم ما إذا كانت المستندات مناسبة للتدريب. يتضمن خط أنابيب تصفية مجموعات البيانات الذي يستخدم المصنفات والقوائم السوداء لإزالة الكلمات أو روابط URL غير المرغوب فيها. يتميز النظام بمجموعة أدوات لإزالة تكرار النصوص تقضي على المحتوى الزائد باستخدام تقنيات المطابقة الدقيقة والتقريبية. تسمح هذه القدرات بتحديد وإزالة المستندات المكررة أو المتطابقة تقريباً عبر مجموعة بيانات.

    Reproduced dataset for training large-scale language models.

    Python
    عرض على GitHub↗4,947
  • esbatmop/mnbvcالصورة الرمزية لـ esbatmop

    esbatmop/MNBVC

    4,123عرض على GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    Massive, diverse Chinese text corpus from internet sources.

    chinesechinese-languagechinese-nlp
    عرض على GitHub↗4,123
  • luodian/otterالصورة الرمزية لـ Luodian

    Luodian/Otter

    3,410عرض على GitHub↗

    🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

    Multimodal in-context instruction tuning data.

    Python
    عرض على GitHub↗3,410
  • opengvlab/internvideoالصورة الرمزية لـ OpenGVLab

    OpenGVLab/InternVideo

    2,292عرض على GitHub↗

    ECCV2024 Video Foundation Models & Data for Multimodal Understanding

    Video-centric instruction dataset for chat-based understanding.

    Pythonaction-recognitionbenchmarkcontrastive-learning
    عرض على GitHub↗2,292
  • microsoft/llava-medالصورة الرمزية لـ microsoft

    microsoft/LLaVA-Med

    2,214عرض على GitHub↗

    Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.

    Biomedical instruction-following dataset for vision-language models.

    Python
    عرض على GitHub↗2,214
  • openmotionlab/motiongptالصورة الرمزية لـ OpenMotionLab

    OpenMotionLab/MotionGPT

    1,938عرض على GitHub↗

    NeurIPS 2023 MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs

    Human motion-related instruction tuning tasks.

    Python3d-generationchatgptgpt
    عرض على GitHub↗1,938
  • lyuchenyang/macaw-llmالصورة الرمزية لـ lyuchenyang

    lyuchenyang/Macaw-LLM

    1,590عرض على GitHub↗

    Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration

    Multi-turn dialogue dataset for multimodal integration.

    Pythondeep-learninglanguage-modelmachine-learning
    عرض على GitHub↗1,590
  • mbzuai-oryx/video-chatgptالصورة الرمزية لـ mbzuai-oryx

    mbzuai-oryx/Video-ChatGPT

    1,504عرض على GitHub↗

    ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

    High-quality video instruction dataset for detailed understanding.

    Pythonchatbotclipgpt-4
    عرض على GitHub↗1,504
  • kakaobrain/coyo-datasetالصورة الرمزية لـ kakaobrain

    kakaobrain/coyo-dataset

    1,256عرض على GitHub↗

    COYO-700M: Large-scale Image-Text Pair Dataset

    Large-scale image-text pairs for pre-training.

    Python
    عرض على GitHub↗1,256
  • cluebenchmark/cluecorpus2020الصورة الرمزية لـ CLUEbenchmark

    CLUEbenchmark/CLUECorpus2020

    1,012عرض على GitHub↗

    Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

    Cleaned 100GB Chinese corpus for pre-training and NLP tasks.

    albertbertchinese
    عرض على GitHub↗1,012
  • plexpt/chatgpt-corpusالصورة الرمزية لـ PlexPt

    PlexPt/chatgpt-corpus

    964عرض على GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    Provides a large-scale Chinese language text collection including dialogue and fiction for training large language models.

    awesomecorpuscorpus-data
    عرض على GitHub↗964
  • optimalscale/detgptالصورة الرمزية لـ OptimalScale

    OptimalScale/DetGPT

    788عرض على GitHub↗

    2023-06-13 Added tuned linear weights for Vicuna-7b. 2023-05-25 Our paper is available at this link. 2023-05-09 We have launched our project website. 2023-05-08 The first version of DetGPT is available now! Try our demo.

    Query-answer pairs for visual detection and reasoning.

    Jupyter Notebook
    عرض على GitHub↗788
  • stevengrove/gpt4toolsالصورة الرمزية لـ StevenGrove

    StevenGrove/GPT4Tools

    771عرض على GitHub↗

    GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.

    Tool-use instruction datasets for language models.

    Python
    عرض على GitHub↗771
  • phellonchen/x-llmالصورة الرمزية لـ phellonchen

    phellonchen/X-LLM

    318عرض على GitHub↗

    X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

    Chinese multimodal instruction dataset for foreign language treatment.

    Python
    عرض على GitHub↗318
  • openlamm/lammالصورة الرمزية لـ OpenLAMM

    OpenLAMM/LAMM

    317عرض على GitHub↗

    NeurIPS 2023 Datasets and Benchmarks Track LAMM: Multi-Modal Large Language Models and Applications as AI Agents

    Comprehensive dataset for language-assisted multimodal tuning.

    Python
    عرض على GitHub↗317
  • fuxiaoliu/lrv-instructionالصورة الرمزية لـ FuxiaoLiu

    FuxiaoLiu/LRV-Instruction

    297عرض على GitHub↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Robust instruction tuning to mitigate model hallucinations.

    Pythonchatgptevaluationevaluation-metrics
    عرض على GitHub↗297
  • mobvoi/seq-monkey-dataالصورة الرمزية لـ mobvoi

    mobvoi/seq-monkey-data

    178عرض على GitHub↗

    Large-scale dataset used for training general-purpose language models.

    عرض على GitHub↗178
  • vt-nlp/multiinstructالصورة الرمزية لـ VT-NLP

    VT-NLP/MultiInstruct

    135عرض على GitHub↗

    MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

    Benchmark dataset for multimodal zero-shot learning.

    Python
    عرض على GitHub↗135
  • icoz69/stablellavaالصورة الرمزية لـ icoz69

    icoz69/StableLLAVA

    95عرض على GitHub↗

    Official repo for StableLLAVA

    Synthesized image-dialogue data for instruction tuning.

    Python
    عرض على GitHub↗95
  • polyu-chenlab/etbenchالصورة الرمزية لـ PolyU-ChenLab

    PolyU-ChenLab/ETBench

    74عرض على GitHub↗

    👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)

    Time-sensitive video understanding instruction dataset.

    Python
    عرض على GitHub↗74
السابق12التالي
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Pre-training Datasets