awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
mobvoi avatar

mobvoi/seq-monkey-data

0
View on GitHub↗
178 نجوم·8 تفرعات·Apache-2.0·6 مشاهدات

Seq Monkey Data

Features

  • Pre-training Datasets - Large-scale dataset used for training general-purpose language models.

سجل النجوم

مخطط تاريخ النجوم لـ mobvoi/seq-monkey-dataمخطط تاريخ النجوم لـ mobvoi/seq-monkey-data

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Seq Monkey Data

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Seq Monkey Data.
  • plexpt/chatgpt-corpusالصورة الرمزية لـ PlexPt

    PlexPt/chatgpt-corpus

    964عرض على GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    awesomecorpuscorpus-data
    عرض على GitHub↗964
  • fuxiaoliu/lrv-instructionالصورة الرمزية لـ FuxiaoLiu

    FuxiaoLiu/LRV-Instruction

    297عرض على GitHub↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Pythonchatgptevaluationevaluation-metrics
    عرض على GitHub↗297
  • guoyang9/unk-vqaالصورة الرمزية لـ guoyang9

    guoyang9/UNK-VQA

    7عرض على GitHub↗

    A VQA dataset that includes unanswerable questions TPAMI 2024.

    عرض على GitHub↗7
  • esbatmop/mnbvcالصورة الرمزية لـ esbatmop

    esbatmop/MNBVC

    4,123عرض على GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    chinesechinese-languagechinese-nlp
    عرض على GitHub↗4,123
عرض جميع البدائل الـ 27 لـ Seq Monkey Data→

الأسئلة الشائعة

ما هي الميزات الرئيسية لـ mobvoi/seq-monkey-data؟

الميزات الرئيسية لـ mobvoi/seq-monkey-data هي: Pre-training Datasets.

ما هي البدائل مفتوحة المصدر لـ mobvoi/seq-monkey-data؟

تشمل البدائل مفتوحة المصدر لـ mobvoi/seq-monkey-data: plexpt/chatgpt-corpus — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of… fuxiaoliu/lrv-instruction — [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning. guoyang9/unk-vqa — A VQA dataset that includes unanswerable questions [TPAMI 2024]. hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models. icoz69/stablellava — Official repo for StableLLAVA. esbatmop/mnbvc — MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and…