awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
mobvoi avatar

mobvoi/seq-monkey-data

0
View on GitHub↗
178 Stars·8 Forks·Apache-2.0·6 Aufrufe

Seq Monkey Data

Features

  • Pre-training Datasets - Large-scale dataset used for training general-purpose language models.

Star-Verlauf

Star-Verlauf für mobvoi/seq-monkey-dataStar-Verlauf für mobvoi/seq-monkey-data

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Seq Monkey Data

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Seq Monkey Data.
  • plexpt/chatgpt-corpusAvatar von PlexPt

    PlexPt/chatgpt-corpus

    964Auf GitHub ansehen↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    awesomecorpuscorpus-data
    Auf GitHub ansehen↗964
  • fuxiaoliu/lrv-instructionAvatar von FuxiaoLiu

    FuxiaoLiu/LRV-Instruction

    297Auf GitHub ansehen↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Pythonchatgptevaluationevaluation-metrics
    Auf GitHub ansehen↗297
  • guoyang9/unk-vqaAvatar von guoyang9

    guoyang9/UNK-VQA

    7Auf GitHub ansehen↗

    A VQA dataset that includes unanswerable questions TPAMI 2024.

    Auf GitHub ansehen↗7
  • esbatmop/mnbvcAvatar von esbatmop

    esbatmop/MNBVC

    4,123Auf GitHub ansehen↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    chinesechinese-languagechinese-nlp
    Auf GitHub ansehen↗4,123
Alle 27 Alternativen zu Seq Monkey Data anzeigen→

Häufig gestellte Fragen

Was sind die Hauptfunktionen von mobvoi/seq-monkey-data?

Die Hauptfunktionen von mobvoi/seq-monkey-data sind: Pre-training Datasets.

Welche Open-Source-Alternativen gibt es zu mobvoi/seq-monkey-data?

Open-Source-Alternativen zu mobvoi/seq-monkey-data sind unter anderem: plexpt/chatgpt-corpus — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of… fuxiaoliu/lrv-instruction — [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning. guoyang9/unk-vqa — A VQA dataset that includes unanswerable questions [TPAMI 2024]. hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models. icoz69/stablellava — Official repo for StableLLAVA. esbatmop/mnbvc — MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and…