awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
CLUEbenchmark avatar

CLUEbenchmark/CLUECorpus2020

0
View on GitHub↗
1,012 stars·83 forks·MIT·4 viewsarxiv.org/abs/2003.01355↗

CLUECorpus2020

Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

Features

  • Datasets and Corpora - High-quality Chinese pre-training corpus for NLP tasks.
  • Pre-training Datasets - Cleaned 100GB Chinese corpus for pre-training and NLP tasks.

Star history

Star history chart for cluebenchmark/cluecorpus2020Star history chart for cluebenchmark/cluecorpus2020

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to CLUECorpus2020

Similar open-source projects, ranked by how many features they share with CLUECorpus2020.
  • esbatmop/mnbvcesbatmop avatar

    esbatmop/MNBVC

    4,123View on GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    chinesechinese-languagechinese-nlp
    View on GitHub↗4,123
  • plexpt/chatgpt-corpusPlexPt avatar

    PlexPt/chatgpt-corpus

    964View on GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    awesomecorpuscorpus-data
    View on GitHub↗964
  • chinawithfrank/chatbotcoursechinawithfrank avatar

    chinawithfrank/ChatBotCourse

    6,018View on GitHub↗

    This project is a development course and learning curriculum focused on building large language model chatbots. It provides a structured series of tutorials for creating conversational agents through the application of natural language processing and deep learning models. The materials include a technical walkthrough for implementing neural networks and word embeddings to handle automated question-answering tasks. It also provides a guide for constructing large-scale conversation corpora from external text sources to train and evaluate dialogue systems. The curriculum covers core text analys

    Python
    View on GitHub↗6,018
  • deepmind/rc-datadeepmind avatar

    deepmind/rc-data

    1,296View on GitHub↗

    Question answering dataset featured in "Teaching Machines to Read and Comprehend

    Python
    View on GitHub↗1,296
See all 30 alternatives to CLUECorpus2020→

Frequently asked questions

What does cluebenchmark/cluecorpus2020 do?

Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

What are the main features of cluebenchmark/cluecorpus2020?

The main features of cluebenchmark/cluecorpus2020 are: Datasets and Corpora, Pre-training Datasets.

What are some open-source alternatives to cluebenchmark/cluecorpus2020?

Open-source alternatives to cluebenchmark/cluecorpus2020 include: esbatmop/mnbvc — MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and… plexpt/chatgpt-corpus — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of… chinawithfrank/chatbotcourse — This project is a development course and learning curriculum focused on building large language model chatbots. It… deepmind/rc-data — Question answering dataset featured in "Teaching Machines to Read and Comprehend. goldsmith/wikipedia — Wikipedia. fuxiaoliu/lrv-instruction — [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.