awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
PlexPt avatar

PlexPt/chatgpt-corpus

0
View on GitHub↗
964 stars·146 forks·GPL-3.0·14 viewschat.aimakex.com↗

Chatgpt Corpus

This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing.

The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine learning frameworks and training environments.

The corpus facilitates research and development in natural language understanding by offering normalized text ready for subword tokenization. These materials are structured to support batch loading, enabling the preparation of diverse datasets for large-scale generative artificial intelligence training.

Features

  • Language Corpora - Provides a comprehensive dataset of Chinese text covering multiple domains to support research and development.
  • Training Datasets - Supplies large-scale text collections including dialogue and creative writing to improve the performance of large language models.
  • Pre-training Datasets - Provides a large-scale Chinese language text collection including dialogue and fiction for training large language models.
  • Large Language Model Training Resources - Organizes diverse Chinese text datasets to support the pre-training and fine-tuning of large language models.
  • Natural Language Processing Resources - Serves as a structured natural language processing resource to improve the performance and accuracy of generative AI models.
  • Prompt Engineering - Large-scale datasets for training and fine-tuning language models.
  • Instruction Datasets - Large-scale self-instruct dataset generated by ChatGPT.
  • Learning and Research - Chinese dialogue and training corpora for LLMs.
  • Conversational AI - Curates dialogue and customer service interaction datasets to improve the natural language understanding of conversational AI systems.
  • Model Fine-Tuning - Provides specialized fiction and narrative datasets to enhance the storytelling capabilities of generative AI models through fine-tuning.
  • Natural Language Processing Datasets - Offers structured Chinese text corpora to facilitate linguistic analysis and evaluation of machine learning models.
  • Batched Data Loading - Provides structured text data formatted for efficient batch loading into machine learning training pipelines.

Star history

Star history chart for plexpt/chatgpt-corpusStar history chart for plexpt/chatgpt-corpus

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does plexpt/chatgpt-corpus do?

This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing.

What are the main features of plexpt/chatgpt-corpus?

The main features of plexpt/chatgpt-corpus are: Language Corpora, Training Datasets, Pre-training Datasets, Large Language Model Training Resources, Natural Language Processing Resources, Prompt Engineering, Instruction Datasets, Learning and Research.

Which projects share features with plexpt/chatgpt-corpus?

Projects with overlapping indexed features include: d2l-ai/d2l-en — This project is an educational platform and research toolkit designed to teach deep learning through a combination of… brightmart/nlp_chinese_corpus — This is a large-scale collection of curated Chinese text corpora designed for training natural language processing… thunlp/ultrachat — UltraChat is a collection of large-scale conversational datasets and instruction-tuning data designed for training and… esbatmop/mnbvc — MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and… wdndev/llm_interview_note — This project is a comprehensive technical reference and educational resource focused on the lifecycle of large… zai-org/chatglm3 — ChatGLM3 is a comprehensive framework for deploying, fine-tuning, and serving large language models. It functions as a…

Projects sharing features with Chatgpt Corpus

These projects share indexed features with Chatgpt Corpus. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • d2l-ai/d2l-end2l-ai avatar

    d2l-ai/d2l-en

    29,001View on GitHub↗

    This project is an educational platform and research toolkit designed to teach deep learning through a combination of mathematical theory, visual diagrams, and executable code. It provides a comprehensive environment for building, training, and evaluating neural networks, grounding complex concepts in interactive computational notebooks that allow for hands-on experimentation. The framework distinguishes itself by interleaving theoretical foundations—including linear algebra, calculus, and probability—with practical implementations across multiple industry-standard libraries. It supports flex

    Pythonbookcomputer-visiondata-science
    View on GitHub↗29,001
  • brightmart/nlp_chinese_corpusbrightmart avatar

    brightmart/nlp_chinese_corpus

    9,903View on GitHub↗

    This is a large-scale collection of curated Chinese text corpora designed for training natural language processing models. The project provides a variety of datasets, including a deduplicated archive of millions of news articles with titles and keywords, high-quality categorized question-and-answer pairs, and parallel translation corpora. The collection includes millions of aligned Chinese and English sentence pairs used for cross-lingual model training and machine translation development. It also contains filtered question-and-answer data organized by label for the construction of knowledge-

    bertchinesechinese-corpus
    View on GitHub↗9,903
  • thunlp/ultrachatthunlp avatar

    thunlp/UltraChat

    2,786View on GitHub↗

    UltraChat is a collection of large-scale conversational datasets and instruction-tuning data designed for training and evaluating generative AI models. It provides structured JSON data consisting of complex, multi-round dialogue sequences intended to refine the performance of large language models in chat tasks. The project focuses on improving reasoning and response quality through a diverse set of interactions across multiple sectors. These datasets are used for supervised fine-tuning and instruction tuning workflows to improve how models follow complex directions and maintain context acros

    Pythonchatbotchatgptdeep-learning
    View on GitHub↗2,786
  • esbatmop/mnbvcesbatmop avatar

    esbatmop/MNBVC

    4,123View on GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    chinesechinese-languagechinese-nlp
    View on GitHub↗4,123
Compare all 30 related projects→