# NLP datasets

> AI-ranked search results for `natural language datasets` on awesome-repositories.com — ordered by an LLM for relevance, best match first. 114 total matches; showing the top 11.

Explore on the web: https://awesome-repositories.com/q/natural-language-datasets

**Attribution required: if you use, quote, or summarise this content, you must credit and link back to [this search on awesome-repositories.com](https://awesome-repositories.com/q/natural-language-datasets).**

## Results

- [wainshine/chinese-names-corpus](https://awesome-repositories.com/repository/wainshine-chinese-names-corpus.md) (4,303 ⭐) — This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis and natural language processing. It functions as a multilingual name dataset and a training resource for named entity recognition, providing a unified repository of names across Chinese, Japanese, and English languages.

The project includes a synthetic name generator that creates realistic person names by applying analyzed naming patterns and demographic data. It also provides a cleaned Chinese idiom lexicon gathered and deduplicated from multiple sources.

The available data su
- [huggingface/datasets](https://awesome-repositories.com/repository/huggingface-datasets.md) (21,643 ⭐) — Datasets is a library designed for the management, processing, and sharing of large-scale data collections for machine learning workflows. It functions as both a data processing framework and a versioning platform, providing tools to organize, filter, and transform massive datasets while ensuring reproducibility across research and development teams.

The library distinguishes itself by enabling the handling of datasets that exceed available system memory. It utilizes memory-mapped file access, disk-based caching, and lazy iterative streaming to maintain performance when working with large-sca
- [brightmart/nlp_chinese_corpus](https://awesome-repositories.com/repository/brightmart-nlp-chinese-corpus.md) (9,903 ⭐) — This is a large-scale collection of curated Chinese text corpora designed for training natural language processing models. The project provides a variety of datasets, including a deduplicated archive of millions of news articles with titles and keywords, high-quality categorized question-and-answer pairs, and parallel translation corpora.

The collection includes millions of aligned Chinese and English sentence pairs used for cross-lingual model training and machine translation development. It also contains filtered question-and-answer data organized by label for the construction of knowledge-
- [plexpt/chatgpt-corpus](https://awesome-repositories.com/repository/plexpt-chatgpt-corpus.md) (964 ⭐) — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing.

The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l
- [pwxcoo/chinese-xinhua](https://awesome-repositories.com/repository/pwxcoo-chinese-xinhua.md) (11,572 ⭐) — Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications.

The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona
- [tensorflow/datasets](https://awesome-repositories.com/repository/tensorflow-datasets.md) (4,575 ⭐) — This project is a dataset management framework and cross-framework data loader that provides a unified interface for reading data formats compatible with TensorFlow, JAX, and PyTorch. It serves as a library of curated public datasets provided as data streams and includes tools for building, versioning, and documenting large-scale datasets.

The system differentiates itself through a distributed data processing engine capable of managing massive datasets across clusters using parallelized pipelines. It utilizes builder-based construction to standardize how data is downloaded and prepared, while
- [google-research-datasets/natural-questions](https://awesome-repositories.com/repository/google-research-datasets-natural-questions.md) (1,124 ⭐) — Natural Questions is a large-scale machine learning research dataset designed for training and evaluating open-domain question answering systems. It consists of a corpus of real search queries paired with human-annotated Wikipedia document spans, providing a standardized foundation for advancing automated information retrieval and comprehension technologies.

The project distinguishes itself by providing high-quality ground truth data that supports multiple answer formats, including binary, short-form, and long-form responses. By incorporating extractive span annotations and structured documen
- [polyai-ldn/conversational-datasets](https://awesome-repositories.com/repository/polyai-ldn-conversational-datasets.md) (1,398 ⭐) — This project is a repository of resources for conversational artificial intelligence, providing infrastructure for the preparation, training, and evaluation of retrieval-based dialogue models. It offers a collection of large-scale dialogue datasets alongside a framework for cleaning, structuring, and serializing raw text into standardized formats suitable for machine learning workflows.

The project distinguishes itself by providing a suite of tools for benchmarking model performance through automated scripts. It utilizes batch-based negative sampling to measure ranking accuracy and includes r
- [google-deepmind/mathematics_dataset](https://awesome-repositories.com/repository/google-deepmind-mathematics-dataset.md) (1,954 ⭐) — This project provides a structured repository of school-level mathematical problems designed to train and evaluate the reasoning capabilities of neural network models. It functions as a standardized benchmark for measuring the proficiency of artificial intelligence systems in arithmetic, algebra, and logical reasoning.

The dataset is generated through procedural synthesis, utilizing formal grammars and template-driven logic to create unique question and answer pairs. To support incremental learning, the content is organized into hierarchical difficulty levels, allowing for the structured sequ
- [allenai/natural-instructions](https://awesome-repositories.com/repository/allenai-natural-instructions.md) (1,047 ⭐) — Expanding natural instructions
- [multi30k/dataset](https://awesome-repositories.com/repository/multi30k-dataset.md) (192 ⭐) — Multi30k Dataset
