awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
deepmind avatar

deepmind/rc-dataArchived

0
View on GitHub↗
1,296 stars·239 forks·Python·Apache-2.0·10 views

Rc Data

Question answering dataset featured in "Teaching Machines to Read and Comprehend

Features

  • Data Science Resources - Datasets for training and evaluating machine reading comprehension models.
  • Datasets and Corpora - Reading comprehension and question answering datasets.
  • Question Answering Datasets - Datasets for machine reading comprehension and reasoning tasks.
  • Text and Language Datasets - Large-scale textual question-answering corpus derived from news documents.

Star history

Star history chart for deepmind/rc-dataStar history chart for deepmind/rc-data

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does deepmind/rc-data do?

Question answering dataset featured in "Teaching Machines to Read and Comprehend

What are the main features of deepmind/rc-data?

The main features of deepmind/rc-data are: Data Science Resources, Datasets and Corpora, Question Answering Datasets, Text and Language Datasets.

What are some open-source alternatives to deepmind/rc-data?

Open-source alternatives to deepmind/rc-data include: karthikncode/nlp-datasets — This is a list of datasets/corpora for NLP tasks, in reverse chronological order. Suggestions and pull requests are… google-research-datasets/natural-questions — Natural Questions is a large-scale machine learning research dataset designed for training and evaluating open-domain… chinawithfrank/chatbotcourse — This project is a development course and learning curriculum focused on building large language model chatbots. It… brightmart/nlp_chinese_corpus — This is a large-scale collection of curated Chinese text corpora designed for training natural language processing… apachecn/recommendersystems — 介绍推荐系统基本知识,相关算法以及实现。. caesar0301/awesome-public-datasets — A topic-centric list of HQ open datasets.

Open-source alternatives to Rc Data

Similar open-source projects, ranked by how many features they share with Rc Data.
  • karthikncode/nlp-datasetskarthikncode avatar

    karthikncode/nlp-datasets

    918View on GitHub↗

    This is a list of datasets/corpora for NLP tasks, in reverse chronological order. Suggestions and pull requests are welcome. The goal is to make this a collaborative effort to maintain an updated list of quality datasets.

    View on GitHub↗918
  • google-research-datasets/natural-questionsgoogle-research-datasets avatar

    google-research-datasets/natural-questions

    1,124View on GitHub↗

    Natural Questions is a large-scale machine learning research dataset designed for training and evaluating open-domain question answering systems. It consists of a corpus of real search queries paired with human-annotated Wikipedia document spans, providing a standardized foundation for advancing automated information retrieval and comprehension technologies. The project distinguishes itself by providing high-quality ground truth data that supports multiple answer formats, including binary, short-form, and long-form responses. By incorporating extractive span annotations and structured documen

    Python
    View on GitHub↗1,124
  • brightmart/nlp_chinese_corpusbrightmart avatar

    brightmart/nlp_chinese_corpus

    9,903View on GitHub↗

    This is a large-scale collection of curated Chinese text corpora designed for training natural language processing models. The project provides a variety of datasets, including a deduplicated archive of millions of news articles with titles and keywords, high-quality categorized question-and-answer pairs, and parallel translation corpora. The collection includes millions of aligned Chinese and English sentence pairs used for cross-lingual model training and machine translation development. It also contains filtered question-and-answer data organized by label for the construction of knowledge-

    bertchinesechinese-corpus
    View on GitHub↗9,903
  • chinawithfrank/chatbotcoursechinawithfrank avatar

    chinawithfrank/ChatBotCourse

    6,018View on GitHub↗

    This project is a development course and learning curriculum focused on building large language model chatbots. It provides a structured series of tutorials for creating conversational agents through the application of natural language processing and deep learning models. The materials include a technical walkthrough for implementing neural networks and word embeddings to handle automated question-answering tasks. It also provides a guide for constructing large-scale conversation corpora from external text sources to train and evaluate dialogue systems. The curriculum covers core text analys

    Python
    View on GitHub↗6,018
  • See all 21 alternatives to Rc Data→