awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to dice-group/nliwod

Projects sharing features with NLIWOD

9 open-source projects similar to dice-group/nliwod, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • google-research-datasets/natural-questionsgoogle-research-datasets avatar

    google-research-datasets/natural-questions

    1,124View on GitHub↗

    Natural Questions is a large-scale machine learning research dataset designed for training and evaluating open-domain question answering systems. It consists of a corpus of real search queries paired with human-annotated Wikipedia document spans, providing a standardized foundation for advancing automated information retrieval and comprehension technologies. The project distinguishes itself by providing high-quality ground truth data that supports multiple answer formats, including binary, short-form, and long-form responses. By incorporating extractive span annotations and structured documen

    Python
    View on GitHub↗1,124
  • brightmart/nlp_chinese_corpusbrightmart avatar

    brightmart/nlp_chinese_corpus

    9,903View on GitHub↗

    This is a large-scale collection of curated Chinese text corpora designed for training natural language processing models. The project provides a variety of datasets, including a deduplicated archive of millions of news articles with titles and keywords, high-quality categorized question-and-answer pairs, and parallel translation corpora. The collection includes millions of aligned Chinese and English sentence pairs used for cross-lingual model training and machine translation development. It also contains filtered question-and-answer data organized by label for the construction of knowledge-

    bertchinesechinese-corpus
    View on GitHub↗9,903
  • deepmind/rc-datadeepmind avatar

    deepmind/rc-data

    1,296View on GitHub↗

    Question answering dataset featured in "Teaching Machines to Read and Comprehend

    Python
    View on GitHub↗1,296

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
facebookresearch/eli5facebookresearch avatar

facebookresearch/ELI5

324View on GitHub↗

Read the Paper: https://arxiv.org/abs/1907.09190

Python
View on GitHub↗324
  • deepmind/narrativeqadeepmind avatar

    deepmind/narrativeqa

    514View on GitHub↗

    This repository contains the NarrativeQA dataset. It includes the list of documents with Wikipedia summaries, links to full stories, and questions and answers.

    Shell
    View on GitHub↗514
  • karthikncode/nlp-datasetskarthikncode avatar

    karthikncode/nlp-datasets

    918View on GitHub↗

    This is a list of datasets/corpora for NLP tasks, in reverse chronological order. Suggestions and pull requests are welcome. The goal is to make this a collaborative effort to maintain an updated list of quality datasets.

    View on GitHub↗918
  • maluuba/newsqaMaluuba avatar

    Maluuba/newsqa

    257View on GitHub↗

    Tools for using Maluuba's news questions and answer data. The code in the repo is used to compile the dataset since it cannot be made directly available due to legal reasons.

    Python
    View on GitHub↗257
  • websail-nu/codahWebsail-NU avatar

    Websail-NU/CODAH

    22View on GitHub↗

    The COmmonsense Dataset Adversarially-authored by Humans (CODAH) is an evaluation set for commonsense question-answering in the sentence completion style of SWAG. As opposed to other automatically generated NLI datasets, CODAH is adversarially constructed by humans who can view feedback from a…

    Python
    View on GitHub↗22
  • ysu1989/graphquestionsysu1989 avatar

    ysu1989/GraphQuestions

    94View on GitHub↗

    GraphQuestions is a characteristic-rich dataset for factoid question answering described in the paper "On Generating Characteristic-rich Question Sets for QA Evaluation" - EMNLP'16.

    ReScript
    View on GitHub↗94