awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
mirfan899 avatar

mirfan899/Urdu

0
View on GitHub↗
72 stars·21 forks·MIT·9 viewsmirfan899.github.io/Urdu↗

Urdu

This a summary dataset. You can train abstractive summarization model using this dataset. It contains 3 files i.e. train, test and val. Data is in jsonl format.

Features

  • Datasets - Dataset for Urdu part-of-speech and named entity recognition.
  • Multilingual Datasets - Annotated corpus for POS tagging and named entity recognition.

Star history

Star history chart for mirfan899/urduStar history chart for mirfan899/urdu

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Urdu

Similar open-source projects, ranked by how many features they share with Urdu.
  • wainshine/chinese-names-corpuswainshine avatar

    wainshine/Chinese-Names-Corpus

    4,303View on GitHub↗

    This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis and natural language processing. It functions as a multilingual name dataset and a training resource for named entity recognition, providing a unified repository of names across Chinese, Japanese, and English languages. The project includes a synthetic name generator that creates realistic person names by applying analyzed naming patterns and demographic data. It also provides a cleaned Chinese idiom lexicon gathered and deduplicated from multiple sources. The available data su

    corpusdatasetdict
    View on GitHub↗4,303
  • allenai/ai2thorallenai avatar

    allenai/ai2thor

    1,668View on GitHub↗
    C#artificial-intelligencecomputer-visioninteraction
    View on GitHub↗1,668
  • andrews2017/kinnews-and-kirnews-corpusA

    Andrews2017/KINNEWS-and-KIRNEWS-Corpus

    0View on GitHub↗
    View on GitHub↗0
  • akngs/petitionsakngs avatar

    akngs/petitions

    40View on GitHub↗

    청와대 국민청원 사이트의 만료된 청원 데이터 모음.

    Python
    View on GitHub↗40
See all 30 alternatives to Urdu→

Frequently asked questions

What does mirfan899/urdu do?

This a summary dataset. You can train abstractive summarization model using this dataset. It contains 3 files i.e. train, test and val. Data is in jsonl format.

What are the main features of mirfan899/urdu?

The main features of mirfan899/urdu are: Datasets, Multilingual Datasets.

What are some open-source alternatives to mirfan899/urdu?

Open-source alternatives to mirfan899/urdu include: wainshine/chinese-names-corpus — This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis… allenai/ai2thor. andrews2017/kinnews-and-kirnews-corpus. anjieyang/vfhq-downloader — VFHQ-downloader is a Python-based utility designed for the easy downloading and processing of videos from the VFHQ… broadinstitute/lincs-profiling-complementarity. akngs/petitions — 청와대 국민청원 사이트의 만료된 청원 데이터 모음.