How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis and natural language processing. It functions as a multilingual name dataset and a training resource for named entity recognition, providing a unified repository of names across Chinese, Japanese, and English languages. The project includes a synthetic name generator that creates realistic person names by applying analyzed naming patterns and demographic data. It also provides a cleaned Chinese idiom lexicon gathered and deduplicated from multiple sources. The available data su
This a summary dataset. You can train abstractive summarization model using this dataset. It contains 3 files i.e. train, test and val. Data is in jsonl format.
The main features of mirfan899/urdu are: Datasets, Multilingual Datasets.
Open-source alternatives to mirfan899/urdu include: wainshine/chinese-names-corpus — This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis… allenai/ai2thor. andrews2017/kinnews-and-kirnews-corpus. anjieyang/vfhq-downloader — VFHQ-downloader is a Python-based utility designed for the easy downloading and processing of videos from the VFHQ… broadinstitute/lincs-profiling-complementarity. akngs/petitions — 청와대 국민청원 사이트의 만료된 청원 데이터 모음.