This project is a comprehensive dataset and archive of classical Chinese poetry, prose, and Confucian classics. It serves as a digital humanities corpus, providing machine-readable access to hundreds of thousands of poems and detailed poet biographies, specifically spanning the Tang and Song dynasties. The collection is distinguished by its scholarly depth, incorporating textual variation annotations to track disputed characters across different source editions. It also includes tonal pattern mapping to describe the rhythmic and phonetic structures of the verse, alongside a popularity ranking
Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications. The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona
SentiBridge: A Knowledge Base for Entity-Sentiment Representation
漢語拆字字典
Las características principales de kfcd/chaizi son: Natural Language Processing, Specialized NLP Datasets, Corpus and Datasets, General Language Corpora, Lexical Analysis Tools.
Las alternativas de código abierto para kfcd/chaizi incluyen: chinese-poetry/chinese-poetry — This project is a comprehensive dataset and archive of classical Chinese poetry, prose, and Confucian classics. It… wainshine/company-names-corpus — 公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。. observerss/textfilter — 敏感词过滤的几种实现+某1w词敏感词库. pwxcoo/chinese-xinhua — Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese… rainarch/sentibridge — SentiBridge: A Knowledge Base for Entity-Sentiment Representation. crownpku/small-chinese-corpus.