How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications. The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona
SentiBridge: A Knowledge Base for Entity-Sentiment Representation
公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。
敏感词过滤的几种实现+某1w词敏感词库
The main features of observerss/textfilter are: Natural Language Processing, Corpus and Datasets, Lexical Analysis Tools.
Projects with overlapping indexed features include: rainarch/sentibridge — SentiBridge: A Knowledge Base for Entity-Sentiment Representation. wainshine/company-names-corpus — 公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。. kfcd/chaizi — 漢語拆字字典. pwxcoo/chinese-xinhua — Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese… chinese-poetry/chinese-poetry — This project is a comprehensive dataset and archive of classical Chinese poetry, prose, and Confucian classics. It… embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It…