1 dépôt
Vector representations tailored to the linguistic characteristics of specific training data sources.
Distinct from Facial Vector Representations: Focuses on corpus-derived text embeddings rather than biometric facial vectors
Explore 1 awesome GitHub repository matching data & databases · Corpus-Specific Vectorizations. Refine with filters or upvote what's useful.
This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu
Generates distinct vector spaces based on the specific linguistic characteristics of different training data sources.