How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
A tokenizer based on the dictionary and Bigram language models for Go. (Now only support chinese segmentation)
The main features of xujiajun/gotokenizer are: Natural Language Processing, Text Tokenization.
Projects with overlapping indexed features include: dchest/stemmer — Stemmer packages for Go programming language. Includes English, German and Dutch stemmers. neurosnap/sentences — A multilingual command line sentence tokenizer in Golang. awsong/mmsego — Chinese word splitting algorithm MMSEG in GO. blevesearch/segment — A Go library for performing Unicode Text Segmentation as described in Unicode Standard Annex #29. go-ego/gse — Go efficient multilingual NLP and text segmentation; support English, Chinese, Japanese and others. osamingo/shamoji — The shamoji (杓文字) is a word filtering package.
A Go library for performing Unicode Text Segmentation as described in Unicode Standard Annex #29
Stemmer packages for Go programming language. Includes English, German and Dutch stemmers.
Go efficient multilingual NLP and text segmentation; support English, Chinese, Japanese and others.