awesome-repositories.comCategoriesBlog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
proycon avatar

proycon/colibri-core

0
View on GitHub↗
130 stars·20 forks·C++·GPL-3.0·10 viewsproycon.github.io/colibri-core↗

Colibri Core

Colibri core is an NLP tool as well as a C++ and Python library for working with basic linguistic constructions such as n-grams and skipgrams (i.e patterns with one or more gaps, either of fixed or dynamic size) in a quick and memory-efficient way. At the core is the tool colibri-patternmodeller whi ch allows you to build, view, manipulate and query pattern models.

Features

  • Natural Language Processing - Efficient extraction and processing of linguistic n-grams.
  • C++ NLP Libraries - Efficient library for extracting n-grams and skipgrams.

Star history

Star history chart for proycon/colibri-coreStar history chart for proycon/colibri-core

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Colibri Core

These projects share indexed features with Colibri Core. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • facebookresearch/starspacefacebookresearch avatar

    facebookresearch/Starspace

    3,954View on GitHub↗

    Starspace is a vector embedding framework designed for training high-dimensional representations of text and images. It functions as a machine learning system for neural ranking, text classification, and knowledge graph embedding, mapping different object types into a shared numerical space to facilitate retrieval and prediction tasks. The system includes specialized tools for knowledge graph completion and link prediction by representing entities and their relationships within a multi-relational vector space. It further provides capabilities for semantic content recommendation and large-scal

    C++
    View on GitHub↗3,954
  • languagemachines/frogLanguageMachines avatar

    LanguageMachines/frog

    81View on GitHub↗

    Frog is an integration of memory-based natural language processing (NLP) modules developed for Dutch. All NLP modules are based on Timbl, the Tilburg memory-based learning software package.

    C++
    View on GitHub↗81
  • bllip/bllip-parserBLLIP avatar

    BLLIP/bllip-parser

    227View on GitHub↗

    BLLIP reranking parser (also known as Charniak-Johnson parser, Charniak parser, Brown reranking parser) See http://pypi.python.org/pypi/bllipparser/ for Python module.

    GAP
    View on GitHub↗227
  • languagemachines/libfoliaLanguageMachines avatar

    LanguageMachines/libfolia

    17View on GitHub↗

    FoLiA library for C++

    C++
    View on GitHub↗17
Compare all 30 related projects→

Frequently asked questions

What does proycon/colibri-core do?

Colibri core is an NLP tool as well as a C++ and Python library for working with basic linguistic constructions such as n-grams and skipgrams (i.e patterns with one or more gaps, either of fixed or dynamic size) in a quick and memory-efficient way. At the core is the tool `colibri-patternmodeller` whi ch allows you to build, view, manipulate and query pattern models.

What are the main features of proycon/colibri-core?

The main features of proycon/colibri-core are: Natural Language Processing, C++ NLP Libraries.

Which projects share features with proycon/colibri-core?

Projects with overlapping indexed features include: languagemachines/frog — Frog is an integration of memory-based natural language processing (NLP) modules developed for Dutch. All NLP modules… languagemachines/ucto — Unicode tokeniser. Ucto tokenizes text files: it separates words from punctuation, and splits sentences. It offers… bllip/bllip-parser — BLLIP reranking parser (also known as Charniak-Johnson parser, Charniak parser, Brown reranking parser) See… facebookresearch/starspace — Starspace is a vector embedding framework designed for training high-dimensional representations of text and images.… languagemachines/libfolia — FoLiA library for C++. meta-toolkit/meta — A Modern C++ Data Sciences Toolkit.