awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to aurelian/ruby-stemmer

Open-source alternatives to Ruby Stemmer

30 open-source projects similar to aurelian/ruby-stemmer, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Ruby Stemmer alternative.

  • ealdent/uea-stemmerAvatar de ealdent

    ealdent/uea-stemmer

    54Ver en GitHub↗

    Ruby port of UEALite Stemmer - a conservative stemmer for search and indexing

    Ruby
    Ver en GitHub↗54
  • 7compass/sentimentalAvatar de 7compass

    7compass/sentimental

    465Ver en GitHub↗

    Simple sentiment analysis with Ruby

    Ruby
    Ver en GitHub↗465
  • alibaba-nlp/daat-cwsAvatar de Alibaba-NLP

    Alibaba-NLP/DAAT-CWS

    23Ver en GitHub↗

    DAAT-CWS

    Python
    Ver en GitHub↗23
  • abadojack/whatlanggoAvatar de abadojack

    abadojack/whatlanggo

    688Ver en GitHub↗

    Natural language detection library for Go

    Go
    Ver en GitHub↗688
  • abitdodgy/gibranAvatar de abitdodgy

    abitdodgy/gibran

    65Ver en GitHub↗

    Gibran is an Elixir natural language processor, and a port of WordsCounted.

    Elixir
    Ver en GitHub↗65
  • abosamoor/polyglotAvatar de aboSamoor

    aboSamoor/polyglot

    2,367Ver en GitHub↗

    Multilingual text (NLP) processing toolkit

    Python
    Ver en GitHub↗2,367

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Find more with AI search
  • aethercortex/llama-xAvatar de AetherCortex

    AetherCortex/Llama-X

    1,605Ver en GitHub↗

    Open Academic Research on Improving LLaMA to SOTA LLM

    Python
    Ver en GitHub↗1,605
  • agonopol/go-stemAvatar de agonopol

    agonopol/go-stem

    81Ver en GitHub↗

    Word Stemming in Go

    Go
    Ver en GitHub↗81
  • ahmedbesbes/character-based-cnnA

    ahmedbesbes/character-based-cnn

    0Ver en GitHub↗
    Ver en GitHub↗0
  • ai-shifu/chatallAvatar de ai-shifu

    ai-shifu/ChatALL

    16,283Ver en GitHub↗

    ChatALL is a desktop application that functions as a multi-model chat client and aggregator for artificial intelligence services. It enables users to send a single prompt to multiple AI models simultaneously, allowing for the side-by-side comparison of generated responses within a unified interface. The application distinguishes itself through a local-first approach to data management, ensuring that all conversation logs and user configurations are stored directly on the user's device. This architecture supports privacy and offline access while providing a centralized system for managing and

    JavaScriptbingchatchatbotchatgpt
    Ver en GitHub↗16,283
  • aigc-audio/audiogptAvatar de AIGC-Audio

    AIGC-Audio/AudioGPT

    10,174Ver en GitHub↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Pythonaudiogptmusic
    Ver en GitHub↗10,174
  • alexrozanski/llamachatAvatar de alexrozanski

    alexrozanski/LlamaChat

    1,510Ver en GitHub↗

    Chat with your favourite LLaMA models in a native macOS app

    Swiftaillamallamacpp
    Ver en GitHub↗1,510
  • alexsergivan/transliteratorA

    alexsergivan/transliterator

    0Ver en GitHub↗
    Ver en GitHub↗0
  • alibaba-edu/simple-effective-text-matching-pytorchA

    alibaba-edu/simple-effective-text-matching-pytorch

    0Ver en GitHub↗
    Ver en GitHub↗0
  • a2800276/porterAvatar de a2800276

    a2800276/porter

    13Ver en GitHub↗

    porter stemmer

    Go
    Ver en GitHub↗13
  • alinapetukhova/textclAvatar de alinapetukhova

    alinapetukhova/textcl

    12Ver en GitHub↗

    Text preprocessing package for use in NLP tasks https://pypi.org/project/textcl/

    Python
    Ver en GitHub↗12
  • alisawuffles/ambientAvatar de alisawuffles

    alisawuffles/ambient

    66Ver en GitHub↗

    Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023)

    Jupyter Notebook
    Ver en GitHub↗66
  • all-hands-ai/openhandsAvatar de All-Hands-AI

    All-Hands-AI/OpenHands

    77,468Ver en GitHub↗

    OpenHands is an autonomous AI software engineer and coding assistant designed to execute software engineering tasks by interacting directly with codebases and development environments. It functions as a platform for running AI agents that can write code and manage files to automate complex development workflows. The system distinguishes itself through a container-based execution environment that isolates agent actions within a sandboxed Linux environment. It employs an autonomous agent loop of observation, planning, and action, supported by a standardized communication protocol that allows it

    Python
    Ver en GitHub↗77,468
  • allenai/allennlpAvatar de allenai

    allenai/allennlp

    11,889Ver en GitHub↗

    AllenNLP is a PyTorch-based research library and deep learning language toolkit designed for developing and training neural network architectures for linguistic tasks. It provides a distributed training system that coordinates data and gradients across multiple GPUs and a framework for integrating pretrained transformer architectures. The system distinguishes itself with a dedicated algorithmic bias mitigation tool used to identify and reduce bias in linguistic model predictions. It also includes model influence analysis to interpret predictions by calculating the influence of specific traini

    Python
    Ver en GitHub↗11,889
  • allenai/mmc4Avatar de allenai

    allenai/mmc4

    953Ver en GitHub↗

    MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.

    Python
    Ver en GitHub↗953
  • allenai/scispacyAvatar de allenai

    allenai/SciSpaCy

    1,968Ver en GitHub↗

    This repository contains custom pipes and models related to using spaCy for scientific documents.

    Python
    Ver en GitHub↗1,968
  • alvations/annotate-questionnaireAvatar de alvations

    alvations/annotate-questionnaire

    59Ver en GitHub↗

    Summary of Responses to Questionnaire on Annotation Platform https://forms.gle/iZk8kehkjAWmB8xe9

    Ver en GitHub↗59
  • anujvyas/natural-language-processing-projectsAvatar de anujvyas

    anujvyas/Natural-Language-Processing-Projects

    254Ver en GitHub↗

    This repository consists of all my NLP Projects

    Jupyter Notebook
    Ver en GitHub↗254
  • arc53/docsgptAvatar de arc53

    arc53/DocsGPT

    17,939Ver en GitHub↗

    DocsGPT is a retrieval-augmented generation platform and private knowledge base used to build AI agents that perform grounded search and analysis. It functions as a multi-model AI orchestrator and enterprise agent builder, allowing for the integration of various local and cloud language models to customize reasoning and text generation. The project provides a visual environment for developing automated assistants using conditional logic and third-party API connectivity. It enables the creation of private AI agents capable of performing enterprise search and detailed document analysis using pr

    Pythonagent-builderagentsai
    Ver en GitHub↗17,939
  • argilla-io/argillaAvatar de argilla-io

    argilla-io/argilla

    5,015Ver en GitHub↗

    Argilla is a collaborative AI feedback tool and data curation management system. It serves as a human-in-the-loop dataset platform designed to coordinate workforce annotators and domain experts in labeling, rating, and refining data samples for machine learning projects. The platform focuses on large language model dataset curation and reinforcement learning from human feedback workflows. It provides a shared workspace for integrating human expertise into AI development to validate model outputs and correct data errors. The system manages the end-to-end machine learning data pipeline, includ

    Python
    Ver en GitHub↗5,015
  • arongdari/python-topic-modelAvatar de arongdari

    arongdari/python-topic-model

    374Ver en GitHub↗

    Implementation of various topic models

    Jupyter Notebook
    Ver en GitHub↗374
  • arongdari/topic-model-lecture-noteAvatar de arongdari

    arongdari/topic-model-lecture-note

    22Ver en GitHub↗

    lecture notes for probabilistic topic models using ipython notebook

    Ver en GitHub↗22
  • artidoro/qloraAvatar de artidoro

    artidoro/qlora

    10,929Ver en GitHub↗

    This project is a quantized fine-tuning framework for large language models. It implements a low-rank adaptation library and a four-bit quantizer to reduce the GPU memory requirements needed to train large models. The framework utilizes four-bit quantization and low-rank adapters to enable model training on consumer-grade hardware. It further reduces the memory footprint through double quantization and a paged optimizer that offloads states to system RAM. The system supports distributed training across multiple GPUs to handle larger parameter scales and includes utilities for custom dataset

    Jupyter Notebook
    Ver en GitHub↗10,929
  • artificiai/multilingual-latent-dirichlet-allocation-ldaAvatar de ArtificiAI

    ArtificiAI/Multilingual-Latent-Dirichlet-Allocation-LDA

    83Ver en GitHub↗

    A Multilingual Latent Dirichlet Allocation (LDA) Pipeline with Stop Words Removal, n-gram features, and Inverse Stemming, in Python.

    Pythonclusteringenglishfrench
    Ver en GitHub↗83
  • 2shou/textgroceryAvatar de 2shou

    2shou/TextGrocery

    683Ver en GitHub↗

    A simple short-text classification tool based on LibLinear

    C++
    Ver en GitHub↗683