awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

37 repository-uri

Awesome GitHub RepositoriesLanguage Detection Tools

Utilities for identifying the language of provided text content.

Distinct from Natural Language Processing: Focuses on language identification for localization, distinct from general NLP libraries or text detection algorithms.

Explore 37 awesome GitHub repositories matching artificial intelligence & ml · Language Detection Tools. Refine with filters or upvote what's useful.

Awesome Language Detection Tools GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • public-apis/public-apisAvatar public-apis

    public-apis/public-apis

    441,986Vezi pe GitHub↗

    Acest proiect este un director curatoriat de comunitate cu endpoint-uri de servicii REST și GraphQL, conceput pentru a ajuta dezvoltatorii să descopere și să integreze surse de date terțe. Funcționează ca un registru centralizat unde serviciile externe sunt organizate pe domenii pentru a facilita prototiparea rapidă a software-ului și dezvoltarea aplicațiilor. Registrul se bazează pe un model de contribuție peer-reviewed, utilizând controlul distribuit al versiunilor pentru a gestiona actualizările și a asigura acuratețea endpoint-urilor listate. Pentru a menține o calitate ridicată a datelor, proiectul folosește validarea bazată pe schemă pentru toate trimiterile primite și compilează datele structurate într-un site web static, ușor de căutat, pentru o regăsire eficientă. Directorul acoperă un spectru larg de capabilități de integrare, inclusiv regăsirea datelor financiare, servicii de geolocalizare și diverse API-uri utilitare pentru sarcini precum detectarea limbajului, procesarea media și verificarea identității. Prin furnizarea unui index centralizat al acestor servicii, proiectul sprijină dezvoltatorii în identificarea furnizorilor de date fiabili pentru diverse cerințe funcționale.

    Identifies the language of text content to enable automated processing and localization.

    Pythonapiapisdataset
    Vezi pe GitHub↗441,986
  • hankcs/hanlpAvatar hankcs

    hankcs/HanLP

    36,413Vezi pe GitHub↗

    HanLP is a natural language processing library and deep learning framework specifically optimized for the Chinese language, while also functioning as a multilingual text processor. It serves as a toolkit for performing linguistic analysis, semantic understanding, and script conversion. The project distinguishes itself through a dedicated focus on Chinese linguistic structures, including a specialized script converter for transforming text between Simplified Chinese, Traditional Chinese, and Pinyin. It further supports domain-specific model training to improve the recognition of professional t

    Provides utilities for identifying the specific language of provided text content.

    Pythondependency-parserhanlpnamed-entity-recognition
    Vezi pe GitHub↗36,413
  • facebookresearch/fairseqAvatar facebookresearch

    facebookresearch/fairseq

    32,228Vezi pe GitHub↗

    Fairseq is a PyTorch toolkit for sequence-to-sequence modeling, specializing in neural machine translation, automatic speech recognition, and large-scale language model training. It provides a framework for processing and aligning diverse data sources, including text, audio, and video, to support tasks such as speech-to-text conversion and multimodal sequence learning. The project is distinguished by its distributed training capabilities, which utilize parameter sharding, mixed-precision training, and CPU offloading to handle models that exceed single-device memory. It also includes specializ

    Classifies the language spoken in an audio sample by extracting embeddings and calculating accuracy.

    Python
    Vezi pe GitHub↗32,228
  • isagalaev/highlight.jsAvatar isagalaev

    isagalaev/highlight.js

    24,937Vezi pe GitHub↗

    highlight.js is a JavaScript syntax highlighter and client-side code formatter that transforms plain text source code into highlighted HTML for web display. It provides syntax highlighting across a wide variety of programming languages. The library includes an automatic language detector that identifies the programming language of a code block to apply the correct highlighting rules without manual tagging. It is designed for web worker compatibility, allowing the highlighting process to run in background threads to prevent the browser interface from freezing during the processing of large vol

    Acts as a specialized library designed to identify the programming language of source code blocks.

    JavaScript
    Vezi pe GitHub↗24,937
  • qeeqbox/social-analyzerAvatar qeeqbox

    qeeqbox/social-analyzer

    21,134Vezi pe GitHub↗

    Social-analyzer is an open-source intelligence framework designed for the automated discovery, correlation, and verification of digital identities across online platforms. It functions as a comprehensive engine for gathering social media intelligence, utilizing distributed browser automation to extract metadata and profile information from hundreds of websites simultaneously. The platform distinguishes itself through its ability to perform cross-platform identity correlation using heuristic-based pattern matching and name permutation generation. It processes these findings through a confidenc

    Detects the primary language of discovered profile content to assist in digital footprint analysis.

    JavaScriptanalysisanalyzercli
    Vezi pe GitHub↗21,134
  • livekit/livekitAvatar livekit

    livekit/livekit

    19,358Vezi pe GitHub↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Identifies the spoken language of an audio stream in real-time, including support for mid-stream language switching.

    Gogolangmedia-serversfu
    Vezi pe GitHub↗19,358
  • modelscope/funasrAvatar modelscope

    modelscope/FunASR

    18,481Vezi pe GitHub↗

    FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken audio into written text across more than fifty languages. It provides a framework for speaker diarization, an OpenAI-compatible transcription API for local server hosting, and speech models compatible with the ONNX format. The project distinguishes itself by supporting high-performance inference on edge hardware via self-contained binaries and portable model exports. It incorporates specialized capabilities for natural speech generation with adjustable timbre and emotional expre

    Identifies the specific language being spoken from a wide range of supported multilingual options.

    Pythonasraudiochinese
    Vezi pe GitHub↗18,481
  • copytranslator/copytranslatorAvatar CopyTranslator

    CopyTranslator/CopyTranslator

    17,749Vezi pe GitHub↗

    CopyTranslator is a clipboard-based translation tool and multi-engine translation client that monitors the system clipboard to provide automatic language conversion. It functions as an assistant that integrates large language models, cloud translation APIs, and digital dictionaries to produce context-aware translations and side-by-side reading views. The application includes a specialized PDF text cleaner to remove formatting artifacts and line breaks from copied content. It also features an optical character recognition extractor to convert images or screen captures into editable text for im

    Integrates a local library to identify the source language of copied text without requiring network access.

    TypeScriptcopytranslatordictionarytranslate
    Vezi pe GitHub↗17,749
  • pot-app/pot-desktopAvatar pot-app

    pot-app/pot-desktop

    17,110Vezi pe GitHub↗

    This application is a cross-platform desktop utility designed for automated translation, optical character recognition, and speech synthesis. It functions as a modular client that integrates various local and remote language services, allowing users to process text through hotkeys, clipboard monitoring, or direct input. The software distinguishes itself through a plugin-based architecture and a built-in automation framework. By exposing a local network interface, it enables external applications and scripts to programmatically trigger its translation and recognition workflows. Users can furth

    Identifies the language of input text using local or remote processing services.

    JavaScriptlinuxmacosocr
    Vezi pe GitHub↗17,110
  • getgrav/gravAvatar getgrav

    getgrav/grav

    15,395Vezi pe GitHub↗

    Grav is a flat-file content management system that eliminates the need for a traditional database by storing site content and configuration in human-readable Markdown and YAML files. Built as a modular PHP web framework, it uses a hierarchical page routing system where the physical directory structure directly determines the site's URL paths. The platform is distinguished by its event-driven plugin architecture and a command-line interface that prioritizes system administration, deployment, and maintenance tasks. It utilizes a blueprint-driven system to generate administrative forms from stru

    Identifies input text language to assist in selecting correct translation targets.

    PHPcmscontentcontent-management
    Vezi pe GitHub↗15,395
  • libretranslate/libretranslateAvatar LibreTranslate

    LibreTranslate/LibreTranslate

    15,201Vezi pe GitHub↗

    LibreTranslate is an open-source, self-hosted machine translation engine that provides a private alternative to proprietary cloud-based translation services. It functions as a portable translation server, allowing users to process text and document translations locally or within their own infrastructure without relying on external providers. The platform distinguishes itself through its focus on privacy and flexible deployment, supporting anonymous network routing to bypass restrictive firewalls and protect user data. It is designed for integration into broader software ecosystems, offering a

    Automatically identifies the primary language of input text to streamline translation workflows.

    Pythonapimachinetranslate
    Vezi pe GitHub↗15,201
  • codelucas/newspaperAvatar codelucas

    codelucas/newspaper

    14,982Vezi pe GitHub↗

    Newspaper is a Python library designed for scraping, parsing, and analyzing web-based information. It functions as a framework for automated news aggregation and large-scale web content extraction, providing tools to download, clean, and structure text, metadata, and media from diverse online sources. The project distinguishes itself through a pipeline-oriented architecture that combines heuristic-based content extraction with natural language processing. It automatically identifies and isolates article bodies from web page boilerplate while simultaneously performing language detection, keywo

    Automatically identifies the language of web pages to ensure accurate parsing of international content.

    HTMLcrawlercrawlingnews
    Vezi pe GitHub↗14,982
  • languagetool-org/languagetoolAvatar languagetool-org

    languagetool-org/languagetool

    14,597Vezi pe GitHub↗

    LanguageTool is a multilingual grammar and style checking engine designed to detect spelling, grammar, and writing errors across multiple languages. It provides automated proofreading capabilities that can be deployed as a self-hosted server or executed as a standalone local desktop application. The project distinguishes itself through a flexible rule development framework, allowing linguistic patterns to be defined via XML or implemented as custom Java classes. It utilizes n-gram frequency modeling for confused word detection and supports neural word embeddings to improve disambiguation betw

    Includes utilities for identifying the language of provided text content to apply the correct grammar rules.

    Javagrammarnatural-languagenatural-language-processing
    Vezi pe GitHub↗14,597
  • github-linguist/linguistAvatar github-linguist

    github-linguist/linguist

    13,546Vezi pe GitHub↗

    Linguist is a programming language detection library designed to identify the languages used within source code files and software repositories. It functions as a repository metadata classifier, providing the automated analysis necessary to generate language statistics and insights for version control platforms. The tool employs a strategy-based detection pipeline that combines multiple identification methods to ensure accuracy. It utilizes heuristic-based pattern matching for file extensions and filenames, supplemented by regex-driven content analysis and Bayesian statistical classification

    Identifies the language of source code files by analyzing extensions, filenames, and content patterns.

    Rubylanguage-grammarslanguage-statisticslinguistic
    Vezi pe GitHub↗13,546
  • k2-fsa/sherpa-onnxAvatar k2-fsa

    k2-fsa/sherpa-onnx

    13,017Vezi pe GitHub↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    Detects the language being spoken in an audio recording using pre-trained models.

    C++aarch64androidarm32
    Vezi pe GitHub↗13,017
  • basedhardware/omiAvatar BasedHardware

    BasedHardware/omi

    12,869Vezi pe GitHub↗

    Omi is an open-source wearable AI platform that captures audio and screen data to provide real-time conversational assistance and memory. It integrates a wearable hardware development kit with a vector memory database and large language model capabilities to create a persistent digital record of user interactions. The platform is distinguished by its BLE audio streaming pipeline, which transmits raw audio from wearable hardware for real-time transcription and speaker identification. It utilizes a plugin-based agent tool framework that allows AI assistants to autonomously invoke custom functio

    Automatically identifies the spoken language from an audio stream without requiring a preset language code.

    Dartaiappbci
    Vezi pe GitHub↗12,869
  • microsoft/windows-universal-samplesAvatar microsoft

    microsoft/Windows-universal-samples

    9,696Vezi pe GitHub↗

    This repository is a comprehensive collection of reference implementations and sample libraries for the Universal Windows Platform. It provides practical examples of how to use Windows Runtime APIs to build cross-device applications, including detailed guidance on XAML-based declarative user interfaces and DirectX-integrated rendering. The project distinguishes itself by providing a wide array of hardware integration suites, covering low-level communication with USB, Serial, I2C, SPI, and GPIO peripherals. It includes specialized implementations for mixed reality holographic rendering, advanc

    Identifies the languages present within a text string and provides confidence-sorted results.

    JavaScript
    Vezi pe GitHub↗9,696
  • niedev/rtranslatorAvatar niedev

    niedev/RTranslator

    9,641Vezi pe GitHub↗

    RTranslator is a speech translation application and background speech processor designed for real-time voice and text translation. It functions as a dual-language audio translator that detects different spoken languages to provide immediate voice translations for face-to-face interactions. The system includes a multi-device translation setup that synchronizes multiple devices over a network to facilitate group conversations through shared speakers. It also provides the ability to maintain audio processing and speech translation tasks while the device is on standby. The project covers a range

    Detects two different spoken languages and provides immediate voice translations for face-to-face interactions.

    C++androidandroid-appbluetooth-le
    Vezi pe GitHub↗9,641
  • livekit/agentsAvatar livekit

    livekit/agents

    9,379Vezi pe GitHub↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Automatically identifies different spoken languages within a single audio stream.

    Pythonagentsaiopenai
    Vezi pe GitHub↗9,379
  • owo-network/deeplxAvatar OwO-Network

    OwO-Network/DeepLX

    8,571Vezi pe GitHub↗

    DeepLX is a self-hosted translation API server that provides free access to DeepL translation services without requiring a paid API token or account. It functions as a token-free proxy, enabling programmatic text translation between languages through a single HTTP endpoint. The server accepts JSON-based requests with source and target language parameters, supporting automatic source language detection and validation of language codes. It processes each translation request independently in a stateless manner, and offers optional bearer token authentication to restrict access to authorized clie

    Provides automatic source language detection for translation requests using standard language codes.

    Gobobplugindeepl
    Vezi pe GitHub↗8,571
Înapoi12Înainte
  1. Home
  2. Artificial Intelligence & ML
  3. Language Detection Tools

Explorează sub-etichetele

  • Language-Agnostic Speech RecognitionPhonetic recognition systems that operate independently of a specific spoken language. **Distinct from Spoken Language Detection:** Focuses on being language-independent for phonetic mapping rather than detecting which language is being spoken.
  • Source Code Language Detectors1 sub-tagLibraries designed to identify the programming language of source code files. **Distinct from Language Detection Tools:** Focuses on source code identification, distinct from general text language detection.
  • Spoken Language Detection1 sub-tagIdentification of the language being spoken in an audio recording. **Distinct from Language Detection Tools:** Focuses on audio-based spoken language identification rather than text-based language detection.
  • System Language Auto-DetectionDetects the user's operating system language automatically and applies internationalized text throughout the application. **Distinct from Language Detection Tools:** Distinct from Language Detection Tools: focuses on detecting the OS locale for UI localization, not identifying the language of arbitrary text.