4 Repos
Libraries designed to identify the programming language of source code files.
Distinct from Language Detection Tools: Focuses on source code identification, distinct from general text language detection.
Explore 4 awesome GitHub repositories matching artificial intelligence & ml · Source Code Language Detectors. Refine with filters or upvote what's useful.
highlight.js is a JavaScript syntax highlighter and client-side code formatter that transforms plain text source code into highlighted HTML for web display. It provides syntax highlighting across a wide variety of programming languages. The library includes an automatic language detector that identifies the programming language of a code block to apply the correct highlighting rules without manual tagging. It is designed for web worker compatibility, allowing the highlighting process to run in background threads to prevent the browser interface from freezing during the processing of large vol
Acts as a specialized library designed to identify the programming language of source code blocks.
Linguist is a programming language detection library designed to identify the languages used within source code files and software repositories. It functions as a repository metadata classifier, providing the automated analysis necessary to generate language statistics and insights for version control platforms. The tool employs a strategy-based detection pipeline that combines multiple identification methods to ensure accuracy. It utilizes heuristic-based pattern matching for file extensions and filenames, supplemented by regex-driven content analysis and Bayesian statistical classification
Identifies the language of source code files by analyzing extensions, filenames, and content patterns.
DeepLX is a self-hosted translation API server that provides free access to DeepL translation services without requiring a paid API token or account. It functions as a token-free proxy, enabling programmatic text translation between languages through a single HTTP endpoint. The server accepts JSON-based requests with source and target language parameters, supporting automatic source language detection and validation of language codes. It processes each translation request independently in a stateless manner, and offers optional bearer token authentication to restrict access to authorized clie
Provides automatic source language detection for translation requests using standard language codes.
Franc ist eine Bibliothek zur Erkennung natürlicher Sprachen und ein Befehlszeilen-Identifikator, der verwendet wird, um die geschriebene Sprache einer Textprobe zu bestimmen. Es fungiert als statistischer Sprachprofiler, der mehrsprachigen Text durch die Analyse von Zeichenverteilungen identifiziert und klassifiziert. Das Tool verwendet ein Trigram-basiertes statistisches Analysesystem, das die Häufigkeit von Drei-Zeichen-Sequenzen in einer Eingabeprobe mit Referenzprofilen vergleicht. Es ordnet potenzielle Sprachübereinstimmungen durch die Berechnung der statistischen Distanz zwischen der Eingabe und diesen Profilen und ermöglicht so die Rückgabe einer Rangliste wahrscheinlicher Sprachen. Das Projekt bietet eine Befehlszeilenschnittstelle für Textanalyse und unterstützt automatisierte Inhaltsklassifizierung. Es nutzt modulare Sprachdatensätze und vorberechnete N-Gramm-Tabellen, um die Erkennungslogik von sprachspezifischen Profildaten zu trennen.
Detects the natural language of text input by analyzing its statistical profile.