5 dépôts
Tools for executing optical character recognition tasks via terminal commands.
Explore 5 awesome GitHub repositories matching artificial intelligence & ml · OCR Command Line Interfaces. Refine with filters or upvote what's useful.
Tesseract is a neural network-based optical character recognition engine designed to convert scanned images and digital documents into machine-readable, searchable text. It functions as both a command-line utility for automating large-scale digitization workflows and a cross-platform library that can be embedded into desktop, mobile, or server-side applications. By utilizing long short-term memory networks, the engine provides robust text extraction across more than one hundred languages and dozens of scripts. The project distinguishes itself through a sophisticated document layout analysis f
Executes character recognition tasks directly from the terminal by specifying input images, language models, and output requirements.
chineseocr_lite is a lightweight Chinese optical character recognition engine designed to detect text regions, analyze orientation, and convert Chinese characters from images into digital text. It supports both horizontal and vertical reading layouts and can be deployed as a web service for image uploads and result visualization. The system utilizes a multi-backend inference framework that supports ncnn, mnn, and tnn, allowing it to run across diverse hardware and platforms. It is specifically engineered for lightweight deployment on mobile and desktop environments through the use of small mo
Provides a command line interface for performing OCR tasks and exporting structured results.
Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo
Performs OCR extraction on files directly from the terminal with configurable backends.
RapidOCR is an offline deep-learning OCR engine that detects and recognizes text in images using ONNX Runtime, operating entirely without an internet connection. It provides a unified inference pipeline that runs across multiple platforms including Windows, Linux, macOS, Android, and Raspberry Pi, with programming language bindings for Python, C++, Java, and C#. The engine separates text detection and recognition into independent modules that can be swapped or fine-tuned individually, and abstracts the inference backend behind a unified interface allowing seamless switching between ONNX Runti
Provides a command-line tool for extracting text from images and URLs with bounding boxes and confidence scores.
Ce projet est un moteur de reconnaissance optique de caractères (OCR) basé sur le terminal qui utilise des modèles de réseaux de neurones pour extraire du texte et des données de mise en page spatiale à partir d'images. Il fonctionne à la fois comme un utilitaire en ligne de commande pour le traitement automatisé de texte et comme une bibliothèque pour intégrer la reconnaissance basée sur l'apprentissage automatique dans des flux de travail plus larges. Le moteur se distingue par un pipeline de traitement modulaire qui prend en charge le chargement de modèles personnalisés et l'initialisation des poids mappés en mémoire pour une exécution efficace. Il préserve la structure du document en suivant les coordonnées géométriques précises pour chaque élément textuel détecté, et permet l'affinement de la sortie via des règles de validation au niveau des caractères. Le système inclut des outils complets pour l'ingestion d'images, y compris la capture directe depuis les presse-papiers système et le contenu du navigateur. Il fournit des capacités de diagnostic en générant des superpositions visuelles et des artefacts de traitement intermédiaires pour vérifier la précision de la reconnaissance et dépanner les performances du pipeline. Le logiciel est distribué sous forme de binaire statique pour garantir la portabilité entre les environnements sans nécessiter de dépendances externes.
Provides a terminal-based utility for processing images and clipboard data into structured text with spatial coordinates.