awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目关于排名机制媒体报道MCP 服务器
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to belval/textrecognitiondatagenerator

Open-source alternatives to TextRecognitionDataGenerator

30 open-source projects similar to belval/textrecognitiondatagenerator, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best TextRecognitionDataGenerator alternative.

  • saurabhdaware/text-to-handwritingsaurabhdaware 的头像

    saurabhdaware/text-to-handwriting

    5,031在 GitHub 上查看↗

    This project is a text-to-handwriting image generator that transforms digital text into images simulating human writing. It functions as a suite of tools for handwriting image generation, physical paper simulation, and PDF document exportation. The system includes a custom font rendering engine that allows users to upload personal handwriting font files to replicate specific individual writing styles. A digital paper simulator provides tools to adjust margins and background images to mimic the appearance of physical paper. The tool covers the conversion of typed text into simulated handwritt

    HTMLassignmentsprojectstext-to-handwriting
    在 GitHub 上查看↗5,031
  • xpixelgroup/basicsrXPixelGroup 的头像

    XPixelGroup/BasicSR

    8,297在 GitHub 上查看↗

    BasicSR is a PyTorch-based image restoration toolbox and framework designed for training and deploying deep learning models to upscale, denoise, and deblur images and videos. It serves as a comprehensive system for image super-resolution and video quality restoration, providing the necessary infrastructure to recover fine visual details and increase pixel density. The project distinguishes itself through specialized toolkits for facial image enhancement and high-fidelity face synthesis, as well as a dedicated video quality restoration suite that utilizes deformable convolutions and generative

    Pythonbasicsrbasicvsrdfdnet
    在 GitHub 上查看↗8,297
  • bradlarson/gpuimage2BradLarson 的头像

    BradLarson/GPUImage2

    4,941在 GitHub 上查看↗

    GPUImage2 is a Swift framework for applying real-time filters and effects to images and video using the GPU. It provides a real-time video filter library, an image geometry manipulation engine, and an OpenGL shading pipeline for processing visual data on graphics hardware. The framework enables the construction of visual effect pipelines by chaining image sources to consumers in sequential flows. It supports the development of custom fragment and vertex shaders for bespoke image processing and offers the ability to bundle these operations into reusable units via graph-based grouping. Capabil

    Swift
    在 GitHub 上查看↗4,941

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • princeton-vl/infinigenprinceton-vl 的头像

    princeton-vl/infinigen

    7,022在 GitHub 上查看↗

    Infinigen is a procedural 3D scene generation framework that creates photorealistic indoor and outdoor environments for computer vision training data. It combines constraint-based object placement, GPU geometry shaders, and ground-truth rendering passes to produce scenes with depth, normals, and segmentation masks alongside final images. The framework distinguishes itself through modular asset composition, a node-graph material system, and physics simulation integration that embeds rigid-body and fluid dynamics directly into the generation pipeline. Procedural rule-based scene composition and

    Python
    在 GitHub 上查看↗7,022
  • paddlepaddle/paddleocrPaddlePaddle 的头像

    PaddlePaddle/PaddleOCR

    82,412在 GitHub 上查看↗

    PaddleOCR is a comprehensive optical character recognition framework designed for detecting and transcribing text from images and documents into structured, machine-readable formats. It provides a modular computer vision pipeline that decouples image preprocessing, text detection, and character recognition into independent, configurable stages. This architecture supports automated document digitization and multilingual text recognition, capable of identifying text in over one hundred languages across diverse environments ranging from scanned documents to industrial scenes. The framework disti

    Pythonai4sciencechineseocrdocument-parsing
    在 GitHub 上查看↗82,412
  • zhengpeng7/birefnetZhengPeng7 的头像

    ZhengPeng7/BiRefNet

    3,173在 GitHub 上查看↗

    BiRefNet is a PyTorch image segmentation framework designed for high-precision binary mask generation. It functions as a bilateral image segmentation model used to isolate foreground objects from complex backgrounds, as well as a specialized tool for camouflaged object detection and industrial defect detection. The project is designed for export to the ONNX format, which facilitates cross-platform deployment and inference. It supports custom model fine-tuning on user-provided image and mask datasets to adapt the model for specialized professional use cases. The system covers high-resolution

    Pythonbackground-removalbirefnetcamouflaged-object-detection
    在 GitHub 上查看↗3,173
  • facebookresearch/maskrcnn-benchmarkfacebookresearch 的头像

    facebookresearch/maskrcnn-benchmark

    9,370在 GitHub 上查看↗

    This project is a modular PyTorch framework for training and evaluating object detection and instance segmentation models. It serves as a computer vision research tool and a deep learning inference engine designed to identify object locations, classes, and pixel-level masks within images. The framework implements a two-stage inference pipeline that utilizes region proposal networks and a symmetric mask-head architecture. It provides specialized capabilities for instance segmentation, object bounding box detection, and human pose estimation via anatomical keypoint detection. The system includ

    Python
    在 GitHub 上查看↗9,370
  • wongkinyiu/yolov9WongKinYiu 的头像

    WongKinYiu/yolov9

    9,534在 GitHub 上查看↗

    YOLOv9 is a real-time computer vision framework and deep learning model designed for image classification, object detection, and instance segmentation. It functions as both a vision model and a trainer, allowing for the optimization of neural network weights on custom datasets using single or multiple GPUs. The framework utilizes programmable gradient information to perform high-speed identification and location of multiple objects within images and video streams. It extends beyond bounding box detection to provide instance segmentation and panoptic segmentation, which labels every pixel in a

    Pythonyolov9
    在 GitHub 上查看↗9,534
  • vercel/geist-fontvercel 的头像

    vercel/geist-font

    3,249在 GitHub 上查看↗

    Geist is an open-source font family and typography collection designed for high legibility in technical interfaces. It consists of a series of web-optimized typefaces, including geometric sans-serif, monospaced, and pixel styles. The collection functions as a variable font library, utilizing coordinate interpolation to allow precise control over weight and style within a single font file. These fonts are built as OpenType typefaces, incorporating standardized layout tables to define advanced typographic behaviors such as kerning and ligatures. The project provides specific implementations fo

    HTMLfontvariable-fonts
    在 GitHub 上查看↗3,249
  • jbarlow83/ocrmypdfjbarlow83 的头像

    jbarlow83/OCRmyPDF

    33,901在 GitHub 上查看↗

    OCRmyPDF is a tool for converting image-based PDF files into machine-readable documents by adding a searchable text layer via optical character recognition. It functions as a multi-language processor capable of detecting and extracting text in over 100 different languages using linguistic data packs. The software includes a PDF image optimizer to remove image artifacts and correct page skew to improve recognition accuracy. It also provides a converter to transform scanned documents into the PDF/A standard for long-term digital archiving. The system manages PDF optimization by compressing emb

    Python
    在 GitHub 上查看↗33,901
  • larsenwork/monoidlarsenwork 的头像

    larsenwork/monoid

    7,962在 GitHub 上查看↗

    Monoid is a monospaced programming font designed for high legibility in code editors and terminals. It is an OpenType feature font optimized for reading and writing source code, providing crisp rendering at small sizes. The typeface utilizes OpenType font-feature settings to provide stylistic alternates and customizable glyph appearances. It specifically implements programming ligatures that combine common coding symbols and operators into single glyphs to improve readability. The project covers monospaced text rendering and stylistic glyph customization, allowing for the swapping of standar

    Python
    在 GitHub 上查看↗7,962
  • matterport/mask_rcnnmatterport 的头像

    matterport/Mask_RCNN

    25,564在 GitHub 上查看↗

    This project is a TensorFlow and Keras implementation of the Mask R-CNN architecture. It provides a framework for performing simultaneous object detection and instance segmentation, transforming raw images into segmented masks and bounding boxes for individual object identification. The toolset enables custom computer vision training through fine-tuning pre-trained weights and integrating user-provided datasets. It includes capabilities for distributed GPU training to accelerate the optimization of large vision models. The framework covers model evaluation using standard precision metrics an

    Pythoninstance-segmentationkerasmask-rcnn
    在 GitHub 上查看↗25,564
  • bing-su/adetailerBing-su 的头像

    Bing-su/adetailer

    4,763在 GitHub 上查看↗

    Adetailer is a Stable Diffusion inpainting extension and automated detail enhancer that identifies specific image regions to improve quality through targeted inpainting. It functions as an AI image masking tool that uses detection models to create precise masks for automated image editing. The system distinguishes itself by integrating structural guides, such as depth and pose, to constrain the inpainting process and maintain anatomical consistency. It also supports object-specific prompt assignment, allowing unique text instructions to be mapped to multiple detected objects within a single i

    Pythonsd-webuistable-diffusion-webuistable-diffusion-webui-plugin
    在 GitHub 上查看↗4,763
  • mozilla/firamozilla 的头像

    mozilla/Fira

    5,151在 GitHub 上查看↗

    Fira is an open-source sans-serif typeface family and digital typography asset. It provides a collection of high-legibility fonts designed for clarity and readability across various screen sizes, resolutions, and operating system interfaces. The project delivers a standardized font resource for web integration and user interface typography. It consists of professional letterforms optimized for digital displays to ensure consistent character rendering. The assets are developed according to the OpenType specification and include Unicode glyph mapping and variable-weight glyphs. The design inco

    CSSabandonedunmaintained
    在 GitHub 上查看↗5,151
  • ub-mannheim/tesseractUB-Mannheim 的头像

    UB-Mannheim/tesseract

    4,111在 GitHub 上查看↗

    Tesseract is an optical character recognition engine and tool designed to convert printed or handwritten text from images into machine-readable digital text. It functions as a multilingual text extractor and a document digitization pipeline that transforms scanned images into structured digital formats. The project includes a framework for training custom scripts and language-specific models, allowing the engine to recognize new languages or unique fonts through custom training data. Its capabilities cover automated text extraction, digital archive digitization, and the export of recognized

    C++lstmocrocr-d
    在 GitHub 上查看↗4,111
  • nmac427/swiftocrNMAC427 的头像

    NMAC427/SwiftOCR

    4,632在 GitHub 上查看↗

    SwiftOCR is a Swift library for performing optical character recognition and text extraction from images. It functions as a neural network text recognizer and OCR model trainer designed for iOS and macOS applications. The project provides tools for custom font training, allowing users to teach neural networks to recognize specific typography or unique character sets by processing custom datasets. This enables the system to identify short alphanumeric sequences based on specific target character mappings. The library includes an image-preprocessing pipeline to clean and transform visual data

    Swift
    在 GitHub 上查看↗4,632
  • turing-project/writegptTuring-Project 的头像

    Turing-Project/WriteGPT

    5,301在 GitHub 上查看↗

    WriteGPT is an end-to-end essay automation system that combines visual recognition and automated text generation to convert images into finished digital documents. It functions as a creative text generator and document processor, utilizing language models to produce long-form written content and essays. The system integrates a neural text fluency evaluator to score the linguistic quality and naturalness of generated prose. It also includes a transformer-based text summarizer to condense long documents into concise summaries. The project provides a pipeline for optical character recognition t

    Python
    在 GitHub 上查看↗5,301
  • qwenlm/qwen2-vlQwenLM 的头像

    QwenLM/Qwen2-VL

    19,404在 GitHub 上查看↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Jupyter Notebook
    在 GitHub 上查看↗19,404
  • breezedeus/pix2textbreezedeus 的头像

    breezedeus/Pix2Text

    3,012在 GitHub 上查看↗

    Pix2Text is an optical character recognition system and document conversion tool designed to transform images and PDFs into Markdown. It functions as a multilingual OCR engine supporting over 80 languages, a LaTeX formula recognizer for mathematical notations, and a parser integrated with vision language models. The project utilizes a hybrid pipeline to separate plain text from mathematical formulas and tabular structures within a single pass. It converts recognized formulas into LaTeX expressions and transforms detected tables and layouts into structured Markdown formatting. The system incl

    Jupyter Notebookimage-to-markdownlatexlatex-pdf
    在 GitHub 上查看↗3,012
  • casia-lmc-lab/fastsamCASIA-LMC-Lab 的头像

    CASIA-LMC-Lab/FastSAM

    8,364在 GitHub 上查看↗

    FastSAM is an image segmentation framework that uses convolutional neural networks to isolate visual elements and generate masks for detectable objects within images. It provides a system for both automatic all-object segmentation and promptable image segmentation. The project utilizes an inference-optimized architecture to reduce computational overhead, enabling faster mask generation and real-time visual analysis. It supports the creation of precise masks through various prompt inputs, including points, bounding boxes, and text descriptions. The framework covers broader computer vision cap

    Python
    在 GitHub 上查看↗8,364
  • clovaai/deep-text-recognition-benchmarkclovaai 的头像

    clovaai/deep-text-recognition-benchmark

    3,938在 GitHub 上查看↗

    This project is a PyTorch-based framework and toolkit for scene text recognition. It provides a deep learning pipeline for extracting characters and words from images of natural environments, covering the full process from training data preparation to model validation. The framework functions as a standardized benchmark for measuring the accuracy and inference speed of text recognition models. It includes tools for calculating recognition accuracy and measuring GPU processing time per image to evaluate model performance across consistent datasets. The system incorporates visual and sequentia

    Jupyter Notebook
    在 GitHub 上查看↗3,938
  • zuruoke/watermark-removalzuruoke 的头像

    zuruoke/watermark-removal

    4,616在 GitHub 上查看↗

    This software is a watermark removal system that uses machine learning and image inpainting to delete unwanted text or logos from images. It reconstructs missing pixels to match the original background, ensuring visual consistency through pretrained models. The project includes a masking utility to isolate specific regions for content replacement using binary masks, bounding boxes, or brush strokes. It also features a batch processor that applies these cleaning tasks to large sets of images via a predefined file list. The system handles image preparation by normalizing dimensions and aspect

    Pythondeep-learningmachine-learningpython
    在 GitHub 上查看↗4,616
  • qwenlm/qwen-vlQwenLM 的头像

    QwenLM/Qwen-VL

    6,535在 GitHub 上查看↗
    Pythonlarge-language-modelsvision-language-model
    在 GitHub 上查看↗6,535
  • rednote-hilab/dots.ocrrednote-hilab 的头像

    rednote-hilab/dots.ocr

    7,695在 GitHub 上查看↗

    dots.ocr is a suite of software utilities for document layout analysis, multilingual optical character recognition, and scene text digitization. It functions as an engine for extracting digital text and structured layout data from images and PDFs across various human scripts. The project includes a specialized transformer for converting charts, diagrams, and chemical formulas from raster images into scalable vector graphics. It also provides a pipeline to transform extracted text and structural layout from documents and web screenshots into formatted Markdown files. The system covers capabil

    Python
    在 GitHub 上查看↗7,695
  • sjvasquez/handwriting-synthesissjvasquez 的头像

    sjvasquez/handwriting-synthesis

    4,779在 GitHub 上查看↗

    This project is a neural script generator that uses a recurrent neural network to synthesize human-like handwriting. It maps ASCII text characters to realistic pen stroke coordinates through an attention mechanism to mimic natural writing patterns. The system allows for handwriting style customization by adjusting priming and biasing parameters to control the neatness and stylistic characteristics of the generated text. Users can also define output formatting, including stroke colors and line widths, for the resulting digital scripts. The project includes a full neural network training workf

    Pythonhandwriting-generationhandwriting-synthesisrecurrent-neural-networks
    在 GitHub 上查看↗4,779
  • liuruoze/easyprliuruoze 的头像

    liuruoze/EasyPR

    6,425在 GitHub 上查看↗

    EasyPR is an automatic license plate recognition system designed to detect vehicle license plates and extract alphanumeric characters from images of Chinese vehicles. It functions as a deep learning OCR tool that converts image regions of license plates into machine-readable text strings. The system includes a specialized detector for identifying vehicle plates within unconstrained environments and complex visual backgrounds. It also provides a synthetic data generator to create artificial image datasets used to train and improve the accuracy of the recognition models. The project covers a m

    C++artificial-intelligenceartificial-neural-networkschinese-characters
    在 GitHub 上查看↗6,425
  • microsoft/unilmmicrosoft 的头像

    microsoft/unilm

    22,030在 GitHub 上查看↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Pythonbeitbeit-3bitnet
    在 GitHub 上查看↗22,030
  • vikhyat/moondreamvikhyat 的头像

    vikhyat/moondream

    9,769在 GitHub 上查看↗

    Moondream is a small-scale vision language model designed to reason across images to generate captions and answer natural language questions. It functions as an edge-optimized system capable of performing visual question answering, image captioning, and object detection. The project distinguishes itself through a lightweight architecture designed for local inference on embedded devices, workstations, and air-gapped hardware. It supports the execution of models on local GPUs and Apple Silicon to ensure data privacy and low latency. The system's capabilities include identifying precise object

    Python
    在 GitHub 上查看↗9,769
  • mono-company/mono-iconsmono-company 的头像

    mono-company/mono-icons

    995在 GitHub 上查看↗

    Mono-icons is a collection of scalable vector graphics and web icon fonts designed for consistent visual representation across digital interfaces. The project provides a repository of resolution-independent assets that maintain visual clarity when scaled for different screen sizes and display densities. The library utilizes OpenType font packaging to bundle vector paths into a unified format, allowing icons to be rendered as text. By mapping semantic class names to specific character codes, the system enables the injection of visual symbols into web elements through stylesheet declarations. T

    HTMLiconssvgsvg-icons
    在 GitHub 上查看↗995
  • tectonic-typesetting/tectonictectonic-typesetting 的头像

    tectonic-typesetting/tectonic

    4,591在 GitHub 上查看↗

    Tectonic is a self-contained TeX typesetting engine and automated distribution system that processes LaTeX source files into formatted documents. It functions as a single-binary executable that removes the requirement for a pre-installed local toolchain. The system implements a zero-configuration workflow by automatically fetching required TeX packages and dependencies from remote repositories on demand during the compilation process. It also provides native support for Unicode and OpenType fonts to render modern typography. The engine can be used as a programmatic library to embed typesetti

    Crusttextex-engine
    在 GitHub 上查看↗4,591