1 dépôt
Dividing large documents into smaller overlapping sections to maintain recognition detail.
Distinct from Optical Character Recognition: Specific tiling strategy for OCR accuracy, distinct from general OCR parsing.
Explore 1 awesome GitHub repository matching graphics & multimedia · Multi-Crop Processing. Refine with filters or upvote what's useful.
GOT-OCR2.0 is an end-to-end optical character recognition system and document text extractor. It utilizes a unified transformer architecture to recognize and extract plain and formatted text from diverse images and documents. The system features a multi-crop processing method that divides high-resolution or dense documents into smaller sections to maintain recognition detail. It also includes a renderer that transforms recognized text into HTML to preserve the original structure and layout of the document. The project provides a framework for fine-tuning pre-trained models on custom datasets
Captures high-detail text across large or dense documents by dividing complex images into smaller sections.