awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
deepseek-ai avatar

deepseek-ai/DeepSeek-OCR

0
View on GitHub↗
22,498 Stars·2,061 Forks·Python·mit·13 Aufrufe

DeepSeek OCR

DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows.

The system distinguishes itself through a high-throughput architecture that utilizes hardware-accelerated batch inference to process large volumes of visual data. It incorporates dynamic resolution scaling to manage the balance between visual detail and token consumption, ensuring that image content is compressed into optimized formats for efficient model ingestion.

The framework includes comprehensive capabilities for scaling inference throughput across distributed backends to maintain consistent performance under heavy traffic. It also integrates automated benchmarking tools to evaluate the accuracy and speed of text extraction across diverse datasets, ensuring reliable output quality during system operations.

Features

  • Document Inference Pipelines - Provides a high-throughput architecture for scaling visual data analysis and text extraction.
  • Optical Character Recognition - Performs optical character recognition to convert visual text into machine-readable tokens.
  • Vision Processing Frameworks - Encodes visual data into compact tokens for efficient document analysis by language models.
  • Optical Character Recognition Engines - Converts image-based text into machine-readable tokens for automated data extraction.
  • Multimodal Large Language Models - Prepares visual data for ingestion into multimodal large language models.
  • Visual Tokenizers - Compresses image content into optimized token representations for visual analysis.
  • Hardware-Accelerated Inference - Executes high-throughput document extraction using hardware-accelerated parallel processing.
  • Model Performance Benchmarking - Provides automated performance and accuracy benchmarking for visual processing models.
  • High-Throughput Inference Services - Scales inference throughput by distributing extraction tasks across high-performance backends.
  • Document Processing Engines - Provides high-performance pipelines for batch processing and text extraction from documents.
  • Visual Encoders - Converts raw pixel data into compressed vector representations for language model ingestion.
  • Multimodal Models - Specialized multimodal model for optical character recognition.
  • Data Extraction and OCR - Optical character recognition for document processing.
  • Distributed Orchestration - Orchestrates distributed compute nodes to maintain high-throughput visual processing.
  • Resolution Scaling - Adjusts input image dimensions at runtime to balance visual detail against token consumption.
  • Visual Token Compression - Encodes image content into compact token representations for efficient model processing.

Star-Verlauf

Star-Verlauf für deepseek-ai/deepseek-ocrStar-Verlauf für deepseek-ai/deepseek-ocr

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu DeepSeek OCR

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit DeepSeek OCR.
  • microsoft/unilmAvatar von microsoft

    microsoft/unilm

    22,030Auf GitHub ansehen↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Pythonbeitbeit-3bitnet
    Auf GitHub ansehen↗22,030
  • sgl-project/sglangAvatar von sgl-project

    sgl-project/sglang

    29,079Auf GitHub ansehen↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Pythonattentionblackwellcuda
    Auf GitHub ansehen↗29,079
  • vikparuchuri/markerAvatar von VikParuchuri

    VikParuchuri/marker

    36,164Auf GitHub ansehen↗

    Marker is an LLM-powered document parser and OCR pipeline designed to convert PDFs and unstructured files into structured markdown, JSON, and HTML. It functions as a data preprocessor that transforms complex documents into machine-readable formats while preserving tables, equations, and layout structures. The system utilizes large language models to refine OCR accuracy, clean mathematical notation, and merge fragmented tables across multiple pages. It employs model-based layout analysis to predict block types and bounding boxes, ensuring a more precise conversion of document elements. Capabi

    Python
    Auf GitHub ansehen↗36,164
  • qwenlm/qwen2-vlAvatar von QwenLM

    QwenLM/Qwen2-VL

    19,404Auf GitHub ansehen↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Jupyter Notebook
    Auf GitHub ansehen↗19,404
Alle 30 Alternativen zu DeepSeek OCR anzeigen→

Häufig gestellte Fragen

Was macht deepseek-ai/deepseek-ocr?

DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows.

Was sind die Hauptfunktionen von deepseek-ai/deepseek-ocr?

Die Hauptfunktionen von deepseek-ai/deepseek-ocr sind: Document Inference Pipelines, Optical Character Recognition, Vision Processing Frameworks, Optical Character Recognition Engines, Multimodal Large Language Models, Visual Tokenizers, Hardware-Accelerated Inference, Model Performance Benchmarking.

Welche Open-Source-Alternativen gibt es zu deepseek-ai/deepseek-ocr?

Open-Source-Alternativen zu deepseek-ai/deepseek-ocr sind unter anderem: microsoft/unilm — This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based… sgl-project/sglang — Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It… vikparuchuri/marker — Marker is an LLM-powered document parser and OCR pipeline designed to convert PDFs and unstructured files into… qwenlm/qwen2-vl — Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text,… clovaai/deep-text-recognition-benchmark — This project is a PyTorch-based framework and toolkit for scene text recognition. It provides a deep learning pipeline… baiyuetribe/paper2gui — Paper2gui is a multi-modal AI toolkit and model GUI wrapper designed to deploy and run various artificial intelligence…