awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
deepseek-ai avatar

deepseek-ai/DeepSeek-OCR

0
View on GitHub↗
22,498 stars·2,061 forks·Python·mit·41 views

DeepSeek OCR

DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows.

The system distinguishes itself through a high-throughput architecture that utilizes hardware-accelerated batch inference to process large volumes of visual data. It incorporates dynamic resolution scaling to manage the balance between visual detail and token consumption, ensuring that image content is compressed into optimized formats for efficient model ingestion.

The framework includes comprehensive capabilities for scaling inference throughput across distributed backends to maintain consistent performance under heavy traffic. It also integrates automated benchmarking tools to evaluate the accuracy and speed of text extraction across diverse datasets, ensuring reliable output quality during system operations.

Features

  • Document Inference Pipelines - Provides a high-throughput architecture for scaling visual data analysis and text extraction.
  • Optical Character Recognition - Performs optical character recognition to convert visual text into machine-readable tokens.
  • Vision Processing Frameworks - Encodes visual data into compact tokens for efficient document analysis by language models.
  • Optical Character Recognition Engines - Converts image-based text into machine-readable tokens for automated data extraction.
  • Multimodal Large Language Models - Prepares visual data for ingestion into multimodal large language models.
  • Visual Tokenizers - Compresses image content into optimized token representations for visual analysis.
  • Hardware-Accelerated Inference - Executes high-throughput document extraction using hardware-accelerated parallel processing.
  • Model Performance Benchmarking - Provides automated performance and accuracy benchmarking for visual processing models.
  • High-Throughput Inference Services - Scales inference throughput by distributing extraction tasks across high-performance backends.
  • Document Processing Engines - Provides high-performance pipelines for batch processing and text extraction from documents.
  • Visual Encoders - Converts raw pixel data into compressed vector representations for language model ingestion.
  • Multimodal Models - Specialized multimodal model for optical character recognition.
  • Data Extraction and OCR - Optical character recognition for document processing.
  • Distributed Orchestration - Orchestrates distributed compute nodes to maintain high-throughput visual processing.
  • Resolution Scaling - Adjusts input image dimensions at runtime to balance visual detail against token consumption.
  • Visual Token Compression - Encodes image content into compact token representations for efficient model processing.

Star history

Star history chart for deepseek-ai/deepseek-ocrStar history chart for deepseek-ai/deepseek-ocr

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with DeepSeek OCR

These projects share indexed features with DeepSeek OCR. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • microsoft/unilmmicrosoft avatar

    microsoft/unilm

    22,030View on GitHub↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Pythonbeitbeit-3bitnet
    View on GitHub↗22,030
  • sgl-project/sglangsgl-project avatar

    sgl-project/sglang

    29,079View on GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Pythonattentionblackwellcuda
    View on GitHub↗29,079
  • vikparuchuri/markerVikParuchuri avatar

    VikParuchuri/marker

    36,164View on GitHub↗

    Marker is an LLM-powered document parser and OCR pipeline designed to convert PDFs and unstructured files into structured markdown, JSON, and HTML. It functions as a data preprocessor that transforms complex documents into machine-readable formats while preserving tables, equations, and layout structures. The system utilizes large language models to refine OCR accuracy, clean mathematical notation, and merge fragmented tables across multiple pages. It employs model-based layout analysis to predict block types and bounding boxes, ensuring a more precise conversion of document elements. Capabi

    Python
    View on GitHub↗36,164
  • qwenlm/qwen2-vlQwenLM avatar

    QwenLM/Qwen2-VL

    19,404View on GitHub↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Jupyter Notebook
    View on GitHub↗19,404
Compare all 30 related projects→

Frequently asked questions

What does deepseek-ai/deepseek-ocr do?

DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows.

What are the main features of deepseek-ai/deepseek-ocr?

The main features of deepseek-ai/deepseek-ocr are: Document Inference Pipelines, Optical Character Recognition, Vision Processing Frameworks, Optical Character Recognition Engines, Multimodal Large Language Models, Visual Tokenizers, Hardware-Accelerated Inference, Model Performance Benchmarking.

Which projects share features with deepseek-ai/deepseek-ocr?

Projects with overlapping indexed features include: microsoft/unilm — This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based… sgl-project/sglang — Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It… vikparuchuri/marker — Marker is an LLM-powered document parser and OCR pipeline designed to convert PDFs and unstructured files into… qwenlm/qwen2-vl — Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text,… clovaai/deep-text-recognition-benchmark — This project is a PyTorch-based framework and toolkit for scene text recognition. It provides a deep learning pipeline… baiyuetribe/paper2gui — Paper2gui is a multi-modal AI toolkit and model GUI wrapper designed to deploy and run various artificial intelligence…