awesome-repositories.com
Blog
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
datalab-to avatar

datalab-to/marker

0
View on GitHub↗
36,137 stars·2,493 forks·Python·GPL-3.0·19 vueswww.datalab.to↗

Marker

Marker is a comprehensive document processing platform designed to automate the conversion, extraction, and structuring of data from complex files. It functions as an orchestration engine that chains modular processing steps into versioned, reusable pipelines, allowing organizations to standardize document handling and automate repetitive business tasks at scale.

The platform distinguishes itself through its support for secure, private infrastructure deployment, enabling users to run containerized services within their own environments to maintain strict data privacy. It features specialized engines for schema-driven data extraction and programmatic form automation, which map unstructured content from PDFs, images, and office files into predefined data structures. Additionally, the system provides robust change tracking and analysis tools to simplify collaborative review cycles by exporting redlines and comments into structured formats.

Beyond core extraction, the platform includes a wide range of operational capabilities for managing document lifecycles. This includes asynchronous task queueing for high-throughput batch processing, granular concurrency and rate-limiting controls to ensure system stability, and event-driven webhook notifications for real-time integration with external systems. The platform also offers built-in usage analytics and monitoring tools to track performance metrics and infrastructure health.

The project provides a complete set of client-side primitives and configuration utilities to manage the entire document processing workflow. Users can interact with the service through a documented API, supported by automatic retry logic and secure credential management to ensure reliable and authorized access to processing capabilities.

Features

  • Intelligent Document Processing - Extracts structured data and text from complex PDFs, images, and office files for downstream applications.
  • Document Processing Platforms - A comprehensive service for converting, extracting, and structuring data from complex files through automated and scalable workflows.
  • Structured Data Extraction - Identifies and extracts specific information like dates or legal clauses from complex documents.
  • Pipeline Orchestration - The platform enables the creation of versioned and reusable configurations by chaining document processors to manage production deployments and iterative workflow updates.
  • Document Automation Tools - Filling out PDF and image-based forms automatically using structured data to eliminate manual entry and increase operational efficiency.
  • Data Extraction Tools - A specialized engine that identifies and maps specific information from unstructured documents into predefined schemas for programmatic use.
  • Workflow Automation - Chains multiple processing steps into versioned pipelines to standardize document handling and automate business tasks.
  • Workflow Orchestration - Chains multiple modular processing steps into versioned configurations to standardize complex document handling.
  • Document Generation Engines - Creates professional documents in standard word processing formats by converting plain text or markdown input.
  • API Authentication - Requires an API key during client initialization to verify identity and authorize access to services.
  • Task Queues - Processes long-running document conversion and extraction jobs in the background to maintain high throughput.
  • Workflow Orchestration Engines - The platform runs specialized processing workflows by referencing unique pipeline identifiers to apply custom logic, validation rules, or automated evaluation steps.
  • Usage Analytics - The platform tracks performance statistics and queue status to evaluate infrastructure health and determine when to scale resources for changing workload demands.
  • Documentation and Processing - High-accuracy PDF and document conversion tool.
  • Data Mapping Utilities - The platform maps structured data to specific fields within PDF or image documents to automate the completion of forms with high accuracy.
  • Versioning & Change Tracking - Identifies and exports tracked changes from word processing files into readable formats.
  • Batch Processing - Handles multiple documents concurrently to increase throughput and improve efficiency.
  • Webhook Notifications - Provides automated notifications via webhooks when document processing tasks finish, enabling event-driven workflows.
  • Credential Management - The platform provides secure credential storage in environment variables with rotation support and spending limits for different environments to prevent unauthorized access.
  • Network Access Controls - The platform controls incoming traffic to private deployments using firewalls and IP allowlisting to ensure that only trusted clients can communicate with the service.
  • Content Extraction Engines - Divides long or batch documents into logical sections by defining a schema that identifies specific parts.
  • Document Processing and Conversion - A conversion utility that translates various file types into structured formats like Markdown, HTML, or JSON for downstream integration.
  • Deployment Automation - Supports installing containerized services within private infrastructure to enable secure document processing.
  • Traffic Management - The platform controls request volume by enforcing rate limits and concurrent connection caps while implementing automated retry strategies for temporary server busy responses.
  • Webhook Security - Validates incoming notifications using HTTPS and request signatures to ensure authenticity and prevent unauthorized event processing.
  • Rate Limiting - Enforces throughput caps and request limits to maintain system stability during high-volume processing.
  • Containerized Environments - Packages processing services into isolated environments to enable secure, private infrastructure execution.
  • Webhooks - Communicates task completion status to external systems by pushing signed JSON payloads to user-defined endpoints.
  • Data Privacy Management - The platform allows users to retrieve results before automatic deletion and configure retention settings to minimize the storage of sensitive information.
  • Private Data Processing Environments - Deploying containerized processing services within private environments to maintain data privacy and control over sensitive document workflows.
  • Asynchronous Task Processing - Executes document tasks in the background to handle multiple files concurrently and improve throughput.
  • Capacity Monitoring - The platform tracks the total number of pages currently being processed to ensure the system stays within defined capacity limits and avoids performance degradation.
  • System Monitoring - Tracks request volumes, performance metrics, and system status to maintain visibility into operational health.

Historique des stars

Graphique de l'historique des stars pour datalab-to/markerGraphique de l'historique des stars pour datalab-to/marker

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à Marker

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Marker.
  • datalab-to/suryaAvatar de datalab-to

    datalab-to/surya

    20,889Voir sur GitHub↗

    Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion. The platform distinguishes itself through a robust pipeline-based architecture that allows users to chain analysis tasks

    Python
    Voir sur GitHub↗20,889
  • unstructured-io/unstructuredAvatar de Unstructured-IO

    Unstructured-IO/unstructured

    14,019Voir sur GitHub↗

    Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t

    HTMLdata-pipelinesdeep-learningdocument-image-analysis
    Voir sur GitHub↗14,019
  • simstudioai/simAvatar de simstudioai

    simstudioai/sim

    28,796Voir sur GitHub↗

    This project is an AI agent orchestration platform that provides a visual environment for building, testing, and deploying complex automation workflows. It functions as a low-code development interface where users can chain discrete functional blocks into dependency-aware pipelines to integrate artificial intelligence with external data and services. The platform supports the creation of intelligent conversational agents, automated business processes, and multi-service API orchestrations within a unified workspace. The platform distinguishes itself through its event-driven integration engine,

    TypeScriptagent-workflowagentic-workflowagents
    Voir sur GitHub↗28,796
  • conductor-oss/conductorAvatar de conductor-oss

    conductor-oss/conductor

    31,962Voir sur GitHub↗

    Conductor is a durable workflow engine designed to orchestrate complex, long-running business processes and autonomous agent loops. It functions as a stateful execution platform that persists the entire history of a process, ensuring that workflows remain reliable and recoverable across infrastructure failures, system restarts, and transient network errors. By managing task lifecycles, worker polling, and state transitions, it provides a centralized coordination layer for distributed systems. The platform distinguishes itself through its specialized support for AI agent orchestration, allowin

    Javadistributed-systemsdurable-executiongrpc
    Voir sur GitHub↗31,962
Voir les 30 alternatives à Marker→

Questions fréquentes

Que fait datalab-to/marker ?

Marker is a comprehensive document processing platform designed to automate the conversion, extraction, and structuring of data from complex files. It functions as an orchestration engine that chains modular processing steps into versioned, reusable pipelines, allowing organizations to standardize document handling and automate repetitive business tasks at scale.

Quelles sont les fonctionnalités principales de datalab-to/marker ?

Les fonctionnalités principales de datalab-to/marker sont : Intelligent Document Processing, Document Processing Platforms, Structured Data Extraction, Pipeline Orchestration, Document Automation Tools, Data Extraction Tools, Workflow Automation, Workflow Orchestration.

Quelles sont les alternatives open-source à datalab-to/marker ?

Les alternatives open-source à datalab-to/marker incluent : datalab-to/surya — Surya is a document processing platform designed to transform unstructured files into structured, machine-readable… unstructured-io/unstructured — Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into… simstudioai/sim — This project is an AI agent orchestration platform that provides a visual environment for building, testing, and… conductor-oss/conductor — Conductor is a durable workflow engine designed to orchestrate complex, long-running business processes and autonomous… kovidgoyal/calibre — Calibre is a comprehensive suite for digital library management, serving as a centralized hub for organizing,… stack-auth/stack-auth — Stack Auth is an open-source authentication and authorization platform that provides pre-built UI components, OAuth…