awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
datalab-to avatar

datalab-to/surya

0
View on GitHub↗
20,889 stars·1,498 forks·Python·Apache-2.0·35 viewswww.datalab.to↗

Surya

Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion.

The platform distinguishes itself through a robust pipeline-based architecture that allows users to chain analysis tasks into versioned, reusable sequences. It supports high-volume operations through batch processing and provides granular control over data extraction via schema management and confidence scoring. For enterprise requirements, it offers containerized deployment options that allow for on-premises execution, ensuring data privacy and security while maintaining consistent performance across environments.

Beyond core analysis, the system includes integrated management for document lifecycles, storage, and event-driven notifications via webhooks. It provides a strongly-typed software development kit to facilitate programmatic interaction, alongside monitoring tools that track system health and usage metrics. Security is maintained through API access controls, request throttling, and payload validation for event notifications.

Features

  • Document Analysis - Performs text recognition, layout analysis, and reading order detection using typed clients and asynchronous requests.
  • Document Conversion - Transforms PDFs, images, and other files into structured formats like markdown, HTML, or JSON for automated data systems.
  • Document and Unstructured Extraction - Transforms unstructured documents like PDFs and images into structured machine-readable formats for business pipelines.
  • Structured Data Extraction - Parses unstructured document content into predefined fields using centralized schemas for consistent machine-readable output.
  • Document Automation Pipelines - Chains multiple analysis tasks into versioned and reusable workflows to automate complex document transformation.
  • Data Pipeline Orchestration - Chains multiple document analysis steps into versioned and reusable sequences to automate complex data extraction workflows.
  • Document Processing Pipelines - Chains document analysis, text recognition, and layout segmentation tasks into versioned, automated, and reusable workflows.
  • Pipeline Orchestration - Connects multiple document processing tasks into versioned and reusable pipelines for complex extraction.
  • Self-Hosted Deployment Platforms - Supports containerized deployment on private infrastructure to provide full control over document processing environments.
  • Form Automation - Programmatically injects structured data into PDF fields and visual overlays to streamline reporting.
  • Form Automation - Injects structured data into native PDF fields or visual document overlays to generate completed forms.
  • Batch Processing - Executes analysis tasks across large document collections simultaneously to improve throughput for high-volume workloads.
  • Data Schema Management - Defines and stores data structures centrally to reference them by identifier across multiple extraction requests.
  • Document Analysis Services - Deploys containerized services to perform local text recognition and layout analysis with strict data privacy.
  • API Access Control - Limits usage and spending by assigning unique keys to different environments and rotating them.
  • Form and Input Management - Programmatically populates digital forms and injects structured data into document overlays for automated reporting.
  • Analysis SDKs - Provides a typed interface for integrating advanced text recognition and document conversion into custom software.
  • Document Management Systems - Manages document lifecycles through centralized storage, batch processing, and automated notifications.
  • Confidence Scoring - Calculates and returns numerical reliability ratings for each extracted field to assess recognition accuracy.
  • Document Processing Platforms - Provides a strongly-typed interface for executing document conversion, structured data extraction, and pipeline management.
  • Service Containerization - Packages document processing logic into isolated containers for consistent local execution and secure on-premises deployment.
  • Webhook Security - Includes a configurable secret in event payloads to allow receiving servers to validate incoming notifications.
  • Documentation Generators - Creates structured DOCX files from markdown input while maintaining support for tracked changes and custom content tags.
  • File Storage Management - Handles file lifecycle operations including uploading, listing, metadata retrieval, and deletion in remote storage.
  • Webhook Notifications - Triggers automated callbacks to specified endpoints upon task completion to eliminate manual polling.
  • Container Isolation - Isolates document processing containers behind reverse proxies and firewalls to restrict network access.
  • Webhook Event Notifications - Triggers automated callbacks to external endpoints upon task completion to eliminate manual polling.
  • Request Throttling - Limits the size and volume of document processing requests to ensure system stability during high-traffic periods.
  • Performance Monitoring - Tracks request volumes, processing latency, and system health to provide visibility into operational status.
  • Revision Extraction - Identifies and extracts revision history, redlines, and tracked changes from word processing files into structured output.
  • Document Segmentation - Identifies and isolates distinct sections within documents to improve data extraction accuracy.
  • Type-Safe Development - Provides a strongly-typed software development kit to simplify programmatic interaction and ensure reliable data structures.
  • Workflow Versioning - Maintains fixed snapshots of processing configurations to ensure production stability during iterative pipeline development.
  • Usage Analytics - Queries historical request volumes and success rates to inform infrastructure capacity planning.

Star history

Star history chart for datalab-to/suryaStar history chart for datalab-to/surya

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Surya

These projects share indexed features with Surya. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • datalab-to/markerdatalab-to avatar

    datalab-to/marker

    36,137View on GitHub↗

    Marker is a comprehensive document processing platform designed to automate the conversion, extraction, and structuring of data from complex files. It functions as an orchestration engine that chains modular processing steps into versioned, reusable pipelines, allowing organizations to standardize document handling and automate repetitive business tasks at scale. The platform distinguishes itself through its support for secure, private infrastructure deployment, enabling users to run containerized services within their own environments to maintain strict data privacy. It features specialized

    Python
    View on GitHub↗36,137
  • unstructured-io/unstructuredUnstructured-IO avatar

    Unstructured-IO/unstructured

    14,019View on GitHub↗

    Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t

    HTMLdata-pipelinesdeep-learningdocument-image-analysis
    View on GitHub↗14,019
  • datalab-to/chandradatalab-to avatar

    datalab-to/chandra

    4,833View on GitHub↗

    sChandra is a document processing platform that converts images, PDFs, Word documents, spreadsheets, and other formats into structured output such as HTML, Markdown, or JSON while preserving layout. It can also extract specific data fields from invoices, contracts, or reports using user-defined JSON schemas, with citations back to source locations. The service supports form filling in PDF and image documents, document generation from Markdown, and extraction of tracked changes from Word files. The platform distinguishes itself with pipeline-based processing chains that combine multiple proces

    Pythonaiocr
    View on GitHub↗4,833
  • zipstack/unstractZipstack avatar

    Zipstack/unstract

    6,669View on GitHub↗

    Unstract is an unstructured data extraction system and ETL pipeline orchestrator that uses large language models to convert documents, images, and scans into structured JSON. It provides a document extraction API for integrating these capabilities into external automation tools and includes a Model Context Protocol server to connect AI agents to structured information retrieval. The system ensures data accuracy through a verification tool featuring dual-model verification and human-in-the-loop review with coordinate-based document highlighting. It utilizes natural language extraction schemas

    Pythonai-agentsdata-engineeringdocument-ai
    View on GitHub↗6,669
Compare all 30 related projects→

Frequently asked questions

What does datalab-to/surya do?

Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion.

What are the main features of datalab-to/surya?

The main features of datalab-to/surya are: Document Analysis, Document Conversion, Document and Unstructured Extraction, Structured Data Extraction, Document Automation Pipelines, Data Pipeline Orchestration, Document Processing Pipelines, Pipeline Orchestration.

Which projects share features with datalab-to/surya?

Projects with overlapping indexed features include: datalab-to/marker — Marker is a comprehensive document processing platform designed to automate the conversion, extraction, and… unstructured-io/unstructured — Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into… datalab-to/chandra — sChandra is a document processing platform that converts images, PDFs, Word documents, spreadsheets, and other formats… zipstack/unstract — Unstract is an unstructured data extraction system and ETL pipeline orchestrator that uses large language models to… elysiajs/elysia — Elysia is a high-performance TypeScript web framework designed for building type-safe backend services. It provides a… google-gemini/cookbook — The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides…