awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
datalab-to avatar

datalab-to/surya

0
View on GitHub↗
20,889 स्टार्स·1,498 फोर्क्स·Python·Apache-2.0·7 व्यूज़www.datalab.to↗

Surya

Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion.

The platform distinguishes itself through a robust pipeline-based architecture that allows users to chain analysis tasks into versioned, reusable sequences. It supports high-volume operations through batch processing and provides granular control over data extraction via schema management and confidence scoring. For enterprise requirements, it offers containerized deployment options that allow for on-premises execution, ensuring data privacy and security while maintaining consistent performance across environments.

Beyond core analysis, the system includes integrated management for document lifecycles, storage, and event-driven notifications via webhooks. It provides a strongly-typed software development kit to facilitate programmatic interaction, alongside monitoring tools that track system health and usage metrics. Security is maintained through API access controls, request throttling, and payload validation for event notifications.

Features

  • Document Analysis - Performs text recognition, layout analysis, and reading order detection using typed clients and asynchronous requests.
  • Document Conversion - Transforms PDFs, images, and other files into structured formats like markdown, HTML, or JSON for automated data systems.
  • Document and Unstructured Extraction - Transforms unstructured documents like PDFs and images into structured machine-readable formats for business pipelines.
  • Structured Data Extraction - Parses unstructured document content into predefined fields using centralized schemas for consistent machine-readable output.
  • Document Automation Pipelines - Chains multiple analysis tasks into versioned and reusable workflows to automate complex document transformation.
  • Data Pipeline Orchestration - Chains multiple document analysis steps into versioned and reusable sequences to automate complex data extraction workflows.
  • Document Processing Pipelines - Chains document analysis, text recognition, and layout segmentation tasks into versioned, automated, and reusable workflows.
  • Pipeline Orchestration - Connects multiple document processing tasks into versioned and reusable pipelines for complex extraction.
  • Self-Hosted Deployment Platforms - Supports containerized deployment on private infrastructure to provide full control over document processing environments.
  • Form Automation - Programmatically injects structured data into PDF fields and visual overlays to streamline reporting.
  • Form Automation - Injects structured data into native PDF fields or visual document overlays to generate completed forms.
  • Batch Processing - Executes analysis tasks across large document collections simultaneously to improve throughput for high-volume workloads.
  • Data Schema Management - Defines and stores data structures centrally to reference them by identifier across multiple extraction requests.
  • Document Analysis Services - Deploys containerized services to perform local text recognition and layout analysis with strict data privacy.
  • API Access Control - Limits usage and spending by assigning unique keys to different environments and rotating them.
  • Form and Input Management - Programmatically populates digital forms and injects structured data into document overlays for automated reporting.
  • Analysis SDKs - Provides a typed interface for integrating advanced text recognition and document conversion into custom software.
  • Document Management Systems - Manages document lifecycles through centralized storage, batch processing, and automated notifications.
  • Confidence Scoring - Calculates and returns numerical reliability ratings for each extracted field to assess recognition accuracy.
  • Document Processing Platforms - Provides a strongly-typed interface for executing document conversion, structured data extraction, and pipeline management.
  • Service Containerization - Packages document processing logic into isolated containers for consistent local execution and secure on-premises deployment.
  • Webhook Security - Includes a configurable secret in event payloads to allow receiving servers to validate incoming notifications.
  • Documentation Generators - Creates structured DOCX files from markdown input while maintaining support for tracked changes and custom content tags.
  • File Storage Management - Handles file lifecycle operations including uploading, listing, metadata retrieval, and deletion in remote storage.
  • Webhook Notifications - Triggers automated callbacks to specified endpoints upon task completion to eliminate manual polling.
  • Container Isolation - Isolates document processing containers behind reverse proxies and firewalls to restrict network access.
  • Webhook Event Notifications - Triggers automated callbacks to external endpoints upon task completion to eliminate manual polling.
  • Request Throttling - Limits the size and volume of document processing requests to ensure system stability during high-traffic periods.
  • Performance Monitoring - Tracks request volumes, processing latency, and system health to provide visibility into operational status.
  • Revision Extraction - Identifies and extracts revision history, redlines, and tracked changes from word processing files into structured output.
  • Document Segmentation - Identifies and isolates distinct sections within documents to improve data extraction accuracy.
  • Type-Safe Development - Provides a strongly-typed software development kit to simplify programmatic interaction and ensure reliable data structures.
  • Workflow Versioning - Maintains fixed snapshots of processing configurations to ensure production stability during iterative pipeline development.
  • Usage Analytics - Queries historical request volumes and success rates to inform infrastructure capacity planning.

स्टार हिस्ट्री

datalab-to/surya के लिए स्टार हिस्ट्री चार्टdatalab-to/surya के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

अक्सर पूछे जाने वाले प्रश्न

datalab-to/surya क्या करता है?

Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion.

datalab-to/surya की मुख्य विशेषताएं क्या हैं?

datalab-to/surya की मुख्य विशेषताएं हैं: Document Analysis, Document Conversion, Document and Unstructured Extraction, Structured Data Extraction, Document Automation Pipelines, Data Pipeline Orchestration, Document Processing Pipelines, Pipeline Orchestration।

datalab-to/surya के कुछ ओपन-सोर्स विकल्प क्या हैं?

datalab-to/surya के ओपन-सोर्स विकल्पों में शामिल हैं: datalab-to/marker — Marker is a comprehensive document processing platform designed to automate the conversion, extraction, and… unstructured-io/unstructured — Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into… datalab-to/chandra — sChandra is a document processing platform that converts images, PDFs, Word documents, spreadsheets, and other formats… zipstack/unstract — Unstract is an unstructured data extraction system and ETL pipeline orchestrator that uses large language models to… elysiajs/elysia — Elysia is a high-performance TypeScript web framework designed for building type-safe backend services. It provides a… google-gemini/cookbook — The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides…

Surya के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Surya के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • datalab-to/markerdatalab-to का अवतार

    datalab-to/marker

    36,137GitHub पर देखें↗

    Marker is a comprehensive document processing platform designed to automate the conversion, extraction, and structuring of data from complex files. It functions as an orchestration engine that chains modular processing steps into versioned, reusable pipelines, allowing organizations to standardize document handling and automate repetitive business tasks at scale. The platform distinguishes itself through its support for secure, private infrastructure deployment, enabling users to run containerized services within their own environments to maintain strict data privacy. It features specialized

    Python
    GitHub पर देखें↗36,137
  • unstructured-io/unstructuredUnstructured-IO का अवतार

    Unstructured-IO/unstructured

    14,019GitHub पर देखें↗

    Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t

    HTMLdata-pipelinesdeep-learningdocument-image-analysis
    GitHub पर देखें↗14,019
  • datalab-to/chandradatalab-to का अवतार

    datalab-to/chandra

    4,833GitHub पर देखें↗

    sChandra is a document processing platform that converts images, PDFs, Word documents, spreadsheets, and other formats into structured output such as HTML, Markdown, or JSON while preserving layout. It can also extract specific data fields from invoices, contracts, or reports using user-defined JSON schemas, with citations back to source locations. The service supports form filling in PDF and image documents, document generation from Markdown, and extraction of tracked changes from Word files. The platform distinguishes itself with pipeline-based processing chains that combine multiple proces

    Pythonaiocr
    GitHub पर देखें↗4,833
  • zipstack/unstractZipstack का अवतार

    Zipstack/unstract

    6,669GitHub पर देखें↗

    Unstract is an unstructured data extraction system and ETL pipeline orchestrator that uses large language models to convert documents, images, and scans into structured JSON. It provides a document extraction API for integrating these capabilities into external automation tools and includes a Model Context Protocol server to connect AI agents to structured information retrieval. The system ensures data accuracy through a verification tool featuring dual-model verification and human-in-the-loop review with coordinate-based document highlighting. It utilizes natural language extraction schemas

    Pythonai-agentsdata-engineeringdocument-ai
    GitHub पर देखें↗6,669
  • Surya के सभी 30 विकल्प देखें→