awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
axa-group avatar

axa-group/Parsr

0
View on GitHub↗
6,178 stars·324 forks·JavaScript·Apache-2.0·31 views

Parsr

Parsr is an unstructured data extractor and document parsing pipeline that converts raw files and images into cleaned, machine-readable formats. It functions as a document layout analyzer and a pipeline for extracting structured data and labels using large language models.

The system includes a document parsing visualizer, providing a graphical interface to upload documents and inspect the resulting structured data output.

The project covers document digitization workflows, including layout analysis to detect headings, tables, and lists, and automated data entry through the cleaning and enrichment of unstructured content.

Features

  • Information Extraction - Uses large language models to map unstructured text segments into predefined structured data schemas.
  • Document Layout Analysis - Detects structural elements like tables and headings to reconstruct the original hierarchy of a document.
  • Optical Character Recognitions - Transforms image-based documents into machine-readable formats using optical character recognition and structural parsing.
  • Document Digitization Tools - Converts physical or digital image files into clean text and organized data for downstream applications.
  • Document Layout Analyzers - Systematically detects structural elements to reconstruct the original hierarchy of a document.
  • Semantic Document Structuring - Detects headings, tables, and lists to regenerate the original semantic hierarchy of a document.
  • Document Parsing Pipelines - Implements a pipeline using large language models to extract structured data and labels from documents.
  • Document and Unstructured Extraction - Turns raw documents and images into structured data formats for automated processing and machine reading.
  • Schema-Driven Extraction - Organizes extracted document fragments into a structured hierarchy based on target data definitions.
  • Structured Data Extraction - Transforms unstructured files and images into enriched data formats for machine reading.
  • Unstructured Data Transformation Tools - Converts raw files and images into cleaned, machine-readable formats for data entry and automation.
  • Automated Data Entry Tools - Prepares unstructured content for direct entry into databases by cleaning and enriching extracted data.
  • Multi-Stage Text Normalizers - Implements a sequential pipeline of normalization and noise removal passes to improve document data quality.
  • Data Cleaning Pipelines - Processes raw document inputs to produce label-enriched and cleaned information for automation.
  • Document Format Converters - Transforms images and files into organized text or data formats for downstream applications.
  • Visual Document Parsing - Provides visual parsing capabilities to verify extracted structured data against original images.
  • Parsing Result Inspectors - Displays a graphical interface to upload documents and visually inspect the resulting structured data.
  • Data Extraction Visualizers - Provides a graphical tool to compare original document layouts against the extracted structured data output.
  • Parsing Visualizers - Offers a graphical interface for uploading documents and inspecting the resulting structured data output.
  • OCR - Document-to-structured-data conversion tool.

Star history

Star history chart for axa-group/parsrStar history chart for axa-group/parsr

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Parsr

These projects share indexed features with Parsr. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • zipstack/unstractZipstack avatar

    Zipstack/unstract

    6,669View on GitHub↗

    Unstract is an unstructured data extraction system and ETL pipeline orchestrator that uses large language models to convert documents, images, and scans into structured JSON. It provides a document extraction API for integrating these capabilities into external automation tools and includes a Model Context Protocol server to connect AI agents to structured information retrieval. The system ensures data accuracy through a verification tool featuring dual-model verification and human-in-the-loop review with coordinate-based document highlighting. It utilizes natural language extraction schemas

    Pythonai-agentsdata-engineeringdocument-ai
    View on GitHub↗6,669
  • cinnamon/kotaemonCinnamon avatar

    Cinnamon/kotaemon

    25,139View on GitHub↗

    Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines. The system distinguishes itself through a highly modular architecture that emphasizes component-based composition and schema-driven data exchange. It supports autonomous agents capable of decomposing complex q

    Pythonchatbotllmsopen-source
    View on GitHub↗25,139
  • rednote-hilab/dots.ocrrednote-hilab avatar

    rednote-hilab/dots.ocr

    7,695View on GitHub↗

    dots.ocr is a suite of software utilities for document layout analysis, multilingual optical character recognition, and scene text digitization. It functions as an engine for extracting digital text and structured layout data from images and PDFs across various human scripts. The project includes a specialized transformer for converting charts, diagrams, and chemical formulas from raster images into scalable vector graphics. It also provides a pipeline to transform extracted text and structural layout from documents and web screenshots into formatted Markdown files. The system covers capabil

    Python
    View on GitHub↗7,695
  • opendatalab/pdf-extract-kitopendatalab avatar

    opendatalab/PDF-Extract-Kit

    9,724View on GitHub↗

    PDF-Extract-Kit is a document extraction toolkit designed to convert PDF documents into structured formats such as Markdown, HTML, and LaTeX. It functions as a multi-stage parsing framework that combines a document layout analyzer, a formula recognition engine, an OCR text extractor, and a table extraction system. The project focuses on recovering complex document elements by translating images of mathematical formulas and tabular structures into editable source code. It utilizes model-driven layout analysis to identify structural elements in reports and textbooks while ignoring noise like wa

    Python
    View on GitHub↗9,724
Compare all 30 related projects→

Frequently asked questions

What does axa-group/parsr do?

Parsr is an unstructured data extractor and document parsing pipeline that converts raw files and images into cleaned, machine-readable formats. It functions as a document layout analyzer and a pipeline for extracting structured data and labels using large language models.

What are the main features of axa-group/parsr?

The main features of axa-group/parsr are: Information Extraction, Document Layout Analysis, Optical Character Recognitions, Document Digitization Tools, Document Layout Analyzers, Semantic Document Structuring, Document Parsing Pipelines, Document and Unstructured Extraction.

Which projects share features with axa-group/parsr?

Projects with overlapping indexed features include: zipstack/unstract — Unstract is an unstructured data extraction system and ETL pipeline orchestrator that uses large language models to… cinnamon/kotaemon — Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document… rednote-hilab/dots.ocr — dots.ocr is a suite of software utilities for document layout analysis, multilingual optical character recognition,… opendatalab/pdf-extract-kit — PDF-Extract-Kit is a document extraction toolkit designed to convert PDF documents into structured formats such as… ds4sd/docling — Docling is a multimodal content converter and document parser designed to transform PDFs, Office files, and HTML into… browserbase/mcp-server-browserbase — This project is an MCP browser automation server that connects large language models to headless cloud browsers. It…