awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to sinelaw/xml-to-json-fast

Open-source alternatives to Xml To Json Fast

30 open-source projects similar to sinelaw/xml-to-json-fast, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Xml To Json Fast alternative.

  • engali94/xmljsonE

    engali94/XMLJson

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • charmparticle/xpeC

    charmparticle/xpe

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • danburzo/hredD

    danburzo/hred

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • trailofbits/graphtageAvatar trailofbits

    trailofbits/graphtage

    2,472Vezi pe GitHub↗

    A semantic diff utility and library for tree-like files such as JSON, JSON5, XML, HTML, YAML, and CSV.

    Pythoncommand-line-tooldiffgraph-algorithms
    Vezi pe GitHub↗2,472
  • jheusser/csvfixJ

    jheusser/csvfix

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • jpmens/joAvatar jpmens

    jpmens/jo

    4,868Vezi pe GitHub↗

    Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell arguments and standard input. It functions as a data processing tool that transforms raw input into structured formats, enabling the generation of complex payloads for APIs, configuration files, and automated data pipelines. The tool distinguishes itself through its ability to resolve hierarchical data structures using delimiter-based path definitions and its integrated type-inference engine, which automatically casts input values into native boolean, numeric, or null types. Users can

    C
    Vezi pe GitHub↗4,868

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Find more with AI search
  • apache-spark-on-k8s/sparkA

    apache-spark-on-k8s/spark

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • allenai/olmocrAvatar allenai

    allenai/olmocr

    17,396Vezi pe GitHub↗

    Olmocr is a distributed document processing framework designed to convert PDF and image files into structured markdown. It functions as a vision-based document parser that utilizes multimodal neural networks to interpret complex visual layouts and translate them into standardized text representations. The system operates as a remote inference orchestrator, offloading heavy document analysis tasks to external servers or cloud APIs to minimize local computational requirements. By employing a stateless worker architecture, it decouples document ingestion from inference, allowing for the distribu

    Python
    Vezi pe GitHub↗17,396
  • aeon-toolkit/aeonAvatar aeon-toolkit

    aeon-toolkit/aeon

    1,404Vezi pe GitHub↗

    A toolkit for time series machine learning and deep learning

    Python
    Vezi pe GitHub↗1,404
  • bjornharrtell/jstsAvatar bjornharrtell

    bjornharrtell/jsts

    1,555Vezi pe GitHub↗

    JavaScript Topology Suite

    JavaScript
    Vezi pe GitHub↗1,555
  • borkdude/jetB

    borkdude/jet

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • bugra9/gdal3.jsAvatar bugra9

    bugra9/gdal3.js

    428Vezi pe GitHub↗

    gdal3.js is a port of Gdal applications (gdaltranslate, ogr2ogr, gdalrasterize, gdalwarp, gdaltransform) to Webassembly. It allows you to convert raster and vector geospatial data to various formats and coordinate systems.

    JavaScript
    Vezi pe GitHub↗428
  • bytedance/dolphinAvatar bytedance

    bytedance/Dolphin

    8,820Vezi pe GitHub↗

    Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers

    Pythondocument-analysislayout-analysisocr
    Vezi pe GitHub↗8,820
  • catchthetornado/pdf-extract-apiC

    CatchTheTornado/pdf-extract-api

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • bitcoinexchangefh/bitcoinexchangefhAvatar BitcoinExchangeFH

    BitcoinExchangeFH/BitcoinExchangeFH

    952Vezi pe GitHub↗

    BitcoinExchangeFH is a slim application to record the price depth and trades in various exchanges. You can set it up quickly and record the all the exchange data in a few minutes!

    Python
    Vezi pe GitHub↗952
  • alibaba/v6dA

    alibaba/v6d

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • conardli/easy-datasetAvatar ConardLi

    ConardLi/easy-dataset

    13,394Vezi pe GitHub↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    JavaScriptdatasetfine-tuningjavascript
    Vezi pe GitHub↗13,394
  • arthur-e/wicketAvatar arthur-e

    arthur-e/Wicket

    591Vezi pe GitHub↗

    Wicket is a lightweight library for translating between Well-Known Text (WKT) and various client-side mapping frameworks: Leaflet (demo) Google Maps API (demo) ESRI ArcGIS JavaScript API Potentially any other web mapping framework through serialization and de-serialization of GeoJSON (with…

    JavaScript
    Vezi pe GitHub↗591
  • datalab-to/chandraAvatar datalab-to

    datalab-to/chandra

    4,833Vezi pe GitHub↗

    sChandra is a document processing platform that converts images, PDFs, Word documents, spreadsheets, and other formats into structured output such as HTML, Markdown, or JSON while preserving layout. It can also extract specific data fields from invoices, contracts, or reports using user-defined JSON schemas, with citations back to source locations. The service supports form filling in PDF and image documents, document generation from Markdown, and extraction of tracked changes from Word files. The platform distinguishes itself with pipeline-based processing chains that combine multiple proces

    Pythonaiocr
    Vezi pe GitHub↗4,833
  • dbohdan/remarshalD

    dbohdan/remarshal

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • dflemstr/rqAvatar dflemstr

    dflemstr/rq

    2,300Vezi pe GitHub↗

    Record Query - A tool for doing record analysis and transformation

    Rustavrocommand-line-tooljavascript
    Vezi pe GitHub↗2,300
  • ds4sd/doclingAvatar DS4SD

    DS4SD/docling

    62,172Vezi pe GitHub↗

    Docling is a multimodal content converter and document parser designed to transform PDFs, Office files, and HTML into structured Markdown or JSON for generative AI applications. It functions as an OCR document processor and a PDF layout analyzer that extracts tables, charts, and hierarchical structures while preserving the original page layout. The system operates as a local-first inference engine, allowing for the processing of sensitive data in air-gapped environments without external network connectivity. It can also be deployed as an API or a Model Context Protocol server to provide parsi

    Python
    Vezi pe GitHub↗62,172
  • ekzhu/datasketchAvatar ekzhu

    ekzhu/datasketch

    2,932Vezi pe GitHub↗

    MinHash, LSH, LSH Forest, Weighted MinHash, HyperLogLog, HyperLogLog++, LSH Ensemble and HNSW

    Pythondata-sketchesdata-summaryhnsw
    Vezi pe GitHub↗2,932
  • emilstenstrom/justhtmlAvatar EmilStenstrom

    EmilStenstrom/justhtml

    1,143Vezi pe GitHub↗

    A pure Python HTML5 parser that just works. No C extensions to compile. No system dependencies to install. No complex API to learn.

    Python
    Vezi pe GitHub↗1,143
  • chatdoc-com/ocrfluxAvatar chatdoc-com

    chatdoc-com/OCRFlux

    2,514Vezi pe GitHub↗

    OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.

    Python
    Vezi pe GitHub↗2,514
  • epistates/treemdAvatar Epistates

    Epistates/treemd

    647Vezi pe GitHub↗

    A (TUI/CLI) markdown navigator with tree-based structural navigation.

    Rust
    Vezi pe GitHub↗647
  • ericchiang/pupAvatar ericchiang

    ericchiang/pup

    8,427Vezi pe GitHub↗

    Pup is a command line tool for extracting and filtering data from HTML documents using CSS selectors. It functions as a parser and selector engine that isolates specific elements based on tags, IDs, classes, and attributes. The project provides utilities for converting selected HTML nodes into plain text, attribute values, or structured JSON objects. It includes a markup formatter that corrects missing tags and applies consistent indentation to improve the readability of HTML documents. The tool handles the retrieval of text content and attributes through a CSS selector engine, supporting co

    HTML
    Vezi pe GitHub↗8,427
  • flatgeobuf/flatgeobufAvatar flatgeobuf

    flatgeobuf/flatgeobuf

    809Vezi pe GitHub↗

    A performant binary encoding for geographic data based on flatbuffers that can hold a collection of Simple Features including circular interpolations as defined by SQL-MM Part 3.

    Rust
    Vezi pe GitHub↗809
  • fslaborg/deedleAvatar fslaborg

    fslaborg/Deedle

    1,004Vezi pe GitHub↗

    Deedle

    F#
    Vezi pe GitHub↗1,004
  • argilla-io/distilabelAvatar argilla-io

    argilla-io/distilabel

    3,277Vezi pe GitHub↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Python
    Vezi pe GitHub↗3,277