awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to modelscope/data-juicer

Open-source alternatives to Data Juicer

30 open-source projects similar to modelscope/data-juicer, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Data Juicer alternative.

  • jpmens/jojpmens का अवतार

    jpmens/jo

    4,868GitHub पर देखें↗

    Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell arguments and standard input. It functions as a data processing tool that transforms raw input into structured formats, enabling the generation of complex payloads for APIs, configuration files, and automated data pipelines. The tool distinguishes itself through its ability to resolve hierarchical data structures using delimiter-based path definitions and its integrated type-inference engine, which automatically casts input values into native boolean, numeric, or null types. Users can

    C
    GitHub पर देखें↗4,868
  • quivrhq/megaparsequivrhq का अवतार

    quivrhq/megaparse

    7,389GitHub पर देखें↗

    Megaparse is a document parsing tool and RAG data preprocessor designed to convert PDFs, Word documents, and presentations into clean text formats. It functions as a vision-based document extractor that recovers high-fidelity information from images and complex layouts to optimize data for large language model ingestion. The system employs multimodal AI and vision models to perform schema-preserving parsing, which maintains structural hierarchies such as tables and headers. It utilizes lossless structural transformation to turn layout-heavy binary files into text sequences while preserving th

    Python
    GitHub पर देखें↗7,389
  • argilla-io/distilabelargilla-io का अवतार

    argilla-io/distilabel

    3,277GitHub पर देखें↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Python
    GitHub पर देखें↗3,277
  • huggingface/llm-swarmhuggingface का अवतार

    huggingface/llm-swarm

    288GitHub पर देखें↗

    Manage scalable open LLM inference endpoints in Slurm clusters

    Python
    GitHub पर देखें↗288

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Find more with AI search
  • modelscope/easydistillM

    modelscope/easydistill

    0GitHub पर देखें↗
    GitHub पर देखें↗0
  • pdf2htmlex/pdf2htmlexpdf2htmlEX का अवतार

    pdf2htmlEX/pdf2htmlEX

    5,412GitHub पर देखें↗

    pdf2htmlEX is a PDF to HTML converter that transforms documents into web pages while preserving the original layout, fonts, and formatting. It functions as a layout engine and text extractor, mapping PDF coordinate data to HTML and CSS to maintain visual fidelity. The tool converts PDF content into searchable and selectable native HTML text by embedding original document fonts. It maintains document interactivity by preserving internal links, bookmarks, and outlines, converting them into functional web navigation. The conversion process supports flexible output structures, allowing documents

    HTMLhtmlpdfpdf-document-processor
    GitHub पर देखें↗5,412
  • bytedance/dolphinbytedance का अवतार

    bytedance/Dolphin

    8,820GitHub पर देखें↗

    Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers

    Pythondocument-analysislayout-analysisocr
    GitHub पर देखें↗8,820
  • chatdoc-com/ocrfluxchatdoc-com का अवतार

    chatdoc-com/OCRFlux

    2,514GitHub पर देखें↗

    OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.

    Python
    GitHub पर देखें↗2,514
  • ekzhu/datasketchekzhu का अवतार

    ekzhu/datasketch

    2,932GitHub पर देखें↗

    MinHash, LSH, LSH Forest, Weighted MinHash, HyperLogLog, HyperLogLog++, LSH Ensemble and HNSW

    Pythondata-sketchesdata-summaryhnsw
    GitHub पर देखें↗2,932
  • huggingface/datatrovehuggingface का अवतार

    huggingface/datatrove

    3,092GitHub पर देखें↗

    Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.

    Python
    GitHub पर देखें↗3,092
  • lm-sys/llm-decontaminatorlm-sys का अवतार

    lm-sys/llm-decontaminator

    324GitHub पर देखें↗

    Code for the paper "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples"

    Python
    GitHub पर देखें↗324
  • minishlab/semhashMinishLab का अवतार

    MinishLab/semhash

    936GitHub पर देखें↗

    Fast Multimodal Semantic Deduplication & Filtering

    Pythondatasetsdeduplicationimage-dataset-cleaning
    GitHub पर देखें↗936
  • opendatalab/mineruopendatalab का अवतार

    opendatalab/MinerU

    67,734GitHub पर देखें↗

    MinerU is a document parsing pipeline designed to transform unstructured files into machine-readable, structured data. It utilizes deep learning models to perform layout analysis, identifying document regions and extracting complex content such as mathematical expressions. By combining these neural network inferences with geometric heuristics, the system reconstructs the reading order and structural hierarchy of documents to ensure accurate data representation. The project distinguishes itself through a multi-stage processing workflow that integrates layout detection, optical character recogn

    Pythonai4sciencedocument-analysisextract-data
    GitHub पर देखें↗67,734
  • opendcai/dataflowOpenDCAI का अवतार

    OpenDCAI/DataFlow

    2,926GitHub पर देखें↗

    DataFlow is an agent-based workflow orchestrator and data pipeline designed to synthesize, clean, and augment large-scale datasets for training large language models. It functions as a synthetic data generator and text curation tool, utilizing an intelligent assistant to assemble modular processing operators into functional pipelines based on user requirements. The project distinguishes itself through a low-code approach, providing a web-based visual interface for designing and monitoring multi-stage execution flows. It features an operator-based registry system that allows for the integratio

    Pythondatadata-agentdata-cleaning
    GitHub पर देखें↗2,926
  • conardli/easy-datasetConardLi का अवतार

    ConardLi/easy-dataset

    13,394GitHub पर देखें↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    JavaScriptdatasetfine-tuningjavascript
    GitHub पर देखें↗13,394
  • allenai/olmocrallenai का अवतार

    allenai/olmocr

    17,396GitHub पर देखें↗

    Olmocr is a distributed document processing framework designed to convert PDF and image files into structured markdown. It functions as a vision-based document parser that utilizes multimodal neural networks to interpret complex visual layouts and translate them into standardized text representations. The system operates as a remote inference orchestrator, offloading heavy document analysis tasks to external servers or cloud APIs to minimize local computational requirements. By employing a stateless worker architecture, it decouples document ingestion from inference, allowing for the distribu

    Python
    GitHub पर देखें↗17,396
  • 599yongyang/datasetloom5

    599yongyang/DatasetLoom

    0GitHub पर देखें↗
    GitHub पर देखें↗0
  • catchthetornado/pdf-extract-apiC

    CatchTheTornado/pdf-extract-api

    0GitHub पर देखें↗
    GitHub पर देखें↗0
  • datalab-to/chandradatalab-to का अवतार

    datalab-to/chandra

    4,833GitHub पर देखें↗

    sChandra is a document processing platform that converts images, PDFs, Word documents, spreadsheets, and other formats into structured output such as HTML, Markdown, or JSON while preserving layout. It can also extract specific data fields from invoices, contracts, or reports using user-defined JSON schemas, with citations back to source locations. The service supports form filling in PDF and image documents, document generation from Markdown, and extraction of tracked changes from Word files. The platform distinguishes itself with pipeline-based processing chains that combine multiple proces

    Pythonaiocr
    GitHub पर देखें↗4,833
  • ds4sd/doclingDS4SD का अवतार

    DS4SD/docling

    62,172GitHub पर देखें↗

    Docling is a multimodal content converter and document parser designed to transform PDFs, Office files, and HTML into structured Markdown or JSON for generative AI applications. It functions as an OCR document processor and a PDF layout analyzer that extracts tables, charts, and hierarchical structures while preserving the original page layout. The system operates as a local-first inference engine, allowing for the processing of sensitive data in air-gapped environments without external network connectivity. It can also be deployed as an API or a Model Context Protocol server to provide parsi

    Python
    GitHub पर देखें↗62,172
  • funstory-ai/babeldocfunstory-ai का अवतार

    funstory-ai/BabelDOC

    7,752GitHub पर देखें↗

    BabelDOC is a technical document translation system designed to translate PDF files while preserving their original layout and styling. It functions as a layout-preserving translator that utilizes large language models to convert content into target languages, specifically tailored for scientific and technical documents. The system distinguishes itself through specialized handling of academic content, including the identification and preservation of mathematical formulas and complex layout structures. It ensures technical accuracy by employing glossary-driven terminology enforcement, using so

    Python
    GitHub पर देखें↗7,752
  • getomni-ai/zeroxgetomni-ai का अवतार

    getomni-ai/zerox

    12,241GitHub पर देखें↗

    Zerox is a multimodal document parser and OCR tool that uses vision models to convert PDF files and images into structured Markdown text. It functions as a visual layout extraction engine, leveraging large multimodal models to digitize documents while maintaining their original structural formatting. The system differentiates itself through the use of coordinate-based element mapping and multimodal layout analysis to identify structural elements like tables, charts, and headers. It utilizes rasterization to convert vector PDF pages into high-resolution bitmaps, ensuring consistent input for t

    TypeScriptocrpdf
    GitHub पर देखें↗12,241
  • jf-tech/omniparserjf-tech का अवतार

    jf-tech/omniparser

    1,085GitHub पर देखें↗

    omniparser: a native Golang ETL streaming parser and transform library for CSV, JSON, XML, EDI, text, etc.

    Go
    GitHub पर देखें↗1,085
  • katanaml/sparrowkatanaml का अवतार

    katanaml/sparrow

    5,162GitHub पर देखें↗

    Sparrow is an LLM document extraction platform and vision-based inference engine designed to convert images and PDFs into validated structured data. It functions as an agentic workflow orchestrator that chains classification, extraction, and validation tasks into multi-step pipelines. The system distinguishes itself through a backend-agnostic inference layer that manages models across local GPUs, Apple Silicon, and cloud providers. It employs coordinate-based visual grounding to map extracted text to precise bounding box coordinates and utilizes hint-based model steering to guide attention an

    Pythonagentic-aicomputer-visiondocumentai
    GitHub पर देखें↗5,162
  • microsoft/markitdownmicrosoft का अवतार

    microsoft/markitdown

    154,485GitHub पर देखें↗

    This project is an AI-powered document processing engine designed to transform diverse file formats into structured Markdown. By leveraging multimodal language models, it performs complex layout analysis and semantic text extraction, allowing for the conversion of both unstructured files and scanned images into machine-readable content. The toolkit distinguishes itself through a modular, plugin-based architecture that orchestrates multi-stage extraction pipelines. Users can steer the parsing behavior by injecting custom instructions, enabling the system to adapt to domain-specific document st

    Pythonautogenautogen-extensionlangchain
    GitHub पर देखें↗154,485
  • mikefarah/yqmikefarah का अवतार

    mikefarah/yq

    14,913GitHub पर देखें↗

    This tool is a command-line processor designed for querying, updating, and transforming structured data files. It functions as a versatile engine for manipulating YAML, JSON, TOML, and XML documents, allowing users to perform complex operations directly from the terminal. By utilizing a path-based expression language, it enables precise navigation and modification of data structures within configuration files and infrastructure-as-code workflows. What distinguishes this tool is its ability to perform in-place document mutations while preserving original formatting, comments, and metadata. It

    Gobashclicsv
    GitHub पर देखें↗14,913
  • opendatalab/doclayout-yoloopendatalab का अवतार

    opendatalab/DocLayout-YOLO

    2,197GitHub पर देखें↗

    DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

    Python
    GitHub पर देखें↗2,197
  • opendatalab/labelllmopendatalab का अवतार

    opendatalab/LabelLLM

    1,241GitHub पर देखें↗

    The Open-Source Data Annotation Platform

    TypeScript
    GitHub पर देखें↗1,241
  • opendatalab/pdf-extract-kitopendatalab का अवतार

    opendatalab/PDF-Extract-Kit

    9,724GitHub पर देखें↗

    PDF-Extract-Kit is a document extraction toolkit designed to convert PDF documents into structured formats such as Markdown, HTML, and LaTeX. It functions as a multi-stage parsing framework that combines a document layout analyzer, a formula recognition engine, an OCR text extractor, and a table extraction system. The project focuses on recovering complex document elements by translating images of mathematical formulas and tabular structures into editable source code. It utilizes model-driven layout analysis to identify structural elements in reports and textbooks while ignoring noise like wa

    Python
    GitHub पर देखें↗9,724
  • raznem/parseraraznem का अवतार

    raznem/parsera

    1,279GitHub पर देखें↗

    Lightweight library for scraping web-sites with LLMs

    Pythonaiai-scrapingdata-extraction
    GitHub पर देखें↗1,279