awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
5

599yongyang/DatasetLoom

0
View on GitHub↗
0 stars·0 forks·22 views

DatasetLoom

Features

  • Data Processing - Platform for building and evaluating multimodal datasets.
  • Data Processing Tools - Platform for building and evaluating multimodal datasets.

Star history

Star history chart for 599yongyang/datasetloomStar history chart for 599yongyang/datasetloom

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What are the main features of 599yongyang/datasetloom?

The main features of 599yongyang/datasetloom are: Data Processing, Data Processing Tools.

Which projects share features with 599yongyang/datasetloom?

Projects with overlapping indexed features include: jpmens/jo — Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell… chatdoc-com/ocrflux — OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion,… allenai/olmocr — Olmocr is a distributed document processing framework designed to convert PDF and image files into structured… argilla-io/distilabel — Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable… catchthetornado/pdf-extract-api. bytedance/dolphin — Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital…

Projects sharing features with DatasetLoom

These projects share indexed features with DatasetLoom. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • jpmens/jojpmens avatar

    jpmens/jo

    4,868View on GitHub↗

    Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell arguments and standard input. It functions as a data processing tool that transforms raw input into structured formats, enabling the generation of complex payloads for APIs, configuration files, and automated data pipelines. The tool distinguishes itself through its ability to resolve hierarchical data structures using delimiter-based path definitions and its integrated type-inference engine, which automatically casts input values into native boolean, numeric, or null types. Users can

    C
    View on GitHub↗4,868
  • bytedance/dolphinbytedance avatar

    bytedance/Dolphin

    8,820View on GitHub↗

    Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers

    Pythondocument-analysislayout-analysisocr
    View on GitHub↗8,820
  • allenai/olmocrallenai avatar

    allenai/olmocr

    17,396View on GitHub↗

    Olmocr is a distributed document processing framework designed to convert PDF and image files into structured markdown. It functions as a vision-based document parser that utilizes multimodal neural networks to interpret complex visual layouts and translate them into standardized text representations. The system operates as a remote inference orchestrator, offloading heavy document analysis tasks to external servers or cloud APIs to minimize local computational requirements. By employing a stateless worker architecture, it decouples document ingestion from inference, allowing for the distribu

    Python
    View on GitHub↗17,396
  • argilla-io/distilabelargilla-io avatar

    argilla-io/distilabel

    3,277View on GitHub↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Python
    View on GitHub↗3,277
Compare all 30 related projects→