awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
IBM avatar

IBM/data-prep-kit

0
View on GitHub↗
940 stars·251 forks·HTML·Apache-2.0·9 viewsdata-prep-kit.github.io/data-prep-kit↗

Data Prep Kit

Open source project for data preparation for GenAI applications

Features

  • Data Generation Frameworks - Prepares large-scale language and code data across distributed computing environments.
  • Data Processing - Toolkit for efficient unstructured data processing.

Star history

Star history chart for ibm/data-prep-kitStar history chart for ibm/data-prep-kit

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does ibm/data-prep-kit do?

Open source project for data preparation for GenAI applications

What are the main features of ibm/data-prep-kit?

The main features of ibm/data-prep-kit are: Data Generation Frameworks, Data Processing.

Which projects share features with ibm/data-prep-kit?

Projects with overlapping indexed features include: bespokelabsai/curator. argilla-io/distilabel — Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable… aimukhin/minfft — A minimalistic Fast Fourier Transform library. ankurchavda/sparklearning — A comprehensive Spark guide collated from multiple sources that can be referred to learn more about Spark or as an… antirez/smaz — Small strings compression library. allenai/olmocr — Olmocr is a distributed document processing framework designed to convert PDF and image files into structured…

Projects sharing features with Data Prep Kit

These projects share indexed features with Data Prep Kit. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • bespokelabsai/curatorbespokelabsai avatar

    bespokelabsai/curator

    1,637View on GitHub↗
    Pythonagentsdeep-learningfine-tuning
    View on GitHub↗1,637
  • argilla-io/distilabelargilla-io avatar

    argilla-io/distilabel

    3,277View on GitHub↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Python
    View on GitHub↗3,277
  • aimukhin/minfftaimukhin avatar

    aimukhin/minfft

    51View on GitHub↗

    A minimalistic Fast Fourier Transform library.

    C
    View on GitHub↗51
  • allenai/olmocrallenai avatar

    allenai/olmocr

    17,396View on GitHub↗

    Olmocr is a distributed document processing framework designed to convert PDF and image files into structured markdown. It functions as a vision-based document parser that utilizes multimodal neural networks to interpret complex visual layouts and translate them into standardized text representations. The system operates as a remote inference orchestrator, offloading heavy document analysis tasks to external servers or cloud APIs to minimize local computational requirements. By employing a stateless worker architecture, it decouples document ingestion from inference, allowing for the distribu

    Python
    View on GitHub↗17,396
Compare all 30 related projects→