awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
argilla-io avatar

argilla-io/distilabel

0
View on GitHub↗
3,277 stars·243 forks·Python·Apache-2.0·13 viewsdistilabel.argilla.io↗

Distilabel

Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

Features

  • Data Generation Frameworks - Generates and augments data for supervised fine-tuning and preference optimization.
  • Data Processing - Pipeline framework for synthetic data generation and AI feedback.
  • Data Processing Tools - Pipeline framework for synthetic data generation and AI feedback.
  • Feature Engineering - Generates synthetic data and AI feedback for model training.

Star history

Star history chart for argilla-io/distilabelStar history chart for argilla-io/distilabel

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does argilla-io/distilabel do?

Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

What are the main features of argilla-io/distilabel?

The main features of argilla-io/distilabel are: Data Generation Frameworks, Data Processing, Data Processing Tools, Feature Engineering.

Which projects share features with argilla-io/distilabel?

Projects with overlapping indexed features include: jpmens/jo — Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell… catchthetornado/pdf-extract-api. 599yongyang/datasetloom. allenai/olmocr — Olmocr is a distributed document processing framework designed to convert PDF and image files into structured… bytedance/dolphin — Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital… bespokelabsai/curator.

Projects sharing features with Distilabel

These projects share indexed features with Distilabel. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • jpmens/jojpmens avatar

    jpmens/jo

    4,868View on GitHub↗

    Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell arguments and standard input. It functions as a data processing tool that transforms raw input into structured formats, enabling the generation of complex payloads for APIs, configuration files, and automated data pipelines. The tool distinguishes itself through its ability to resolve hierarchical data structures using delimiter-based path definitions and its integrated type-inference engine, which automatically casts input values into native boolean, numeric, or null types. Users can

    C
    View on GitHub↗4,868
  • bespokelabsai/curatorbespokelabsai avatar

    bespokelabsai/curator

    1,637View on GitHub↗
    Pythonagentsdeep-learningfine-tuning
    View on GitHub↗1,637
  • 599yongyang/datasetloom5

    599yongyang/DatasetLoom

    0View on GitHub↗
    View on GitHub↗0
  • allenai/olmocrallenai avatar

    allenai/olmocr

    17,396View on GitHub↗

    Olmocr is a distributed document processing framework designed to convert PDF and image files into structured markdown. It functions as a vision-based document parser that utilizes multimodal neural networks to interpret complex visual layouts and translate them into standardized text representations. The system operates as a remote inference orchestrator, offloading heavy document analysis tasks to external servers or cloud APIs to minimize local computational requirements. By employing a stateless worker architecture, it decouples document ingestion from inference, allowing for the distribu

    Python
    View on GitHub↗17,396
Compare all 30 related projects→