awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Génération de données synthétiques pour LLM

Classement mis à jour le 30 juin 2026

For outil de génération de datasets synthétiques, the strongest matches are conardli/easy-dataset (Easy-dataset is an end-to-end platform for managing ML datasets), tatsu-lab/stanford_alpaca (Stanford Alpaca is an end-to-end framework that uses LLMs) and meta-llama/synthetic-data-kit (This tool from Meta is purpose-built for generating high-quality). camel-ai/owl and opendcai/dataflow round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Frameworks et outils pour créer des jeux de données synthétiques de haute qualité afin d'entraîner et de fine-tuner les grands modèles de langage.

Génération de données synthétiques pour LLM

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • conardli/easy-datasetAvatar de ConardLi

    ConardLi/easy-dataset

    13,394Voir sur GitHub↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    Easy-dataset is an end-to-end platform for managing ML datasets that explicitly covers automated synthetic data generation with LLM integration, quality filtering, and batch processing, directly matching the search for a synthetic data generation framework.

    JavaScriptSynthetic Data PipelinesSynthetic Data Generators
    Voir sur GitHub↗13,394
  • tatsu-lab/stanford_alpacaAvatar de tatsu-lab

    tatsu-lab/stanford_alpaca

    30,266Voir sur GitHub↗

    This project provides an end-to-end framework for adapting large language models to follow user instructions through supervised fine-tuning. It functions as a comprehensive training pipeline that enables the creation of specialized assistant models by minimizing the difference between predicted outputs and target responses within structured instruction datasets. The framework distinguishes itself by integrating synthetic data generation with memory-efficient training techniques. It utilizes powerful language models to iteratively expand small sets of human-written seeds into diverse, high-qua

    Stanford Alpaca is an end-to-end framework that uses LLMs to iteratively expand small seed sets into diverse synthetic instruction-following datasets, making it a clear fit for generating synthetic training data with LLMs.

    PythonSynthetic Data Generators
    Voir sur GitHub↗30,266
  • meta-llama/synthetic-data-kitAvatar de meta-llama

    meta-llama/synthetic-data-kit

    1,602Voir sur GitHub↗

    The synthetic data kit is an integrated framework designed to generate, curate, and format training datasets for language models. It provides an end-to-end pipeline that transforms raw source documents into structured data suitable for fine-tuning, reasoning, and tool-use model training. The framework distinguishes itself through a modular orchestration engine that manages the entire lifecycle of data preparation. It supports multimodal input by extracting both text and image content from various file formats, while employing context-aware chunking to maintain semantic coherence. The generati

    This tool from Meta is purpose-built for generating high-quality synthetic datasets using LLMs, fitting the requirement for an open-source framework that integrates LLM APIs and supports prompt-driven generation.

    PythonSynthetic Data Generators
    Voir sur GitHub↗1,602
  • camel-ai/owlAvatar de camel-ai

    camel-ai/owl

    19,864Voir sur GitHub↗

    Owl is a framework for agentic workflow automation and multi-agent orchestration. It functions as a system for coordinating autonomous large language model agents to decompose and execute complex tasks through shared communication and collaborative planning. The project distinguishes itself through a multi-modal toolset for processing images, audio, and video, alongside a synthetic data generator that produces domain-specific datasets using self-instruct and verifier loops. It further incorporates a retrieval-augmented generation pipeline framework that integrates long-term memory and real-ti

    Owl is an agent orchestration framework that includes a synthetic data generator using LLM self-instruct and verifier loops, so it can produce domain-specific datasets — fitting the core need for synthetic data generation, though its primary focus is broader.

    PythonSynthetic Data Generators
    Voir sur GitHub↗19,864
  • opendcai/dataflowAvatar de OpenDCAI

    OpenDCAI/DataFlow

    2,926Voir sur GitHub↗

    DataFlow is an agent-based workflow orchestrator and data pipeline designed to synthesize, clean, and augment large-scale datasets for training large language models. It functions as a synthetic data generator and text curation tool, utilizing an intelligent assistant to assemble modular processing operators into functional pipelines based on user requirements. The project distinguishes itself through a low-code approach, providing a web-based visual interface for designing and monitoring multi-stage execution flows. It features an operator-based registry system that allows for the integratio

    DataFlow is an agent-based workflow orchestrator for synthesizing, cleaning, and augmenting large-scale LLM training datasets, with LLM backend integration (vLLM, SGLang) and a low-code visual pipeline designer, directly addressing your need for a synthetic data generation framework.

    PythonText Quality Filtering
    Voir sur GitHub↗2,926
  • datajuicer/data-juicerAvatar de datajuicer

    datajuicer/data-juicer

    6,574Voir sur GitHub↗

    Data-Juicer is an open-source framework for cleaning, filtering, deduplicating, and transforming multimodal datasets to prepare them for training large language and vision models. It functions as a distributed data pipeline engine that runs processing jobs across Ray clusters, handling billions of samples with automatic operator fusion and adaptive parallelism. The framework provides a library of operators that leverage large language models for semantic extraction, filtering, and data synthesis within processing pipelines. The project distinguishes itself through a YAML-based data recipe sys

    Data-Juicer is a pipeline framework that includes LLM-powered operators for data synthesis and filtering, making it a valid tool for generating synthetic training data, though its primary focus is on broader multimodal data processing rather than dedicated generation.

    PythonText Quality Filtering
    Voir sur GitHub↗6,574
  • argilla-io/distilabelAvatar de argilla-io

    argilla-io/distilabel

    3,277Voir sur GitHub↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Distilabel is a framework for building scalable synthetic data and AI feedback pipelines using LLMs, directly matching the search for an open-source tool that generates training data with LLM integration and batch processing.

    PythonData Generation FrameworksData ProcessingData Processing Tools
    Voir sur GitHub↗3,277
  • bespokelabsai/curatorAvatar de bespokelabsai

    bespokelabsai/curator

    1,637Voir sur GitHub↗

    Curator is a Python framework for generating synthetic datasets using LLMs, fitting the search for a synthetic data generation tool with LLM integration.

    PythonData Generation FrameworksData Processing
    Voir sur GitHub↗1,637

Related searches

  • an open source tool for generating data
  • générer des données synthétiques réalistes
  • toolkit pour le red-teaming de modèles de langage
  • un framework pour évaluer la qualité des sorties LLM
  • framework pour l'évaluation d'applications LLM
  • toolkit pour créer des agents IA capables d'utiliser des outils
  • a platform for building generative AI applications
  • un framework pour le fine-tuning de grands modèles de langage