awesome-repositories.com
Blog
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Open-Source Synthetic Data Generators

Classement mis à jour le 13 juil. 2026

For an open source tool for generating data, the strongest matches are hitsz-ids/synthetic-data-generator (This framework is a dedicated engine for generating synthetic), data-centric-ai-community/fg-data-synthetic (This tool is a dedicated synthetic data generator that) and sdv-dev/sdv (This library provides a comprehensive suite of tools for). joke2k/faker and instruction-tuning-with-gpt-4/gpt-4-llm round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Nous sélectionnons les dépôts GitHub open-source correspondant à « open source synthetic data generators ». Les résultats sont classés par pertinence par rapport à votre recherche — utilisez les filtres ci-dessous pour affiner, ou utilisez l'IA.

Open-Source Synthetic Data Generators

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • hitsz-ids/synthetic-data-generatorAvatar de hitsz-ids

    hitsz-ids/synthetic-data-generator

    2,422Voir sur GitHub↗

    This project is a framework for generating synthetic tabular data that preserves the statistical properties and relational integrity of original source datasets. It functions as a metadata-driven engine, utilizing language models to synthesize information even when original training samples are restricted. The system is designed to maintain logical consistency across complex, multi-table structures while ensuring that generated outputs adhere to defined schema requirements. The platform distinguishes itself through a focus on privacy-preserving synthesis, integrating tools to quantify and mit

    This framework is a dedicated engine for generating synthetic tabular data that explicitly incorporates privacy-preserving techniques like differential privacy and supports complex relational structures for machine learning workflows.

    PythonData SynthesisRelational SynthesisDifferential Privacy Aggregators
    Voir sur GitHub↗2,422
  • data-centric-ai-community/fg-data-syntheticAvatar de Data-Centric-AI-Community

    Data-Centric-AI-Community/fg-data-synthetic

    1,642Voir sur GitHub↗

    This project is a synthetic data generator designed to create realistic tabular and time-series datasets for machine learning and testing workflows. It functions as a privacy-preserving platform that models the underlying statistical distributions of source data to produce new records that maintain the original statistical properties and structural integrity. The tool distinguishes itself by utilizing CPU-optimized statistical sampling, allowing for high-performance data generation on standard hardware without the need for specialized graphics processing units. It employs a configuration-driv

    This tool is a dedicated synthetic data generator that supports tabular and time-series data, incorporates privacy-preserving techniques, and provides the necessary interfaces for integration into machine learning workflows.

    Jupyter NotebookData SynthesisDifferential Privacy Noise Injection
    Voir sur GitHub↗1,642
  • sdv-dev/sdvAvatar de sdv-dev

    sdv-dev/SDV

    3,508Voir sur GitHub↗

    Synthetic data generation for tabular data

    This library provides a comprehensive suite of tools for modeling the statistical distributions of tabular data and generating synthetic versions, making it a direct fit for training machine learning models while preserving privacy.

    PythonData Annotation and SynthesisData Loading Extraction
    Voir sur GitHub↗3,508
  • joke2k/fakerAvatar de joke2k

    joke2k/faker

    19,278Voir sur GitHub↗

    Faker is a Python library designed to generate realistic synthetic data for software testing, database prototyping, and privacy-preserving anonymization. It provides a comprehensive suite of tools to create diverse information types, including personal identities, financial records, geographic locations, and technical system metadata, allowing developers to populate environments with mock data that mimics real-world structures. The library is built on a modular provider architecture that supports dynamic method dispatch, enabling users to extend functionality by registering custom data genera

    This library provides a robust framework for generating realistic synthetic tabular data and personal information, making it a practical tool for creating mock datasets for testing and privacy-focused prototyping.

    PythonAnonymization Services
    Voir sur GitHub↗19,278
  • instruction-tuning-with-gpt-4/gpt-4-llmAvatar de Instruction-Tuning-with-GPT-4

    Instruction-Tuning-with-GPT-4/GPT-4-LLM

    4,335Voir sur GitHub↗

    This project is an instruction tuning framework and synthetic data generator that uses high-capacity teacher models to produce instruction-following pairs for training smaller student models. It provides datasets and tools for supervised instruction tuning and reinforcement learning from human feedback. The framework specializes in cross-lingual tuning, offering high-quality instruction-following examples in English and Chinese to improve model generalization across different scripts. It includes a reward modeling tool for creating preference datasets and comparative ratings used to train rew

    This project functions as a synthetic data generator specifically for instruction-tuning LLMs by using teacher models to create training pairs, fitting the category for model-based data augmentation.

    HTMLInstruction Tuning FrameworksSupervised Instruction Fine-TuningCross-Lingual Instruction Datasets
    Voir sur GitHub↗4,335
  • conardli/easy-datasetAvatar de ConardLi

    ConardLi/easy-dataset

    13,394Voir sur GitHub↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    This platform provides automated synthetic data generation and augmentation workflows specifically for fine-tuning machine learning models, fitting the category despite its broader focus on the full dataset lifecycle.

    JavaScriptAI Model BenchmarkingModel Evaluation SuitesSynthetic Data Generation
    Voir sur GitHub↗13,394
  • camel-ai/owlAvatar de camel-ai

    camel-ai/owl

    19,864Voir sur GitHub↗

    Owl is a framework for agentic workflow automation and multi-agent orchestration. It functions as a system for coordinating autonomous large language model agents to decompose and execute complex tasks through shared communication and collaborative planning. The project distinguishes itself through a multi-modal toolset for processing images, audio, and video, alongside a synthetic data generator that produces domain-specific datasets using self-instruct and verifier loops. It further incorporates a retrieval-augmented generation pipeline framework that integrates long-term memory and real-ti

    While primarily an agent orchestration framework, this tool includes a synthetic data generator that utilizes self-instruct and verifier loops to produce domain-specific datasets, fitting the category's core capability for programmatic data generation.

    PythonHierarchical Agent OrchestrationMulti-Agent Orchestration SystemsAgentic Workflow Automation
    Voir sur GitHub↗19,864
  • tatsu-lab/stanford_alpacaAvatar de tatsu-lab

    tatsu-lab/stanford_alpaca

    30,266Voir sur GitHub↗

    This project provides an end-to-end framework for adapting large language models to follow user instructions through supervised fine-tuning. It functions as a comprehensive training pipeline that enables the creation of specialized assistant models by minimizing the difference between predicted outputs and target responses within structured instruction datasets. The framework distinguishes itself by integrating synthetic data generation with memory-efficient training techniques. It utilizes powerful language models to iteratively expand small sets of human-written seeds into diverse, high-qua

    This framework provides tools for generating synthetic instruction datasets using large language models to fine-tune other models, fitting the category of synthetic data generation for machine learning pipelines.

    PythonInstruction Fine-Tuning FrameworksInstruction TuningInstruction Tuning Frameworks
    Voir sur GitHub↗30,266
  • ydataai/ydata-syntheticAvatar de ydataai

    ydataai/ydata-synthetic

    1,642Voir sur GitHub↗

    Synthetic data generators for tabular and time-series data

    This tool provides synthetic data generation for tabular and time-series datasets using generative models, making it a direct fit for creating privacy-preserving training data.

    Jupyter NotebookData Annotation and SynthesisData Centric AI
    Voir sur GitHub↗1,642
  • gretelai/gretel-syntheticsAvatar de gretelai

    gretelai/gretel-synthetics

    679Voir sur GitHub↗

    Synthetic data generators for structured and unstructured text, featuring differentially private learning.

    This tool provides a Python-based library for generating synthetic tabular and text data using differentially private learning, making it a direct fit for privacy-preserving dataset generation.

    PythonData Annotation and Synthesis
    Voir sur GitHub↗679
  • opendcai/dataflowAvatar de OpenDCAI

    OpenDCAI/DataFlow

    2,926Voir sur GitHub↗

    DataFlow is an agent-based workflow orchestrator and data pipeline designed to synthesize, clean, and augment large-scale datasets for training large language models. It functions as a synthetic data generator and text curation tool, utilizing an intelligent assistant to assemble modular processing operators into functional pipelines based on user requirements. The project distinguishes itself through a low-code approach, providing a web-based visual interface for designing and monitoring multi-stage execution flows. It features an operator-based registry system that allows for the integratio

    DataFlow is a synthetic data generator that uses agent-based workflows to synthesize and augment datasets for machine learning, specifically targeting LLM training pipelines.

    PythonTraining Data GenerationTraining DatasetsAgent-Based Pipeline Assembly
    Voir sur GitHub↗2,926
Comparez le top 10 en un coup d'œil
DépôtStarsLangageLicenceDernier push
hitsz-ids/synthetic-data-generator2.4KPythonApache-2.025 mai 2026
data-centric-ai-community/fg-data-synthetic1.6KJupyter NotebookMIT23 avr. 2026
sdv-dev/sdv3.5KPythonNOASSERTION15 juin 2026
joke2k/faker19.3KPythonMIT10 juin 2026
instruction-tuning-with-gpt-4/gpt-4-llm4.3KHTMLApache-2.011 juin 2023
conardli/easy-dataset13.4KJavaScriptother20 févr. 2026
camel-ai/owl19.9KPython—12 juin 2026
tatsu-lab/stanford_alpaca30.3KPythonapache-2.017 juil. 2024
ydataai/ydata-synthetic1.6KJupyter NotebookMIT23 avr. 2026
gretelai/gretel-synthetics679PythonNOASSERTION24 juin 2025

Related searches

  • outil de génération de datasets synthétiques
  • générer des données synthétiques réalistes
  • générer des données de test pour ma base de données
  • an open source tool for data visualization
  • an open source model for code generation
  • toolkit pour le nettoyage et la curation de datasets
  • une plateforme open source pour les catalogues de données
  • toolkit pour la génération de musique par IA