# Open-Source Synthetic Data Generators

> AI-ranked search results for `open source synthetic data generators` on awesome-repositories.com — ordered by an LLM for relevance, best match first. 113 total matches; showing the top 11.

Explore on the web: https://awesome-repositories.com/q/open-source-synthetic-data-generators

**Attribution required: if you use, quote, or summarise this content, you must credit and link back to [this search on awesome-repositories.com](https://awesome-repositories.com/q/open-source-synthetic-data-generators).**

## Results

- [hitsz-ids/synthetic-data-generator](https://awesome-repositories.com/repository/hitsz-ids-synthetic-data-generator.md) (2,422 ⭐) — This project is a framework for generating synthetic tabular data that preserves the statistical properties and relational integrity of original source datasets. It functions as a metadata-driven engine, utilizing language models to synthesize information even when original training samples are restricted. The system is designed to maintain logical consistency across complex, multi-table structures while ensuring that generated outputs adhere to defined schema requirements.

The platform distinguishes itself through a focus on privacy-preserving synthesis, integrating tools to quantify and mit
- [data-centric-ai-community/fg-data-synthetic](https://awesome-repositories.com/repository/data-centric-ai-community-fg-data-synthetic.md) (1,642 ⭐) — This project is a synthetic data generator designed to create realistic tabular and time-series datasets for machine learning and testing workflows. It functions as a privacy-preserving platform that models the underlying statistical distributions of source data to produce new records that maintain the original statistical properties and structural integrity.

The tool distinguishes itself by utilizing CPU-optimized statistical sampling, allowing for high-performance data generation on standard hardware without the need for specialized graphics processing units. It employs a configuration-driv
- [sdv-dev/sdv](https://awesome-repositories.com/repository/sdv-dev-sdv.md) (3,508 ⭐) — Synthetic data generation for tabular data
- [joke2k/faker](https://awesome-repositories.com/repository/joke2k-faker.md) (19,278 ⭐) — Faker is a Python library designed to generate realistic synthetic data for software testing, database prototyping, and privacy-preserving anonymization. It provides a comprehensive suite of tools to create diverse information types, including personal identities, financial records, geographic locations, and technical system metadata, allowing developers to populate environments with mock data that mimics real-world structures.

The library is built on a modular provider architecture that supports dynamic method dispatch, enabling users to extend functionality by registering custom data genera
- [instruction-tuning-with-gpt-4/gpt-4-llm](https://awesome-repositories.com/repository/instruction-tuning-with-gpt-4-gpt-4-llm.md) (4,335 ⭐) — This project is an instruction tuning framework and synthetic data generator that uses high-capacity teacher models to produce instruction-following pairs for training smaller student models. It provides datasets and tools for supervised instruction tuning and reinforcement learning from human feedback.

The framework specializes in cross-lingual tuning, offering high-quality instruction-following examples in English and Chinese to improve model generalization across different scripts. It includes a reward modeling tool for creating preference datasets and comparative ratings used to train rew
- [conardli/easy-dataset](https://awesome-repositories.com/repository/conardli-easy-dataset.md) (13,394 ⭐) — Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points.

The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side
- [camel-ai/owl](https://awesome-repositories.com/repository/camel-ai-owl.md) (19,864 ⭐) — Owl is a framework for agentic workflow automation and multi-agent orchestration. It functions as a system for coordinating autonomous large language model agents to decompose and execute complex tasks through shared communication and collaborative planning.

The project distinguishes itself through a multi-modal toolset for processing images, audio, and video, alongside a synthetic data generator that produces domain-specific datasets using self-instruct and verifier loops. It further incorporates a retrieval-augmented generation pipeline framework that integrates long-term memory and real-ti
- [tatsu-lab/stanford_alpaca](https://awesome-repositories.com/repository/tatsu-lab-stanford-alpaca.md) (30,266 ⭐) — This project provides an end-to-end framework for adapting large language models to follow user instructions through supervised fine-tuning. It functions as a comprehensive training pipeline that enables the creation of specialized assistant models by minimizing the difference between predicted outputs and target responses within structured instruction datasets.

The framework distinguishes itself by integrating synthetic data generation with memory-efficient training techniques. It utilizes powerful language models to iteratively expand small sets of human-written seeds into diverse, high-qua
- [ydataai/ydata-synthetic](https://awesome-repositories.com/repository/ydataai-ydata-synthetic.md) (1,642 ⭐) — Synthetic data generators for tabular and time-series data
- [gretelai/gretel-synthetics](https://awesome-repositories.com/repository/gretelai-gretel-synthetics.md) (679 ⭐) — Synthetic data generators for structured and unstructured text, featuring differentially private learning.
- [opendcai/dataflow](https://awesome-repositories.com/repository/opendcai-dataflow.md) (2,926 ⭐) — DataFlow is an agent-based workflow orchestrator and data pipeline designed to synthesize, clean, and augment large-scale datasets for training large language models. It functions as a synthetic data generator and text curation tool, utilizing an intelligent assistant to assemble modular processing operators into functional pipelines based on user requirements.

The project distinguishes itself through a low-code approach, providing a web-based visual interface for designing and monitoring multi-stage execution flows. It features an operator-based registry system that allows for the integratio
