awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
hitsz-ids avatar

hitsz-ids/synthetic-data-generator

0
View on GitHub↗
2,422 stars·391 forks·Python·Apache-2.0·30 views

Synthetic Data Generator

This project is a framework for generating synthetic tabular data that preserves the statistical properties and relational integrity of original source datasets. It functions as a metadata-driven engine, utilizing language models to synthesize information even when original training samples are restricted. The system is designed to maintain logical consistency across complex, multi-table structures while ensuring that generated outputs adhere to defined schema requirements.

The platform distinguishes itself through a focus on privacy-preserving synthesis, integrating tools to quantify and mitigate re-identification risks through differential privacy and anonymization techniques. It supports modular extensibility, allowing for the integration of custom generation models and data connectors. Furthermore, the framework includes automated validation routines that compare the distribution and correlation patterns of synthetic outputs against source data to verify statistical fidelity.

Beyond core generation, the system provides capabilities for data enrichment and feature engineering by deriving new columns from learned patterns. It incorporates operational oversight tools to monitor resource utilization and processing efficiency during high-volume tasks. The library is designed to handle large-scale datasets through memory-efficient stream processing and iterative batching to ensure stability.

Features

  • Tabular - Creates privacy-preserving synthetic tabular records that maintain the statistical properties of original source data.
  • Data Synthesis - Leverages language models and metadata to synthesize information when original training samples are restricted.
  • Synthetic Data Generators - Creates privacy-preserving structured datasets that maintain the statistical properties of original source data.
  • Synthetic Data Generators - Generates synthetic datasets using metadata and language model knowledge without requiring original training samples.
  • Relational Synthesis - Generates consistent synthetic datasets across multiple related tables while preserving structural integrity and logical relationships.
  • Statistical Fidelity Evaluators - Validates synthetic data quality by comparing distribution and correlation patterns against original source datasets.
  • Privacy and Data Protection - Protects sensitive information using differential privacy and masking techniques to ensure compliance.
  • Differential Privacy Aggregators - Provides a toolkit for quantifying and mitigating re-identification risks in synthetic datasets.
  • Differential Privacy Noise Injection - Injects mathematical noise into synthetic data generation to ensure individual records cannot be re-identified.
  • Schema-Driven Generators - Constructs synthetic datasets by interpreting structured schema definitions rather than relying on raw input samples.
  • Relational Data Generators - Generates consistent synthetic data across multiple related tables while preserving structural integrity.
  • Synthetic Data Generation - Enriches datasets by generating new columns based on learned patterns to improve data depth.
  • Data Enrichment - Enriches datasets by deriving new columns from learned patterns and external knowledge.
  • Security & Privacy - Identifies personal identifiers and applies privacy techniques to secure generated information.
  • Memory-Efficient Data Streaming - Processes large-scale datasets in memory-efficient chunks to maintain system stability during high-volume generation.
  • Graph-Relational Models - Maintains logical consistency across multi-table structures by enforcing structural dependencies during data generation.
  • Relational Data Modeling - Models complex multi-table relationships to ensure logical consistency during synthetic data generation.
  • Generation Framework Extensions - Allows integration of custom generation models, data processing routines, and external connectors through a modular plugin architecture.

Star history

Star history chart for hitsz-ids/synthetic-data-generatorStar history chart for hitsz-ids/synthetic-data-generator

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Synthetic Data Generator

These projects share indexed features with Synthetic Data Generator. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • data-centric-ai-community/fg-data-syntheticData-Centric-AI-Community avatar

    Data-Centric-AI-Community/fg-data-synthetic

    1,642View on GitHub↗

    This project is a synthetic data generator designed to create realistic tabular and time-series datasets for machine learning and testing workflows. It functions as a privacy-preserving platform that models the underlying statistical distributions of source data to produce new records that maintain the original statistical properties and structural integrity. The tool distinguishes itself by utilizing CPU-optimized statistical sampling, allowing for high-performance data generation on standard hardware without the need for specialized graphics processing units. It employs a configuration-driv

    Jupyter Notebookdatagenerationdatageneratordeep-learning
    View on GitHub↗1,642
  • priorlabs/tabpfnPriorLabs avatar

    PriorLabs/TabPFN

    7,408View on GitHub↗
    Pythondata-sciencefoundation-modelsmachine-learning
    View on GitHub↗7,408
  • wiseodd/generative-modelswiseodd avatar

    wiseodd/generative-models

    7,497View on GitHub↗

    This is a generative AI model library containing a collection of PyTorch and TensorFlow implementations for creating synthetic data and modeling complex probability distributions. It serves as a multi-framework repository of deep learning models designed for learning and replicating data patterns. The project provides specialized implementation suites for several generative architectures. This includes Generative Adversarial Networks using competing generator and discriminator models, Variational Autoencoder frameworks that map data to a latent space, and Restricted Boltzmann Machine and Deep

    Python
    View on GitHub↗7,497
  • conardli/easy-datasetConardLi avatar

    ConardLi/easy-dataset

    13,394View on GitHub↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    JavaScriptdatasetfine-tuningjavascript
    View on GitHub↗13,394
Compare all 30 related projects→

Frequently asked questions

What does hitsz-ids/synthetic-data-generator do?

This project is a framework for generating synthetic tabular data that preserves the statistical properties and relational integrity of original source datasets. It functions as a metadata-driven engine, utilizing language models to synthesize information even when original training samples are restricted. The system is designed to maintain logical consistency across complex, multi-table structures while ensuring that generated outputs adhere to defined schema requirements.

What are the main features of hitsz-ids/synthetic-data-generator?

The main features of hitsz-ids/synthetic-data-generator are: Tabular, Data Synthesis, Synthetic Data Generators, Relational Synthesis, Statistical Fidelity Evaluators, Privacy and Data Protection, Differential Privacy Aggregators, Differential Privacy Noise Injection.

Which projects share features with hitsz-ids/synthetic-data-generator?

Projects with overlapping indexed features include: data-centric-ai-community/fg-data-synthetic — This project is a synthetic data generator designed to create realistic tabular and time-series datasets for machine… wiseodd/generative-models — This is a generative AI model library containing a collection of PyTorch and TensorFlow implementations for creating… priorlabs/tabpfn. conardli/easy-dataset — Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets,… limix-ldm-ai/limix — LimiX is a tabular foundation model and a suite of tools for structured data, providing a transformer-based system for… camel-ai/camel — This project is a comprehensive framework for building and managing autonomous agent systems. It provides a unified…

Curated searches featuring Synthetic Data Generator

Hand-picked collections where Synthetic Data Generator appears.
  • Synthetic Data Generation Tools
  • LLM Synthetic Data Generation
  • Synthetic Database Data Generators