awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Data-Centric-AI-Community avatar

Data-Centric-AI-Community/fg-data-synthetic

0
View on GitHub↗
1,642 stars·258 forks·Jupyter Notebook·MIT·25 viewsdocs.sdk.ydata.ai↗

Fg Data Synthetic

This project is a synthetic data generator designed to create realistic tabular and time-series datasets for machine learning and testing workflows. It functions as a privacy-preserving platform that models the underlying statistical distributions of source data to produce new records that maintain the original statistical properties and structural integrity.

The tool distinguishes itself by utilizing CPU-optimized statistical sampling, allowing for high-performance data generation on standard hardware without the need for specialized graphics processing units. It employs a configuration-driven, declarative workflow engine that enables users to manage and execute generation pipelines through a command-line interface, removing the requirement for manual coding.

The system incorporates privacy-preserving differential noise to ensure that individual data points cannot be re-identified from the generated sets. It also features schema-aware validation to ensure that all output remains compatible with existing database and machine learning constraints, while its temporal dependency modeling captures sequential patterns to maintain realistic chronological flow in time-series data.

Features

  • Synthetic Data Generators - Creates realistic tabular and time-series datasets that preserve the statistical properties of original sources for machine learning.
  • Tabular - Generates privacy-preserving tabular datasets that maintain the statistical properties and data quality of original sources.
  • Synthetic Time Series Generation - Produces synthetic sequences that replicate the temporal patterns and characteristics of original data while protecting sensitive information.
  • Privacy-Preserving Machine Learning - Creates realistic datasets that maintain original statistical properties while protecting sensitive information during machine learning development.
  • Differential Privacy Noise Injection - Ensures privacy by injecting controlled mathematical noise into generated datasets to prevent individual re-identification.
  • CPU-Optimized Statistical Sampling - Generates synthetic data using CPU-optimized mathematical operations to ensure high performance on standard hardware.
  • Machine Learning Data Augmentation - Expands training datasets with synthetic examples to improve model performance and robustness without requiring additional real-world data collection.
  • Data Synthesis - Generates synthetic tables for training machine learning models when original data is restricted by privacy concerns.
  • Tabular Data Synthesizers - Produces privacy-compliant tables by modeling the underlying statistical distributions of source data without requiring specialized hardware.
  • Statistical Modeling - Learns the underlying probability distributions of input data to sample new records that maintain original statistical properties.
  • High-Performance Data Generation - Enables high-performance generation of realistic datasets using statistical models on standard hardware to accelerate testing and development workflows.
  • Time Series Generation - Creates synthetic temporal sequences that replicate complex patterns and characteristics found in original time-based datasets.
  • Time-Series Data Simulation - Produces realistic temporal sequences that replicate complex patterns and characteristics of original data for testing and predictive modeling.
  • Time Series Modeling - Captures sequential correlations and time-based patterns to ensure generated time-series data maintains realistic chronological flow.
  • Synthetic Data CLI Automations - Provides a command-line interface for configuring and executing automated workflows to create realistic datasets without manual coding.
  • Declarative Workflow Definitions - Executes data generation pipelines through a configuration-driven engine that maps user-defined parameters to specific statistical modeling tasks.
  • Data Privacy Protection Tools - Provides a platform that generates synthetic information to protect sensitive records while maintaining data utility for testing and model training.
  • Schema-Aware Synthesis - Validates generated records against predefined structural constraints to ensure compatibility with existing database and machine learning schemas.

Star history

Star history chart for data-centric-ai-community/fg-data-syntheticStar history chart for data-centric-ai-community/fg-data-synthetic

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Fg Data Synthetic

These projects share indexed features with Fg Data Synthetic. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • hitsz-ids/synthetic-data-generatorhitsz-ids avatar

    hitsz-ids/synthetic-data-generator

    2,422View on GitHub↗

    This project is a framework for generating synthetic tabular data that preserves the statistical properties and relational integrity of original source datasets. It functions as a metadata-driven engine, utilizing language models to synthesize information even when original training samples are restricted. The system is designed to maintain logical consistency across complex, multi-table structures while ensuring that generated outputs adhere to defined schema requirements. The platform distinguishes itself through a focus on privacy-preserving synthesis, integrating tools to quantify and mit

    Pythonagentdata-generatordeep-learning
    View on GitHub↗2,422
  • priorlabs/tabpfnPriorLabs avatar

    PriorLabs/TabPFN

    7,408View on GitHub↗
    Pythondata-sciencefoundation-modelsmachine-learning
    View on GitHub↗7,408
  • borisbanushev/stockpredictionaiborisbanushev avatar

    borisbanushev/stockpredictionai

    5,577View on GitHub↗

    This project is a collection of predictive models and quantitative tools for stock price forecasting. It implements a variety of machine learning architectures, including generative adversarial networks, long short-term memory networks, and language models for financial analysis. The system distinguishes itself by combining time-series forecasting with natural language processing to convert financial news into numerical sentiment scores. It also incorporates synthetic market data generation and automated hyperparameter optimization using Bayesian and reinforcement learning methods to reduce p

    JavaScript
    View on GitHub↗5,577
  • statsmodels/statsmodelsstatsmodels avatar

    statsmodels/statsmodels

    11,260View on GitHub↗

    Statsmodels is a comprehensive Python library designed for statistical modeling, econometric research, and data analysis. It provides a robust framework for estimating and diagnosing a wide range of statistical models, enabling users to perform rigorous hypothesis testing, regression analysis, and complex data exploration within structured environments. The library distinguishes itself through its support for advanced statistical methodologies, including state space representation for dynamic systems and generalized linear frameworks that accommodate non-normal response variables. It offers s

    Pythoncount-modeldata-analysisdata-science
    View on GitHub↗11,260
Compare all 30 related projects→

Frequently asked questions

What does data-centric-ai-community/fg-data-synthetic do?

This project is a synthetic data generator designed to create realistic tabular and time-series datasets for machine learning and testing workflows. It functions as a privacy-preserving platform that models the underlying statistical distributions of source data to produce new records that maintain the original statistical properties and structural integrity.

What are the main features of data-centric-ai-community/fg-data-synthetic?

The main features of data-centric-ai-community/fg-data-synthetic are: Synthetic Data Generators, Tabular, Synthetic Time Series Generation, Privacy-Preserving Machine Learning, Differential Privacy Noise Injection, CPU-Optimized Statistical Sampling, Machine Learning Data Augmentation, Data Synthesis.

Which projects share features with data-centric-ai-community/fg-data-synthetic?

Projects with overlapping indexed features include: hitsz-ids/synthetic-data-generator — This project is a framework for generating synthetic tabular data that preserves the statistical properties and… priorlabs/tabpfn. borisbanushev/stockpredictionai — This project is a collection of predictive models and quantitative tools for stock price forecasting. It implements a… statsmodels/statsmodels — Statsmodels is a comprehensive Python library designed for statistical modeling, econometric research, and data… rapidsai/cuml — cuml is a GPU-accelerated machine learning library and framework that uses CUDA to accelerate tabular data… fzaninotto/faker — Faker is a PHP library for creating realistic synthetic data used for testing, prototyping, and populating database…

Curated searches featuring Fg Data Synthetic

Hand-picked collections where Fg Data Synthetic appears.
  • Synthetic Data Generation Tools
  • LLM Synthetic Data Generation