awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
vanderschaarlab avatar

vanderschaarlab/synthcity

0
View on GitHub↗
665 stars·94 forks·Python·Apache-2.0·9 viewswww.vanderschaar-lab.com↗

Synthcity

A library for generating and evaluating synthetic tabular data for privacy, fairness and data augmentation.

Features

  • Data Annotation and Synthesis - Library for generating and evaluating synthetic tabular data.

Star history

Star history chart for vanderschaarlab/synthcityStar history chart for vanderschaarlab/synthcity

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Synthcity

These projects share indexed features with Synthcity. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • cleanlab/cleanlabcleanlab avatar

    cleanlab/cleanlab

    11,513View on GitHub↗

    Cleanlab is a data-centric AI library and toolkit designed to improve machine learning model performance by detecting label errors and increasing overall dataset quality. It implements a confident learning framework that iteratively refines label noise estimates by comparing model predictions with estimated label probabilities to identify mislabeled examples. The project provides specialized utilities for active learning optimization, allowing for the selection of the most impactful examples for labeling or re-labeling. It also includes an outlier detection tool to identify atypical data poin

    Pythonactive-learningannotationanomaly-detection
    View on GitHub↗11,513
  • code-kern-ai/refinerycode-kern-ai avatar

    code-kern-ai/refinery

    1,470View on GitHub↗

    The data scientist's open-source choice to scale, assess and maintain natural language data. Treat training data like a software artifact.

    Python
    View on GitHub↗1,470
  • cvat-ai/cvatcvat-ai avatar

    cvat-ai/cvat

    15,317View on GitHub↗

    CVAT is an open-source, web-based platform designed for annotating images, videos, and 3D point clouds to create high-quality training datasets for machine learning. It functions as a containerized server that orchestrates the entire lifecycle of computer vision data, from initial task creation and manual labeling to quality assurance and final dataset export. The platform distinguishes itself through deep integration with machine learning models, allowing users to deploy custom AI models as serverless functions for automated object detection, tracking, and skeleton annotation. It supports co

    Pythonannotationannotation-toolannotations
    View on GitHub↗15,317
  • argilla-io/argillaargilla-io avatar

    argilla-io/argilla

    5,015View on GitHub↗

    Argilla is a collaborative AI feedback tool and data curation management system. It serves as a human-in-the-loop dataset platform designed to coordinate workforce annotators and domain experts in labeling, rating, and refining data samples for machine learning projects. The platform focuses on large language model dataset curation and reinforcement learning from human feedback workflows. It provides a shared workspace for integrating human expertise into AI development to validate model outputs and correct data errors. The system manages the end-to-end machine learning data pipeline, includ

    Python
    View on GitHub↗5,015
Compare all 15 related projects→

Frequently asked questions

What does vanderschaarlab/synthcity do?

A library for generating and evaluating synthetic tabular data for privacy, fairness and data augmentation.

What are the main features of vanderschaarlab/synthcity?

The main features of vanderschaarlab/synthcity are: Data Annotation and Synthesis.

Which projects share features with vanderschaarlab/synthcity?

Projects with overlapping indexed features include: cleanlab/cleanlab — Cleanlab is a data-centric AI library and toolkit designed to improve machine learning model performance by detecting… code-kern-ai/refinery — The data scientist's open-source choice to scale, assess and maintain natural language data. Treat training data like… cvat-ai/cvat — CVAT is an open-source, web-based platform designed for annotating images, videos, and 3D point clouds to create… diyago/tabular-data-generation — We well know GANs for success in the realistic image generation. However, they can be applied in tabular data… doccano/doccano — Doccano is a collaborative data labeling platform and machine learning dataset management system. It provides a… argilla-io/argilla — Argilla is a collaborative AI feedback tool and data curation management system. It serves as a human-in-the-loop…