3 Repos
Using rules and templates to generate massive quantities of synthetic training data.
Distinct from Large-Scale Data Computation: Candidates focus on distributed computation or storage; this is specifically about the generative process at scale.
Explore 3 awesome GitHub repositories matching artificial intelligence & ml · Large-Scale Synthetic Data Generation. Refine with filters or upvote what's useful.
Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo
Generates synthetic examples using rules, templates, and constraints to produce large-scale datasets.
Dieses Projekt ist eine kuratierte Sammlung chinesischer Namen, Nachnamen und Verwandtschaftsbegriffe, die für linguistische Analysen und Natural Language Processing konzipiert wurde. Es fungiert als mehrsprachiger Namensdatensatz und Trainingsressource für Named Entity Recognition und bietet ein einheitliches Repository von Namen in chinesischer, japanischer und englischer Sprache. Das Projekt enthält einen synthetischen Namensgenerator, der realistische Personennamen durch Anwendung analysierter Namensmuster und demografischer Daten erstellt. Es bietet zudem ein bereinigtes Lexikon chinesischer Idiome, das aus mehreren Quellen zusammengetragen und dedupliziert wurde. Die verfügbaren Daten unterstützen eine Vielzahl von NLP-Aufgaben, einschließlich Wortsegmentierung und Text-Preprocessing. Das Korpus enthält kategorisierte Register geschlechtsspezifischer Namen, Familiennamen und Verwandtschaftstitel, um bei der Entitäts-Tagging- und linguistischen Forschung zu unterstützen.
Creates realistic synthetic names using structural rules and statistical distributions derived from a large corpus.
ruvector is a Rust-based vector store and graph database designed for local inference and nearest neighbor searches. It utilizes a vector graph database architecture and a graph neural network index to refine search rankings through structural attention. The system includes a hardware-accelerated quantum circuit simulator for executing state-vector simulations and complex search patterns, alongside a WebAssembly inference engine for running vector search and model execution directly in web browsers. The project employs a cognitive container format that bundles models, data, and a bootable mic
Creates large-scale structured time-series or embedding datasets through programmatic optimization.