awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Generare de date sintetice pentru LLM-uri

Clasament actualizat la 30 iun. 2026

For instrument pentru generarea de seturi de date sintetice, the strongest matches are conardli/easy-dataset (Easy-dataset is an end-to-end platform for managing ML datasets), tatsu-lab/stanford_alpaca (Stanford Alpaca is an end-to-end framework that uses LLMs) and meta-llama/synthetic-data-kit (This tool from Meta is purpose-built for generating high-quality). camel-ai/owl and opendcai/dataflow round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Framework-uri și instrumente pentru crearea de seturi de date sintetice de înaltă calitate, destinate antrenării și fine-tuning-ului modelelor de limbaj mari.

Generare de date sintetice pentru LLM-uri

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • conardli/easy-datasetAvatar ConardLi

    ConardLi/easy-dataset

    13,394Vezi pe GitHub↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    Easy-dataset is an end-to-end platform for managing ML datasets that explicitly covers automated synthetic data generation with LLM integration, quality filtering, and batch processing, directly matching the search for a synthetic data generation framework.

    JavaScriptSynthetic Data PipelinesSynthetic Data Generators
    Vezi pe GitHub↗13,394
  • tatsu-lab/stanford_alpacaAvatar tatsu-lab

    tatsu-lab/stanford_alpaca

    30,266Vezi pe GitHub↗

    This project provides an end-to-end framework for adapting large language models to follow user instructions through supervised fine-tuning. It functions as a comprehensive training pipeline that enables the creation of specialized assistant models by minimizing the difference between predicted outputs and target responses within structured instruction datasets. The framework distinguishes itself by integrating synthetic data generation with memory-efficient training techniques. It utilizes powerful language models to iteratively expand small sets of human-written seeds into diverse, high-qua

    Stanford Alpaca is an end-to-end framework that uses LLMs to iteratively expand small seed sets into diverse synthetic instruction-following datasets, making it a clear fit for generating synthetic training data with LLMs.

    PythonSynthetic Data Generators
    Vezi pe GitHub↗30,266
  • meta-llama/synthetic-data-kitAvatar meta-llama

    meta-llama/synthetic-data-kit

    1,602Vezi pe GitHub↗

    The synthetic data kit is an integrated framework designed to generate, curate, and format training datasets for language models. It provides an end-to-end pipeline that transforms raw source documents into structured data suitable for fine-tuning, reasoning, and tool-use model training. The framework distinguishes itself through a modular orchestration engine that manages the entire lifecycle of data preparation. It supports multimodal input by extracting both text and image content from various file formats, while employing context-aware chunking to maintain semantic coherence. The generati

    This tool from Meta is purpose-built for generating high-quality synthetic datasets using LLMs, fitting the requirement for an open-source framework that integrates LLM APIs and supports prompt-driven generation.

    PythonSynthetic Data Generators
    Vezi pe GitHub↗1,602
  • camel-ai/owlAvatar camel-ai

    camel-ai/owl

    19,864Vezi pe GitHub↗

    Owl is a framework for agentic workflow automation and multi-agent orchestration. It functions as a system for coordinating autonomous large language model agents to decompose and execute complex tasks through shared communication and collaborative planning. The project distinguishes itself through a multi-modal toolset for processing images, audio, and video, alongside a synthetic data generator that produces domain-specific datasets using self-instruct and verifier loops. It further incorporates a retrieval-augmented generation pipeline framework that integrates long-term memory and real-ti

    Owl is an agent orchestration framework that includes a synthetic data generator using LLM self-instruct and verifier loops, so it can produce domain-specific datasets — fitting the core need for synthetic data generation, though its primary focus is broader.

    PythonSynthetic Data Generators
    Vezi pe GitHub↗19,864
  • opendcai/dataflowAvatar OpenDCAI

    OpenDCAI/DataFlow

    2,926Vezi pe GitHub↗

    DataFlow is an agent-based workflow orchestrator and data pipeline designed to synthesize, clean, and augment large-scale datasets for training large language models. It functions as a synthetic data generator and text curation tool, utilizing an intelligent assistant to assemble modular processing operators into functional pipelines based on user requirements. The project distinguishes itself through a low-code approach, providing a web-based visual interface for designing and monitoring multi-stage execution flows. It features an operator-based registry system that allows for the integratio

    DataFlow is an agent-based workflow orchestrator for synthesizing, cleaning, and augmenting large-scale LLM training datasets, with LLM backend integration (vLLM, SGLang) and a low-code visual pipeline designer, directly addressing your need for a synthetic data generation framework.

    PythonText Quality Filtering
    Vezi pe GitHub↗2,926
  • datajuicer/data-juicerAvatar datajuicer

    datajuicer/data-juicer

    6,574Vezi pe GitHub↗

    Data-Juicer is an open-source framework for cleaning, filtering, deduplicating, and transforming multimodal datasets to prepare them for training large language and vision models. It functions as a distributed data pipeline engine that runs processing jobs across Ray clusters, handling billions of samples with automatic operator fusion and adaptive parallelism. The framework provides a library of operators that leverage large language models for semantic extraction, filtering, and data synthesis within processing pipelines. The project distinguishes itself through a YAML-based data recipe sys

    Data-Juicer is a pipeline framework that includes LLM-powered operators for data synthesis and filtering, making it a valid tool for generating synthetic training data, though its primary focus is on broader multimodal data processing rather than dedicated generation.

    PythonText Quality Filtering
    Vezi pe GitHub↗6,574
  • argilla-io/distilabelAvatar argilla-io

    argilla-io/distilabel

    3,277Vezi pe GitHub↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Distilabel is a framework for building scalable synthetic data and AI feedback pipelines using LLMs, directly matching the search for an open-source tool that generates training data with LLM integration and batch processing.

    PythonData Generation FrameworksData ProcessingData Processing Tools
    Vezi pe GitHub↗3,277
  • bespokelabsai/curatorAvatar bespokelabsai

    bespokelabsai/curator

    1,637Vezi pe GitHub↗

    Curator is a Python framework for generating synthetic datasets using LLMs, fitting the search for a synthetic data generation tool with LLM integration.

    PythonData Generation FrameworksData Processing
    Vezi pe GitHub↗1,637

Related searches

  • an open source tool for generating data
  • date sintetice cu aspect realist
  • toolkit pentru red-teaming-ul modelelor de limbaj
  • un framework pentru evaluarea calității output-ului LLM
  • framework pentru evaluarea aplicațiilor LLM
  • toolkit pentru construirea de agenți AI care utilizează instrumente
  • a platform for building generative AI applications
  • un framework pentru fine-tuning-ul modelelor de limbaj mari