3 repository-uri
Validation of data imports to ensure they adhere to a standardized conversation format.
Distinct from Structural Data Validators: Unlike general structural validation, this specifically targets the conversation schema required for LLM training.
Explore 3 awesome GitHub repositories matching data & databases · Conversation Structure Validation. Refine with filters or upvote what's useful.
Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo
Converts various file formats into a standardized conversation structure while validating schema compatibility.
Harmony este un SDK AI conceput pentru tokenizarea conversațiilor, formatarea layout-urilor de raționament, parsarea output-urilor brute și definirea schemelor de apelare a instrumentelor. Acesta oferă un sistem pentru conversia dialogului structurat și a apelurilor de instrumente în secvențe de token-uri necesare pentru inferența și antrenarea modelelor de limbaj mari. Proiectul include un formator de output care structurează lanțurile de raționament și output-urile multi-canal în layout-uri consistente pentru a preveni pierderea de token-uri. De asemenea, dispune de un parser de răspunsuri care transformă token-urile brute de completare și fluxurile live înapoi în obiecte de mesaje structurate și roluri. SDK-ul gestionează integrarea instrumentelor printr-un framework pentru definirea funcțiilor apelabile și a namespace-urilor. Oferă, de asemenea, capabilități pentru parsarea token-urilor în timp real, configurarea comportamentului modelului și serializarea conversațiilor cu stare.
Organizes messages between users and assistants into formal data structures to facilitate consistent interaction with a model.
This project is a training pipeline and framework for developing Chinese language models based on the Llama 2 architecture. It functions as a distributed GPU trainer and dataset preprocessing toolkit designed for both the initial pre-training of baseline models and subsequent supervised fine-tuning. The system distinguishes itself through a specialized workflow for Chinese text, incorporating a data curation pipeline that uses similarity hashing for deduplication and a tokenization process that converts raw text into memory-mapped binary files for efficient disk access. It implements a superv
Filters and cleans raw conversational datasets into structured formats by enforcing prompt and answer length constraints.