2 مستودعات
Systems for versioning, splitting, and documenting large-scale datasets using standardized schemas and metadata.
Distinct from Large-Scale Dataset Management: Covers the full lifecycle of dataset management including versioning and metadata, not just high-capacity storage.
Explore 2 awesome GitHub repositories matching data & databases · Dataset Management Frameworks. Refine with filters or upvote what's useful.
This project is a dataset management framework and cross-framework data loader that provides a unified interface for reading data formats compatible with TensorFlow, JAX, and PyTorch. It serves as a library of curated public datasets provided as data streams and includes tools for building, versioning, and documenting large-scale datasets. The system differentiates itself through a distributed data processing engine capable of managing massive datasets across clusters using parallelized pipelines. It utilizes builder-based construction to standardize how data is downloaded and prepared, while
Offers a framework for versioning, splitting, and documenting large-scale datasets with standardized metadata and feature schemas.
This repository serves as the documentation source for the Hugging Face Hub, a collaborative platform designed for hosting, versioning, and discovering machine learning models, datasets, and interactive applications. It provides the foundational infrastructure for managing machine learning assets through Git-based repositories, which support large file storage, branching, and comprehensive commit history. The platform distinguishes itself by integrating metadata-driven discovery and structured management systems that allow users to attach licensing, task categories, and performance metrics to
Manages and versions large-scale datasets with structured metadata, schema definitions, and streaming capabilities.