6 مستودعات
Mechanisms for linking external configuration files to database records to display dataset attributes.
Distinct from File Storage and Metadata Management: Focuses on mapping external text files to database metadata for display, rather than general file system metadata services.
Explore 6 awesome GitHub repositories matching data & databases · Dataset Metadata Mapping. Refine with filters or upvote what's useful.
Datasette is a tool for publishing and sharing SQLite databases as public websites. It functions as a data publishing system that provides searchable interfaces and JSON APIs to expose the contents of SQLite files. The project enables both server-side and client-side execution. It can operate as an API server or as a database browser that runs entirely within a web browser using WebAssembly, allowing for serverless database access. The system supports a variety of deployment strategies, including containerized images for cloud hosting and a local development server for testing. It includes c
Provides a mechanism to attach licensing and source information to datasets via external configuration files.
BasicSR is a PyTorch-based image restoration toolbox and framework designed for training and deploying deep learning models to upscale, denoise, and deblur images and videos. It serves as a comprehensive system for image super-resolution and video quality restoration, providing the necessary infrastructure to recover fine visual details and increase pixel density. The project distinguishes itself through specialized toolkits for facial image enhancement and high-fidelity face synthesis, as well as a dedicated video quality restoration suite that utilizes deformable convolutions and generative
Includes tools to generate metadata files that organize image paths and properties for training.
This project is a research data sharing framework and provenance protocol designed to ensure computational reproducibility. It provides a standardized set of guidelines for transforming raw source data into tidy formats through documented processing scripts and cleaning workflows. The framework distinguishes itself by emphasizing a strict provenance-based packaging system. It requires the organization of raw data, processing recipes, and code books into a single package, ensuring that original unmodified sources are preserved to allow for independent verification of all transformation steps.
Decouples metadata from datasets by storing variable units and experimental design details in separate reference files.
Parabolic is a graphical frontend for the yt-dlp engine, serving as a web media downloader and extractor. It provides a visual interface for saving high-quality video and audio content from various web platforms into local files. The application functions as a multi-format media exporter, allowing content to be saved into diverse audio and video containers. It includes a media metadata manager to capture and store associated information, such as subtitles and technical metadata, alongside the downloaded files. The system supports batch content acquisition through concurrent download manageme
Maps raw technical metadata from external tools into a structured format for consistent storage and display.
Argilla هي أداة تعاونية للتغذية الراجعة للذكاء الاصطناعي ونظام إدارة تنظيم البيانات. تعمل كمنصة لمجموعات البيانات التي تعتمد على الإنسان في الحلقة (Human-in-the-loop) مصممة لتنسيق القائمين على التعليق التوضيحي والخبراء في المجال في تصنيف وتقييم وتحسين عينات البيانات لمشاريع تعلم الآلة. تركز المنصة على تنظيم مجموعات بيانات نماذج اللغات الكبيرة وسير عمل التعلم التعزيزي من التغذية الراجعة البشرية. توفر مساحة عمل مشتركة لدمج الخبرة البشرية في تطوير الذكاء الاصطناعي للتحقق من مخرجات النموذج وتصحيح أخطاء البيانات. يدير النظام خط أنابيب بيانات تعلم الآلة من البداية إلى النهاية، بما في ذلك استيراد مجموعات البيانات من المراكز الخارجية، وتحديد مخططات التغذية الراجعة المخصصة للتصنيفات والترتيبات، وتصدير البيانات المشروحة. يدعم إدارة البيانات البرمجية وإنشاء سير عمل آلي لتحسين أداء النموذج بشكل تكراري.
Links external dataset columns to internal feedback schemas to ensure data integrity during import and export.
This repository serves as the documentation source for the Hugging Face Hub, a collaborative platform designed for hosting, versioning, and discovering machine learning models, datasets, and interactive applications. It provides the foundational infrastructure for managing machine learning assets through Git-based repositories, which support large file storage, branching, and comprehensive commit history. The platform distinguishes itself by integrating metadata-driven discovery and structured management systems that allow users to attach licensing, task categories, and performance metrics to
Allows attaching structured information like licensing and task categories to datasets to improve discoverability.