awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

15 repository-uri

Awesome GitHub RepositoriesDataset Management

Explore 15 awesome GitHub repositories matching artificial intelligence & ml · Dataset Management. Refine with filters or upvote what's useful.

Awesome Dataset Management GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • ultralytics/ultralyticsAvatar ultralytics

    ultralytics/ultralytics

    58,468Vezi pe GitHub↗

    Ultralytics is a comprehensive computer vision framework designed for training, validating, and deploying deep learning models across a wide range of visual recognition tasks. It provides a unified interface for core operations including object detection, instance segmentation, pose estimation, and image classification. By utilizing a modular architecture, the platform allows users to swap model components to balance inference speed and accuracy requirements for diverse applications. The framework distinguishes itself through its support for real-time processing and flexible deployment. It in

    Facilitates the organization of training data and the conversion of models into standard file formats for broad compatibility.

    Pythonclicomputer-visiondeep-learning
    Vezi pe GitHub↗58,468
  • ultralytics/yolov5Avatar ultralytics

    ultralytics/yolov5

    57,528Vezi pe GitHub↗

    YOLOv5 is a comprehensive computer vision framework designed for end-to-end deep learning, specializing in real-time object detection, image classification, and instance segmentation. It provides a unified toolkit that manages the entire lifecycle of a model, from initial dataset configuration and hyperparameter tuning to high-speed inference and deployment. The framework utilizes a modular neural architecture, allowing users to swap backbone and head components to tailor models for specific visual tasks. What distinguishes this project is its focus on production-ready deployment and model ef

    Standardizes data organization through structured files that map class labels, file paths, and validation splits for training.

    Pythoncoremldeep-learningios
    Vezi pe GitHub↗57,528
  • roboflow/supervisionAvatar roboflow

    roboflow/supervision

    44,437Vezi pe GitHub↗

    Supervision is a computer vision toolset for normalizing model outputs, managing datasets, and visualizing annotations. It provides a framework to convert predictions from various classification and detection models into a standardized data format to ensure interoperability across different computer vision pipelines. The library features a post-processor for filtering, counting, and tracking detected objects across image frames and video streams. It includes capabilities for large image tiling to improve the detection of small objects and tools for assigning persistent identities to objects t

    Provides utilities for converting computer vision datasets between common formats to ensure model compatibility.

    Pythonclassificationcococomputer-vision
    Vezi pe GitHub↗44,437
  • huggingface/lerobotAvatar huggingface

    huggingface/lerobot

    21,687Vezi pe GitHub↗

    This project is a comprehensive research platform designed for the end-to-end lifecycle of robotic learning. It provides a modular framework for training neural network policies—specifically through imitation and reinforcement learning—and deploying them onto physical robotic hardware. By offering a unified interface for hardware abstraction, the platform decouples high-level control logic from the specific sensors and actuators of diverse robotic systems. The framework distinguishes itself through a standardized approach to data and policy management. It utilizes a consistent schema for reco

    Standardizes data storage using synchronized video and state files for efficient dataset management.

    Python
    Vezi pe GitHub↗21,687
  • opencv/cvatAvatar opencv

    opencv/cvat

    16,086Vezi pe GitHub↗

    CVAT este un instrument open-source de adnotare a viziunii computerizate și o platformă de gestionare a seturilor de date vizuale. Acesta oferă o interfață auto-găzduită pentru etichetarea imaginilor, videoclipurilor și datelor 3D pentru a crea seturi de date pentru modelele AI de viziune. Platforma dispune de etichetarea datelor asistată de AI pentru a automatiza crearea de măști și casete de delimitare, utilizând un sistem de plug-in-uri pentru a conecta modele externe de învățare automată. Include un sistem de asigurare a calității bazat pe consens care verifică acuratețea etichetelor prin compararea adnotărilor independente. Sistemul acoperă gestionarea colaborativă a echipei, organizarea proiectelor prin descompunerea sarcinilor și integrarea stocării la distanță în cloud. De asemenea, oferă un API REST pentru controlul programatic al fluxului de lucru și importul/exportul de date în formate standard din industrie.

    Implements utilities for organizing, annotating, and converting visual datasets to support machine learning training pipelines.

    Python
    Vezi pe GitHub↗16,086
  • cvat-ai/cvatAvatar cvat-ai

    cvat-ai/cvat

    15,317Vezi pe GitHub↗

    CVAT is an open-source, web-based platform designed for annotating images, videos, and 3D point clouds to create high-quality training datasets for machine learning. It functions as a containerized server that orchestrates the entire lifecycle of computer vision data, from initial task creation and manual labeling to quality assurance and final dataset export. The platform distinguishes itself through deep integration with machine learning models, allowing users to deploy custom AI models as serverless functions for automated object detection, tracking, and skeleton annotation. It supports co

    Exports annotated datasets into structured formats including geometric shapes, attributes, and tracking identifiers for machine learning training.

    Pythonannotationannotation-toolannotations
    Vezi pe GitHub↗15,317
  • modelscope/ms-swiftAvatar modelscope

    modelscope/ms-swift

    14,597Vezi pe GitHub↗

    This project is a comprehensive toolkit designed for the full lifecycle management of large language and multimodal models. It functions as a unified orchestrator that handles the entire development process, ranging from dataset preparation and supervised fine-tuning to advanced reinforcement learning alignment and production-ready inference deployment. The platform distinguishes itself through a specialized reinforcement learning library that supports complex optimization algorithms, including group relative policy optimization and leave-one-out techniques, to improve model instruction-follo

    Provides access to a curated library of datasets with pre-calculated token statistics for model training.

    Pythondeepseek-r1embeddinggrpo
    Vezi pe GitHub↗14,597
  • conardli/easy-datasetAvatar ConardLi

    ConardLi/easy-dataset

    13,394Vezi pe GitHub↗

    Easy-dataset is a comprehensive platform designed for the end-to-end management of machine learning datasets, specifically tailored for language and vision model fine-tuning. It functions as a centralized environment for the entire data lifecycle, encompassing the automated generation of synthetic training data, the structural organization of document collections, and the systematic annotation of individual data points. The platform distinguishes itself through its integrated evaluation and orchestration capabilities. It provides a dedicated suite for benchmarking models, featuring blind side

    Provides a centralized interface to organize, maintain, and structure collections of documents and annotations for model training.

    JavaScriptdatasetfine-tuningjavascript
    Vezi pe GitHub↗13,394
  • axolotl-ai-cloud/axolotlAvatar axolotl-ai-cloud

    axolotl-ai-cloud/axolotl

    12,059Vezi pe GitHub↗

    Axolotl is a configuration-driven framework designed for the fine-tuning, evaluation, and quantization of large language models. It functions as a comprehensive orchestrator for distributed training, enabling users to manage complex workflows across multi-node and multi-GPU environments. By utilizing structured configuration files, the platform streamlines the setup of training parameters, dataset paths, and hardware distribution strategies. The project distinguishes itself through its support for diverse training methodologies, including full-parameter tuning, parameter-efficient adaptation,

    Structures training data to include dedicated thinking roles within system and assistant messages for reasoning models.

    Pythonfine-tuningllm
    Vezi pe GitHub↗12,059
  • doccano/doccanoAvatar doccano

    doccano/doccano

    10,674Vezi pe GitHub↗

    Doccano is a collaborative data labeling platform and machine learning dataset management system. It provides a web-based interface for teams to import raw text, mark datasets, and export structured annotations for model training. The project specifically supports text annotation for classification and named entity recognition tasks. It enables teams to coordinate multiple users on a single project to maintain consistent labeling guidelines and increase the speed of dataset creation. The system includes tools for data management and team coordination, providing the ability to import raw data

    Implements tools for importing raw data and exporting structured annotations for machine learning workflows.

    Python
    Vezi pe GitHub↗10,674
  • getmoto/motoAvatar getmoto

    getmoto/moto

    8,550Vezi pe GitHub↗

    Moto is a cloud service mockery framework and API mock server that simulates AWS infrastructure locally. It allows developers to test cloud-dependent code and verify infrastructure-as-code templates without deploying real resources or incurring costs. The project functions as an SDK interceptor that can patch existing service clients to redirect requests to a local mock environment. It can also be run as a standalone HTTP server, enabling any programming language to interact with the simulated endpoints. The framework covers a vast array of simulated capabilities, including data storage, com

    Simulates the management of dataset groups used for cloud-based forecasting and machine learning.

    Pythonawsbotoec2
    Vezi pe GitHub↗8,550
  • open-mmlab/mmposeAvatar open-mmlab

    open-mmlab/mmpose

    7,374Vezi pe GitHub↗

    MMPose is a PyTorch-based pose estimation toolbox and deep learning training pipeline designed for detecting 2D and 3D keypoints on humans, animals, and faces. It serves as a computer vision model zoo and a framework for both 2D pose estimation and 3D pose lifting. The project is distinguished by its modular architecture and extensibility, employing a registry-based system and hierarchical configurations to allow for custom algorithm integration and model pipeline customization. It supports diverse estimation paradigms, including top-down, bottom-up, and two-stage pose lifting workflows. The

    Uses standardized dataset configurations to define keypoint properties and skeleton connectivity.

    Pythonanimal-pose-estimationbenchmarkcpm
    Vezi pe GitHub↗7,374
  • kohya-ss/sd-scriptsAvatar kohya-ss

    kohya-ss/sd-scripts

    7,133Vezi pe GitHub↗

    sd-scripts is a suite of utilities designed for fine-tuning generative models, preprocessing datasets, and converting model weights. It provides a collection of scripts for executing Stable Diffusion training through methods such as DreamBooth, textual inversion, and full fine-tuning, alongside a framework for creating and managing Low-Rank Adaptation weights. The project features specialized capabilities for model weight conversion between different architectures and precision formats. It includes tools for merging adaptation weights into base models, extracting weights from trained models,

    Uses configuration files to define training data, high-resolution image settings, and aspect ratio bucketing.

    Python
    Vezi pe GitHub↗7,133
  • open-mmlab/mmocrAvatar open-mmlab

    open-mmlab/mmocr

    4,739Vezi pe GitHub↗

    mmocr este un framework de recunoaștere optică a caracterelor (OCR) bazat pe PyTorch, conceput pentru antrenarea și deployment-ul modelelor de detectare a textului, recunoaștere și extragere a informațiilor cheie. Servește ca un toolkit cuprinzător pentru detectarea și recunoașterea textului în scene, oferind biblioteci specializate pentru localizarea regiunilor de text și convertirea textului vizual în șiruri de caractere codificate de mașină. Proiectul se distinge printr-un framework de cercetare pentru extragerea informațiilor cheie și capabilități avansate de text spotting. Acestea includ spotting bazat pe puncte folosind transformatoare și utilizarea curbelor Bezier parametrizate pentru a identifica și transcrie text cu forme arbitrare. Framework-ul acoperă o suprafață largă de capabilități de viziune artificială, inclusiv gestionarea pipeline-ului de date pentru augmentarea și standardizarea seturilor de date OCR diverse, antrenarea modelelor cu scalare distribuită și evaluarea performanței folosind metrici OCR standard. Oferă, de asemenea, utilitare pentru manipularea poligoanelor geometrice și vizualizarea rezultatelor pentru auditarea predicțiilor față de adnotările ground truth. Sistemul este implementat în Python și suportă instalarea prin împachetarea mediului Docker.

    Generates Python configuration files that define data roots and annotation paths for training pipelines.

    Pythonabcnetabinetcrnn
    Vezi pe GitHub↗4,739
  • tensorflow/datasetsAvatar tensorflow

    tensorflow/datasets

    4,575Vezi pe GitHub↗

    This project is a dataset management framework and cross-framework data loader that provides a unified interface for reading data formats compatible with TensorFlow, JAX, and PyTorch. It serves as a library of curated public datasets provided as data streams and includes tools for building, versioning, and documenting large-scale datasets. The system differentiates itself through a distributed data processing engine capable of managing massive datasets across clusters using parallelized pipelines. It utilizes builder-based construction to standardize how data is downloaded and prepared, while

    Provides command-line utilities for building and versioning datasets to ensure reproducible data management.

    Python
    Vezi pe GitHub↗4,575
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Dataset Management

Explorează sub-etichetele

  • Dataset Configurations1 sub-tagStandardized file formats used to define data paths, class labels, and training/validation splits for machine learning models.
  • Dataset Management ToolsUtilities for organizing, annotating, and converting datasets for machine learning training.