awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
iterative avatar

iterative/dvc

0
View on GitHub↗
15,680 Stars·1,302 Forks·Python·Apache-2.0·17 Aufrufedvc.org↗

Dvc

DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache.

The project serves as an experiment tracker and remote storage synchronizer, enabling the execution and comparison of machine learning iterations based on hyperparameters and performance metrics. It provides a bridge for pushing and pulling these large data artifacts between local environments and cloud or on-premises storage.

The tool covers data pipeline automation through the definition and execution of computational graphs, ensuring only components impacted by changes are rerun. It further supports model reproducibility by reconstructing specific experiment states and syncing the corresponding data and code versions.

Features

  • Dataset Versioning Systems - Tracks large datasets and models using lightweight metadata in version control while storing binaries in an external cache.
  • Pointer-Based Tracking - Tracks large datasets using lightweight meta-files in version control while storing binaries in an external cache.
  • Experiment Tracking - Logs and tracks combinations of hyperparameters and performance metrics to compare machine learning model iterations.
  • Machine Learning Experiment Trackers - Provides systems for monitoring metrics and hyperparameters across multiple machine learning model iterations.
  • Model Reproducibility Tools - Ensures model reproducibility by syncing exact data and code versions to reconstruct specific experiment states.
  • Model Versioning Systems - Tracks and manages iterations of machine learning models and their associated data artifacts for reproducibility.
  • Content-Addressable Storage - Implements a content-addressable storage system using hashes to deduplicate large data artifacts.
  • Data Pipeline Automation - Executes structured data processing workflows and automatically reruns only modified components.
  • Data Pipeline Orchestration - Allows the definition and orchestration of complex data processing sequences through computational graphs.
  • Data Pipeline Orchestrators - Automates complex sequences of data processing tasks using computational graphs with automatic change detection.
  • Dataset Versioning Platforms - Provides a platform for versioning large research datasets and ML models to ensure training reproducibility.
  • Hash-Based Change Detection - Uses cryptographic checksums to detect changes in data or code and determine if pipeline stages need updating.
  • Workflow Orchestration - Provides DAG-based pipeline execution to orchestrate data processing steps and optimize re-execution.
  • State Reconstruction - Enables the reconstruction of specific experiment states and data versions to reproduce results.
  • Cloud Storage Sync Tools - Synchronizes local data caches with cloud platforms or on-premises network storage.
  • Dataset Comparators - Analyzes differences and statistical drift between different versions of datasets, models, and parameters.
  • Storage Synchronization Services - Implements automated synchronization of large datasets and models between local caches and remote cloud or on-premises storage.
  • Remote Build Caches - Provides a remote cache for pushing and pulling large data artifacts to facilitate team collaboration.
  • Machine Learning - CLI tool for version control of machine learning data.
  • Machine Learning Operations - Version control system for machine learning projects.
  • MLOps and Infrastructure - Data version control for ML projects.
  • Model Management - Data version control for ML projects.
  • Datenmanagement - Versions data and models for ML experiment management.
  • Data Management Systems - Version control system for data in machine learning projects.
  • Data Science and ML - Support and discussion for open-source data version control systems.
  • Experiment and Data Management - Git-based version control system for ML models and data.
  • Infrastructure and Serving - Version control for large files.
  • MLOps and Pipelines - Version control system for data and models.
  • Data Science Tooling - Version control system for data and models.
  • Data Science Tools - Version control system for data science projects.
  • Experimentation Tracking - Provides version control for data, models, and experiment pipelines.
  • Project Documentation Examples - Uses a website-like menu and animation for workflows.
  • Versionskontrolle - Versions datasets and machine learning models.
  • Version Control Systems - Version control for data and machine learning models.

Star-Verlauf

Star-Verlauf für iterative/dvcStar-Verlauf für iterative/dvc

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Dvc

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Dvc.
  • treeverse/dvcAvatar von treeverse

    treeverse/dvc

    15,679Auf GitHub ansehen↗

    DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models using external storage and metadata pointers. It integrates with Git by utilizing placeholders to keep heavy artifacts out of the repository while maintaining a versioned link between code and data. The system manages remote data caches through a synchronization layer that connects local environments to cloud storage or network filesystems. It also functions as an experiment tracker, recording hyperparameters and metrics to compare the performance of different model iterations.

    Pythonaidata-sciencedata-version-control
    Auf GitHub ansehen↗15,679
  • wandb/wandbAvatar von wandb

    wandb/wandb

    10,844Auf GitHub ansehen↗

    Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow orchestration. It provides a comprehensive suite of tools for logging, visualizing, and versioning training metrics, model artifacts, and hyperparameter sweeps to ensure reproducibility across development cycles. The platform also functions as an observability tool for large language model applications, enabling the tracing of execution steps, token usage, and reasoning processes. The project distinguishes itself through its event-driven automation capabilities, which allow users

    Pythonaicollaborationdata-science
    Auf GitHub ansehen↗10,844
  • mlflow/mlflowAvatar von mlflow

    mlflow/mlflow

    26,554Auf GitHub ansehen↗
    Pythonagentopsagentsai
    Auf GitHub ansehen↗26,554
  • iterative/mlemAvatar von iterative

    iterative/mlem

    718Auf GitHub ansehen↗

    🐶 A tool to package, serve, and deploy any ML model on any platform. Archived to be resurrected one day🤞

    Python
    Auf GitHub ansehen↗718
Alle 30 Alternativen zu Dvc anzeigen→

Häufig gestellte Fragen

Was macht iterative/dvc?

DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache.

Was sind die Hauptfunktionen von iterative/dvc?

Die Hauptfunktionen von iterative/dvc sind: Dataset Versioning Systems, Pointer-Based Tracking, Experiment Tracking, Machine Learning Experiment Trackers, Model Reproducibility Tools, Model Versioning Systems, Content-Addressable Storage, Data Pipeline Automation.

Welche Open-Source-Alternativen gibt es zu iterative/dvc?

Open-Source-Alternativen zu iterative/dvc sind unter anderem: treeverse/dvc — DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models… wandb/wandb — Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow… mlflow/mlflow. iterative/mlem — 🐶 A tool to package, serve, and deploy any ML model on any platform. Archived to be resurrected one day🤞. christoschristofidis/awesome-deep-learning — This project is a curated directory of resources, libraries, and frameworks designed to support the development,… allegroai/clearml — ClearML is a comprehensive MLOps platform designed to manage the entire machine learning lifecycle. It functions as an…