awesome-repositories.com
ब्लॉग
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
treeverse avatar

treeverse/dvc

0
View on GitHub↗
15,679 स्टार्स·1,302 फोर्क्स·Python·Apache-2.0·6 व्यूज़dvc.org↗

Dvc

DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models using external storage and metadata pointers. It integrates with Git by utilizing placeholders to keep heavy artifacts out of the repository while maintaining a versioned link between code and data.

The system manages remote data caches through a synchronization layer that connects local environments to cloud storage or network filesystems. It also functions as an experiment tracker, recording hyperparameters and metrics to compare the performance of different model iterations.

The framework supports the definition of reproducible computational graphs by managing dependencies between code and commands. This capability enables the tracking of model lineage and the validation of data versioning consistency through commit-stage hooks.

Features

  • Data Pipeline Orchestration - Orchestrates complex sequences of data processing tasks by managing dependencies between code and data.
  • Model Lineage Trackers - Maintains a consistent provenance link between specific data versions, code, and hyperparameters used to produce a model.
  • Machine Learning Experiment Trackers - Monitors metrics and hyperparameters across multiple model iterations to identify optimal performance.
  • Data-Code Version Linking - Links specific versions of large datasets and models to the exact git commit of the code that produced them.
  • Content-Addressable Storage - Implements content-addressable storage using cryptographic hashes to ensure data integrity and deduplicate large artifacts.
  • Remote Dataset Caching - Provides a synchronization layer to cache large remote datasets locally using hash-based integrity verification.
  • Dataset Versioning Platforms - Provides a workflow for tracking historical versions of large-scale datasets to ensure machine learning reproducibility.
  • Artifact Versioning - Tracks large datasets and machine learning models using external caches and repository placeholders.
  • Git-Integrated Data Versioning - Integrates large file tracking with Git by using placeholders to keep heavy artifacts out of the repository.
  • Pointer-Based Tracking - Uses lightweight pointer files in Git to track large binary assets stored in an external cache.
  • Data Pipeline Definitions - Allows users to define data processing pipelines as version-controlled code to ensure reproducibility.
  • Directed Acyclic Graph Engines - Provides a DAG-based execution engine to manage computational dependencies between data and code.
  • ML Pipeline Reproducibility - Defines dependencies between data and code to ensure computational graphs are rebuilt reliably across environments.
  • Experiment Tracking - Records hyperparameters and metrics to compare the performance of different model iterations and training workflows.
  • Cloud Storage Sync Tools - Synchronizes local data caches with remote cloud storage providers using standard transfer protocols.
  • Backup Storage Backends - Provides drivers and configurations to offload large data caches to cloud or network storage providers.
  • Experiment Result Comparators - Records hyperparameters and performance metrics in structured files to enable comparative analysis of model iterations.

स्टार हिस्ट्री

treeverse/dvc के लिए स्टार हिस्ट्री चार्टtreeverse/dvc के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

अक्सर पूछे जाने वाले प्रश्न

treeverse/dvc क्या करता है?

DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models using external storage and metadata pointers. It integrates with Git by utilizing placeholders to keep heavy artifacts out of the repository while maintaining a versioned link between code and data.

treeverse/dvc की मुख्य विशेषताएं क्या हैं?

treeverse/dvc की मुख्य विशेषताएं हैं: Data Pipeline Orchestration, Model Lineage Trackers, Machine Learning Experiment Trackers, Data-Code Version Linking, Content-Addressable Storage, Remote Dataset Caching, Dataset Versioning Platforms, Artifact Versioning।

treeverse/dvc के कुछ ओपन-सोर्स विकल्प क्या हैं?

treeverse/dvc के ओपन-सोर्स विकल्पों में शामिल हैं: iterative/dvc — DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models.… wandb/wandb — Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow… apache/incubator-airflow — This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author,… aimhubio/aim — Aim is an open-source platform for logging, visualizing, and comparing machine learning training runs and LLM traces.… polyaxon/polyaxon — Polyaxon is a Kubernetes-native machine learning orchestration platform and MLOps pipeline orchestrator. It serves as… maiot-io/zenml — ZenML is an extensible machine learning orchestration framework designed to manage the end-to-end lifecycle of data…

Dvc के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Dvc के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • iterative/dvciterative का अवतार

    iterative/dvc

    15,680GitHub पर देखें↗

    DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache. The project serves as an experiment tracker and remote storage synchronizer, enabling the execution and comparison of machine learning iterations based on hyperparameters and performance metrics. It provides a bridge for pushing and pulling these large data artifacts between local environments and cloud or on-premi

    Python
    GitHub पर देखें↗15,680
  • wandb/wandbwandb का अवतार

    wandb/wandb

    10,844GitHub पर देखें↗

    Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow orchestration. It provides a comprehensive suite of tools for logging, visualizing, and versioning training metrics, model artifacts, and hyperparameter sweeps to ensure reproducibility across development cycles. The platform also functions as an observability tool for large language model applications, enabling the tracing of execution steps, token usage, and reasoning processes. The project distinguishes itself through its event-driven automation capabilities, which allow users

    Pythonaicollaborationdata-science
    GitHub पर देखें↗10,844
  • apache/incubator-airflowapache का अवतार

    apache/incubator-airflow

    45,840GitHub पर देखें↗

    This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author, schedule, and monitor complex data pipelines. It functions as a directed acyclic graph manager and scheduler, allowing users to define data movement and transformation tasks as code to ensure precise execution order and maintainability. The platform distinguishes itself by treating workflows as code, enabling pipelines to be versioned and tested through a standard programming language. It utilizes a system of extensible operators to encapsulate integration logic and employs a templat

    Python
    GitHub पर देखें↗45,840
  • aimhubio/aimaimhubio का अवतार

    aimhubio/aim

    6,159GitHub पर देखें↗

    Aim is an open-source platform for logging, visualizing, and comparing machine learning training runs and LLM traces. It provides a remote tracking server and a comparison UI, functioning as an ML experiment tracker, AI workflow logger, and LLM trace recorder that captures prompts, generations, and tool calls from AI applications. The platform distinguishes itself through a run-based data model with local SQLite storage, real-time metric streaming, and a plugin-based explorer system that supports specialized visual analysis of metrics, images, audio, and text. It offers a Python SDK with cont

    Python
    GitHub पर देखें↗6,159
Dvc के सभी 30 विकल्प देखें→