awesome-repositories.com
Blog
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Versionnage de données et d'artefacts ML

Classement mis à jour le 30 juin 2026

For système de contrôle de version pour données ML, the strongest matches are treeverse/dvc (DVC is a Git-integrated data versioning tool and pipeline), treeverse/lakefs (lakeFS is a data lake versioning system that provides) and iterative/dvc (DVC is the leading open-source data version control tool). attic-labs/noms and wandb/client round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Outils pour suivre, versionner et gérer de larges jeux de données de machine learning et des artefacts de modèles comme du code.

Versionnage de données et d'artefacts ML

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • treeverse/dvcAvatar de treeverse

    treeverse/dvc

    15,679Voir sur GitHub↗

    DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models using external storage and metadata pointers. It integrates with Git by utilizing placeholders to keep heavy artifacts out of the repository while maintaining a versioned link between code and data. The system manages remote data caches through a synchronization layer that connects local environments to cloud storage or network filesystems. It also functions as an experiment tracker, recording hyperparameters and metrics to compare the performance of different model iterations.

    DVC is a Git-integrated data versioning tool and pipeline orchestrator that tracks datasets and ML models via external storage, provides experiment tracking with metrics comparison, and supports cloud storage sync — directly matching the request for Git-like version control of data and ML artifacts.

    PythonGit-Integrated Data Versioning
    Voir sur GitHub↗15,679
  • treeverse/lakefsAvatar de treeverse

    treeverse/lakeFS

    5,406Voir sur GitHub↗

    lakeFS is a data lake versioning system that provides Git-like branching and commits for large datasets stored in object storage. It functions as a version control layer, enabling the creation of immutable snapshots, atomic commits, and zero-copy branching to create isolated environments for data experimentation without duplicating physical files. The system serves as an S3-compatible storage gateway and an Iceberg REST catalog, allowing standard cloud storage protocols and compatible clients to manage versioned tables. It acts as a data quality gatekeeper by using an event-driven hook system

    lakeFS is a data lake versioning system that provides Git-like branching and commits for large datasets stored in object storage, making it a direct fit for version controlling datasets with Git workflows; while it focuses on data lakes rather than ML-specific artifacts, it covers dataset versioning, snapshotting, and cloud integration.

    GoDifferential Dataset ComparisonsData Versioning SystemsS3-Compatible Storage Adapters
    Voir sur GitHub↗5,406
  • iterative/dvcAvatar de iterative

    iterative/dvc

    15,680Voir sur GitHub↗

    DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache. The project serves as an experiment tracker and remote storage synchronizer, enabling the execution and comparison of machine learning iterations based on hyperparameters and performance metrics. It provides a bridge for pushing and pulling these large data artifacts between local environments and cloud or on-premi

    DVC is the leading open-source data version control tool that tracks datasets and ML artifacts with Git-like workflows, supporting large file storage, pipeline orchestration, model registry, dataset snapshotting, cloud storage sync, and diff/compare—covering everything needed for this search.

    PythonDataset Versioning SystemsPointer-Based TrackingContent-Addressable Storage
    Voir sur GitHub↗15,680
  • attic-labs/nomsAvatar de attic-labs

    attic-labs/noms

    7,422Voir sur GitHub↗

    Noms is a distributed version control database and content-addressable data store. It identifies data by cryptographic hashes to ensure integrity and deduplication, while tracking dataset state changes through a sequence of immutable commits to enable branching, forking, and historical recovery. The system functions as a peer-to-peer data synchronizer, reconciling state between disconnected database instances to ensure all nodes converge on the same data. It distinguishes itself as a schema-flexible document store that supports self-describing types, allowing schemas to evolve and widen as ne

    Noms is a distributed version-control database that uses commits and branching to version datasets, directly matching the Git-like workflow you need for data versioning, though it does not include ML-specific pipeline tracking or a model registry.

    GoVersioned Dataset Snapshots
    Voir sur GitHub↗7,422
  • wandb/clientAvatar de wandb

    wandb/client

    11,128Voir sur GitHub↗

    This project is a collection of utilities designed for machine learning experiment tracking, data versioning, and the observability of large language model applications. It provides a client for recording hyperparameters and metrics during training to visualize performance trends and compare different model versions. The tool includes a model evaluation framework that uses custom scorers and automated judges to assess the quality of generated text outputs. It also provides observability tools to monitor and debug the execution flow and runtime behavior of language model applications. The sys

    This repository is the Python client for Weights & Biases, a platform that provides Git-like versioning for datasets and ML artifacts, experiment tracking, model registry, and dataset snapshotting — exactly the kind of tool this search is after, though you'll need the wandb service to store and retrieve versions.

    PythonExperiment TrackingArtifact VersioningData Lineage
    Voir sur GitHub↗11,128

Related searches

  • registre pour le versioning de modèles ML
  • a version control system for software development
  • le versioning façon Git pour ma base de données
  • projet pour apprendre Git en le réimplémentant
  • un outil de versioning et de gestion de prompts
  • un système de contrôle de version pour le code source
  • un système de contrôle de version pour le développement logiciel
  • un CMS basé sur Git