awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to iterative/mlem

Open-source alternatives to Mlem

30 open-source projects similar to iterative/mlem, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Mlem alternative.

  • iterative/dvcAvatar von iterative

    iterative/dvc

    15,680Auf GitHub ansehen↗

    DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache. The project serves as an experiment tracker and remote storage synchronizer, enabling the execution and comparison of machine learning iterations based on hyperparameters and performance metrics. It provides a bridge for pushing and pulling these large data artifacts between local environments and cloud or on-premi

    Python
    Auf GitHub ansehen↗15,680
  • comet-ml/comet-examplesAvatar von comet-ml

    comet-ml/comet-examples

    174Auf GitHub ansehen↗

    Examples of Machine Learning code using Comet.ml

    Jupyter Notebook
    Auf GitHub ansehen↗174
  • feast-dev/feastAvatar von feast-dev

    feast-dev/feast

    6,727Auf GitHub ansehen↗

    Feast is an open-source feature store for machine learning that provides a central platform for defining, storing, and serving features across both training and inference workflows. It operates as a declarative system where feature definitions are written as code in Python files, synchronized to a central registry, and made available for low-latency online retrieval or point-in-time correct historical joins for training datasets. The project abstracts storage behind a pluggable architecture, allowing offline and online backends to be swapped without changing retrieval logic, and coordinates ma

    Pythonbig-datadata-engineeringdata-quality
    Auf GitHub ansehen↗6,727
  • dslp/dslp-repo-templateAvatar von dslp

    dslp/dslp-repo-template

    202Auf GitHub ansehen↗

    Template repository for data science lifecycle project

    Python
    Auf GitHub ansehen↗202

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Find more with AI search
  • dslp/dslpAvatar von dslp

    dslp/dslp

    527Auf GitHub ansehen↗

    The Data Science Lifecycle Process is a process for taking data science teams from Idea to Value repeatedly and sustainably. The process is documented in this repo.

    Auf GitHub ansehen↗527
  • asavinov/lambdoAvatar von asavinov

    asavinov/lambdo

    26Auf GitHub ansehen↗

    Feature engineering and machine learning: together at last!

    Python
    Auf GitHub ansehen↗26
  • ml-tooling/ml-workspaceAvatar von ml-tooling

    ml-tooling/ml-workspace

    3,540Auf GitHub ansehen↗

    🛠 All-in-one web-based IDE specialized for machine learning and data science.

    Jupyter Notebook
    Auf GitHub ansehen↗3,540
  • minerva-ml/steppy-toolkitAvatar von minerva-ml

    minerva-ml/steppy-toolkit

    23Auf GitHub ansehen↗

    Curated set of transformers that make your work with steppy faster and more effective :telescope:

    Python
    Auf GitHub ansehen↗23
  • tensorchord/envdAvatar von tensorchord

    tensorchord/envd

    2,211Auf GitHub ansehen↗

    🏕️ Reproducible development environment for humans and agents

    Go
    Auf GitHub ansehen↗2,211
  • iterative/cmlAvatar von iterative

    iterative/cml

    4,178Auf GitHub ansehen↗

    CML is a pipeline automation tool for training and evaluating machine learning models, functioning as a CI/CD system for machine learning. It serves as a cloud compute orchestrator and Git-based workflow manager that automates model training cycles through branch management, automated commits, and integrated reporting. The project distinguishes itself by provisioning ephemeral cloud instances or Kubernetes nodes to provide specialized hardware for compute-heavy tasks. It also manages remote compute runners, allowing the connection of self-hosted GPU clusters or on-premise machines to execute

    JavaScript
    Auf GitHub ansehen↗4,178
  • minerva-ml/steppyAvatar von minerva-ml

    minerva-ml/steppy

    136Auf GitHub ansehen↗

    Lightweight, Python library for fast and reproducible experimentation :microscope:

    Python
    Auf GitHub ansehen↗136
  • comet-ml/comet-llmAvatar von comet-ml

    comet-ml/comet-llm

    19,673Auf GitHub ansehen↗

    Comet LLM is an observability platform and evaluation framework designed for large language model applications and agentic workflows. It functions as a system for tracing, monitoring, and debugging execution flows while providing tools for prompt optimization and the enforcement of AI safety guardrails. The platform distinguishes itself through a combination of model-based scoring and heuristic metrics to quantify output quality and detect hallucinations. It includes a dedicated prompt and agent optimizer with an interactive playground for refining templates and tool configurations. For retri

    Python
    Auf GitHub ansehen↗19,673
  • allegroai/clearmlAvatar von allegroai

    allegroai/clearml

    6,733Auf GitHub ansehen↗

    ClearML is a comprehensive MLOps platform designed to manage the entire machine learning lifecycle. It functions as an experiment tracking tool, a data versioning system, and a pipeline orchestrator, while providing infrastructure for GPU cluster management and model serving. The platform is distinguished by its ability to handle hybrid-cloud compute scheduling and fractional GPU allocation, allowing multiple workloads to share a single hardware accelerator. It employs a metadata-based approach to data versioning, using virtual views to track large datasets and artifacts without duplicating r

    Python
    Auf GitHub ansehen↗6,733
  • benedekrozemberczki/littleballoffurAvatar von benedekrozemberczki

    benedekrozemberczki/littleballoffur

    715Auf GitHub ansehen↗

    Little Ball of Fur - A graph sampling extension library for NetworKit and NetworkX (CIKM 2020)

    Python
    Auf GitHub ansehen↗715
  • mindsdb/mindsdbAvatar von mindsdb

    mindsdb/mindsdb

    39,313Auf GitHub ansehen↗

    MindsDB is an AI-native database engine that treats machine learning models and autonomous agents as virtual tables. By mapping external data sources, predictive models, and third-party services directly into the database schema, it enables users to perform inference, data retrieval, and complex orchestration using standard SQL syntax. The platform distinguishes itself through an autonomous agent orchestrator that executes iterative reasoning loops, allowing agents to plan data access and synthesize natural language responses from connected knowledge bases. It functions as a federated data ga

    Makefileagentsaianalytics
    Auf GitHub ansehen↗39,313
  • benedekrozemberczki/karateclubAvatar von benedekrozemberczki

    benedekrozemberczki/karateclub

    2,284Auf GitHub ansehen↗

    Karate Club: An API Oriented Open-source Python Framework for Unsupervised Learning on Graphs (CIKM 2020)

    Python
    Auf GitHub ansehen↗2,284
  • linealabs/lineapyAvatar von LineaLabs

    LineaLabs/lineapy

    670Auf GitHub ansehen↗

    Move fast from data science prototype to pipeline. Capture, analyze, and transform messy notebooks into data pipelines with just two lines of code.

    Jupyter Notebook
    Auf GitHub ansehen↗670
  • hydrospheredata/mistAvatar von Hydrospheredata

    Hydrospheredata/mist

    324Auf GitHub ansehen↗

    Serverless proxy for Spark cluster

    Scala
    Auf GitHub ansehen↗324
  • adrotog/pandasguiAvatar von adrotog

    adrotog/PandasGUI

    3,259Auf GitHub ansehen↗

    A GUI for Pandas DataFrames

    Python
    Auf GitHub ansehen↗3,259
  • logicalclocks/hopsworksAvatar von logicalclocks

    logicalclocks/hopsworks

    1,302Auf GitHub ansehen↗

    Hopsworks - Data-Intensive AI platform with a Feature Store

    Java
    Auf GitHub ansehen↗1,302
  • awslabs/autogluonAvatar von awslabs

    awslabs/autogluon

    10,481Auf GitHub ansehen↗

    AutoGluon is an automated machine learning framework designed to optimize model selection and hyperparameter tuning across tabular, text, image, and time series data. It functions as an ensemble learning library and a tabular data prediction engine, aiming to build high-accuracy predictive models without manual algorithm selection. The framework integrates multimodal machine learning pipelines that combine disparate data types into a single representation using specialized encoders. It also includes a probabilistic time series forecaster that fits multiple statistical and deep learning models

    Python
    Auf GitHub ansehen↗10,481
  • alteryx/featuretoolsAvatar von alteryx

    alteryx/featuretools

    7,658Auf GitHub ansehen↗

    Featuretools is an automated feature engineering library and data transformation framework written in Python. It automatically generates machine learning feature vectors from multi-table datasets by applying synthesis patterns to relational and timestamped data. The system functions as a distributed feature synthesis engine, allowing the process of creating feature vectors to scale across multiple cores or clusters to handle large-scale datasets. The library supports the synthesis of multi-table datasets, time series feature generation, and the creation of custom machine learning primitives

    Python
    Auf GitHub ansehen↗7,658
  • cleanlab/cleanlabAvatar von cleanlab

    cleanlab/cleanlab

    11,513Auf GitHub ansehen↗

    Cleanlab is a data-centric AI library and toolkit designed to improve machine learning model performance by detecting label errors and increasing overall dataset quality. It implements a confident learning framework that iteratively refines label noise estimates by comparing model predictions with estimated label probabilities to identify mislabeled examples. The project provides specialized utilities for active learning optimization, allowing for the selection of the most impactful examples for labeling or re-labeling. It also includes an outlier detection tool to identify atypical data poin

    Pythonactive-learningannotationanomaly-detection
    Auf GitHub ansehen↗11,513
  • astrazeneca/rexmexAvatar von AstraZeneca

    AstraZeneca/rexmex

    278Auf GitHub ansehen↗

    A general purpose recommender metrics library for fair evaluation.

    Python
    Auf GitHub ansehen↗278
  • astrazeneca/chemicalxAvatar von AstraZeneca

    AstraZeneca/chemicalx

    781Auf GitHub ansehen↗

    A PyTorch and TorchDrug based deep learning library for drug pair scoring. (KDD 2022)

    Python
    Auf GitHub ansehen↗781
  • julialang/ijulia.jlAvatar von JuliaLang

    JuliaLang/IJulia.jl

    2,889Auf GitHub ansehen↗

    Julia kernel for Jupyter

    Julia
    Auf GitHub ansehen↗2,889
  • albumentations-team/albumentationsAvatar von albumentations-team

    albumentations-team/albumentations

    15,308Auf GitHub ansehen↗

    Albumentations is a computer vision image augmentation library designed to increase training data diversity for deep learning models. It provides a toolset for applying geometric and color transformations to images and annotations, including a specialized collection of 3D operations for volumetric data used in medical and scientific imaging. The library functions as an image mask and bounding box transformer, automatically updating masks, bounding boxes, and keypoints when images undergo geometric changes. This ensures that spatial alterations remain synchronized across images and their assoc

    Python
    Auf GitHub ansehen↗15,308
  • bentoml/bentomlAvatar von bentoml

    bentoml/BentoML

    8,456Auf GitHub ansehen↗

    BentoML is a machine learning model serving framework and GPU-accelerated inference server designed to package, deploy, and scale AI models as production-ready REST APIs. It functions as an AI model lifecycle manager and an inference graph orchestrator, enabling the chaining of multiple models and custom logic into complex pipelines for advanced task sequences. The framework distinguishes itself through a dynamic batching engine that optimizes GPU throughput and an artifact-based packaging system that bundles model weights and dependencies into immutable archives for consistent deployment. It

    Pythonai-inferencedeep-learninggenerative-ai
    Auf GitHub ansehen↗8,456
  • benedekrozemberczki/shapleyAvatar von benedekrozemberczki

    benedekrozemberczki/shapley

    226Auf GitHub ansehen↗

    The official implementation of "The Shapley Value of Classifiers in Ensemble Games" (CIKM 2021).

    Python
    Auf GitHub ansehen↗226
  • hi-primus/optimusAvatar von hi-primus

    hi-primus/optimus

    1,534Auf GitHub ansehen↗

    :truck: Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark

    Python
    Auf GitHub ansehen↗1,534