awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to hi-primus/optimus

Open-source alternatives to Optimus

30 open-source projects similar to hi-primus/optimus, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Optimus alternative.

  • cleanlab/cleanlabAvatar de cleanlab

    cleanlab/cleanlab

    11,513Ver en GitHub↗

    Cleanlab is a data-centric AI library and toolkit designed to improve machine learning model performance by detecting label errors and increasing overall dataset quality. It implements a confident learning framework that iteratively refines label noise estimates by comparing model predictions with estimated label probabilities to identify mislabeled examples. The project provides specialized utilities for active learning optimization, allowing for the selection of the most impactful examples for labeling or re-labeling. It also includes an outlier detection tool to identify atypical data poin

    Pythonactive-learningannotationanomaly-detection
    Ver en GitHub↗11,513
  • albumentations-team/albumentationsAvatar de albumentations-team

    albumentations-team/albumentations

    15,308Ver en GitHub↗

    Albumentations is a computer vision image augmentation library designed to increase training data diversity for deep learning models. It provides a toolset for applying geometric and color transformations to images and annotations, including a specialized collection of 3D operations for volumetric data used in medical and scientific imaging. The library functions as an image mask and bounding box transformer, automatically updating masks, bounding boxes, and keypoints when images undergo geometric changes. This ensures that spatial alterations remain synchronized across images and their assoc

    Python
    Ver en GitHub↗15,308
  • alteryx/featuretoolsAvatar de alteryx

    alteryx/featuretools

    7,658Ver en GitHub↗

    Featuretools is an automated feature engineering library and data transformation framework written in Python. It automatically generates machine learning feature vectors from multi-table datasets by applying synthesis patterns to relational and timestamped data. The system functions as a distributed feature synthesis engine, allowing the process of creating feature vectors to scale across multiple cores or clusters to handle large-scale datasets. The library supports the synthesis of multi-table datasets, time series feature generation, and the creation of custom machine learning primitives

    Python
    Ver en GitHub↗7,658

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Find more with AI search
  • adrotog/pandasguiAvatar de adrotog

    adrotog/PandasGUI

    3,259Ver en GitHub↗

    A GUI for Pandas DataFrames

    Python
    Ver en GitHub↗3,259
  • astrazeneca/rexmexAvatar de AstraZeneca

    AstraZeneca/rexmex

    278Ver en GitHub↗

    A general purpose recommender metrics library for fair evaluation.

    Python
    Ver en GitHub↗278
  • towhee-io/towheeAvatar de towhee-io

    towhee-io/towhee

    3,447Ver en GitHub↗

    Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.

    Pythoncomputer-visionconvolutional-networksembedding-vectors
    Ver en GitHub↗3,447
  • astrazeneca/chemicalxAvatar de AstraZeneca

    AstraZeneca/chemicalx

    781Ver en GitHub↗

    A PyTorch and TorchDrug based deep learning library for drug pair scoring. (KDD 2022)

    Python
    Ver en GitHub↗781
  • linealabs/lineapyAvatar de LineaLabs

    LineaLabs/lineapy

    670Ver en GitHub↗

    Move fast from data science prototype to pipeline. Capture, analyze, and transform messy notebooks into data pipelines with just two lines of code.

    Jupyter Notebook
    Ver en GitHub↗670
  • benedekrozemberczki/karateclubAvatar de benedekrozemberczki

    benedekrozemberczki/karateclub

    2,284Ver en GitHub↗

    Karate Club: An API Oriented Open-source Python Framework for Unsupervised Learning on Graphs (CIKM 2020)

    Python
    Ver en GitHub↗2,284
  • benedekrozemberczki/littleballoffurAvatar de benedekrozemberczki

    benedekrozemberczki/littleballoffur

    715Ver en GitHub↗

    Little Ball of Fur - A graph sampling extension library for NetworKit and NetworkX (CIKM 2020)

    Python
    Ver en GitHub↗715
  • asavinov/lambdoAvatar de asavinov

    asavinov/lambdo

    26Ver en GitHub↗

    Feature engineering and machine learning: together at last!

    Python
    Ver en GitHub↗26
  • hydrospheredata/mistAvatar de Hydrospheredata

    Hydrospheredata/mist

    324Ver en GitHub↗

    Serverless proxy for Spark cluster

    Scala
    Ver en GitHub↗324
  • iterative/cmlAvatar de iterative

    iterative/cml

    4,178Ver en GitHub↗

    CML is a pipeline automation tool for training and evaluating machine learning models, functioning as a CI/CD system for machine learning. It serves as a cloud compute orchestrator and Git-based workflow manager that automates model training cycles through branch management, automated commits, and integrated reporting. The project distinguishes itself by provisioning ephemeral cloud instances or Kubernetes nodes to provide specialized hardware for compute-heavy tasks. It also manages remote compute runners, allowing the connection of self-hosted GPU clusters or on-premise machines to execute

    JavaScript
    Ver en GitHub↗4,178
  • iterative/mlemAvatar de iterative

    iterative/mlem

    718Ver en GitHub↗

    🐶 A tool to package, serve, and deploy any ML model on any platform. Archived to be resurrected one day🤞

    Python
    Ver en GitHub↗718
  • benedekrozemberczki/shapleyAvatar de benedekrozemberczki

    benedekrozemberczki/shapley

    226Ver en GitHub↗

    The official implementation of "The Shapley Value of Classifiers in Ensemble Games" (CIKM 2021).

    Python
    Ver en GitHub↗226
  • minerva-ml/steppyAvatar de minerva-ml

    minerva-ml/steppy

    136Ver en GitHub↗

    Lightweight, Python library for fast and reproducible experimentation :microscope:

    Python
    Ver en GitHub↗136
  • ml-tooling/ml-workspaceAvatar de ml-tooling

    ml-tooling/ml-workspace

    3,540Ver en GitHub↗

    🛠 All-in-one web-based IDE specialized for machine learning and data science.

    Jupyter Notebook
    Ver en GitHub↗3,540
  • julialang/ijulia.jlAvatar de JuliaLang

    JuliaLang/IJulia.jl

    2,889Ver en GitHub↗

    Julia kernel for Jupyter

    Julia
    Ver en GitHub↗2,889
  • dslp/dslp-repo-templateAvatar de dslp

    dslp/dslp-repo-template

    202Ver en GitHub↗

    Template repository for data science lifecycle project

    Python
    Ver en GitHub↗202
  • kevinschaich/pyspark-cheatsheetAvatar de kevinschaich

    kevinschaich/pyspark-cheatsheet

    689Ver en GitHub↗

    🐍 Quick reference guide to common patterns & functions in PySpark.

    Ver en GitHub↗689
  • khangich/machine-learning-interviewAvatar de khangich

    khangich/machine-learning-interview

    12,624Ver en GitHub↗

    This project is a curated collection of technical reference materials and study guides designed for machine learning interview preparation. It provides comprehensive resources for candidates pursuing engineering roles, focusing on deep learning, production infrastructure, and large-scale system design. The repository distinguishes itself through an architecture that combines theoretical research with industrial case studies. It utilizes a pattern-based approach to system design, breaking down complex deployments—such as recommendation engines, search ranking, and ad click prediction—into reus

    Ver en GitHub↗12,624
  • jinglescode/python-signal-processingAvatar de jinglescode

    jinglescode/python-signal-processing

    85Ver en GitHub↗

    splearn: package for signal processing and machine learning with Python. Contains tutorials on understanding and applying signal processing.

    Jupyter Notebook
    Ver en GitHub↗85
  • dslp/dslpAvatar de dslp

    dslp/dslp

    527Ver en GitHub↗

    The Data Science Lifecycle Process is a process for taking data science teams from Idea to Value repeatedly and sustainably. The process is documented in this repo.

    Ver en GitHub↗527
  • handcraftsman/geneticalgorithmswithpythonAvatar de handcraftsman

    handcraftsman/GeneticAlgorithmsWithPython

    1,255Ver en GitHub↗

    source code from the book Genetic Algorithms with Python by Clinton Sheppard

    Python
    Ver en GitHub↗1,255
  • jacopotagliabue/mlsys-nyu-2022Avatar de jacopotagliabue

    jacopotagliabue/MLSys-NYU-2022

    558Ver en GitHub↗

    Slides, scripts and materials for the Machine Learning in Finance Course at NYU Tandon, 2022

    Jupyter Notebook
    Ver en GitHub↗558
  • comet-ml/comet-llmAvatar de comet-ml

    comet-ml/comet-llm

    19,673Ver en GitHub↗

    Comet LLM is an observability platform and evaluation framework designed for large language model applications and agentic workflows. It functions as a system for tracing, monitoring, and debugging execution flows while providing tools for prompt optimization and the enforcement of AI safety guardrails. The platform distinguishes itself through a combination of model-based scoring and heuristic metrics to quantify output quality and detect hallucinations. It includes a dedicated prompt and agent optimizer with an interactive playground for refining templates and tool configurations. For retri

    Python
    Ver en GitHub↗19,673
  • comet-ml/comet-examplesAvatar de comet-ml

    comet-ml/comet-examples

    174Ver en GitHub↗

    Examples of Machine Learning code using Comet.ml

    Jupyter Notebook
    Ver en GitHub↗174
  • iterative/dvcAvatar de iterative

    iterative/dvc

    15,680Ver en GitHub↗

    DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache. The project serves as an experiment tracker and remote storage synchronizer, enabling the execution and comparison of machine learning iterations based on hyperparameters and performance metrics. It provides a bridge for pushing and pulling these large data artifacts between local environments and cloud or on-premi

    Python
    Ver en GitHub↗15,680
  • jadianes/data-science-your-wayAvatar de jadianes

    jadianes/data-science-your-way

    616Ver en GitHub↗

    Ways of doing Data Science Engineering and Machine Learning in R and Python

    Jupyter Notebook
    Ver en GitHub↗616
  • minerva-ml/steppy-toolkitAvatar de minerva-ml

    minerva-ml/steppy-toolkit

    23Ver en GitHub↗

    Curated set of transformers that make your work with steppy faster and more effective :telescope:

    Python
    Ver en GitHub↗23